A full Stokes subgrid model for simulation of grounding line migration in ice sheets
专题命中 视觉定位与Grounding :grounding(title,abstract)
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :grounding(title,abstract)
专题命中 视觉定位与Grounding :grounding(title,abstract)
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments ACL 2019
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments 8 pages, 11 figures, 1 table
Journal ref IEEE Robotics and Automation Letters, vol. 4, no. 2, pp. 351-358, April 2019
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments Published in ICJAI 2017
专题命中 视觉定位与Grounding :grounding(title,abstract)
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments Paper presented at the 34nd International Conference on Logic Programming (ICLP 2018), Oxford, UK, July 14 to July 17, 2018 18 pages, LaTeX
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments Submitted to the Journal of Artificial Intelligence Research
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments 2 pages
专题命中 视觉定位与Grounding :grounding(title,abstract)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI、cs.LG
Comments 6 pages, 4 figures, ICDL-Epirob 2016
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments 10 pages, 7 figures, 3 tables. To appear in ACL-IJCNLP 2015
专题命中 视觉定位与Grounding :grounding(title,abstract)
Journal ref Journal of Artificial Intelligence Research, feb 2015, volume 52, pages 235-286
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments 4 pages, 11 figures
Journal ref Proceedings of the 27th Symposium On Fusion Technology (SOFT-27); Liege, Belgium, September 24-28, 2012. Fusion Engineering and Design, Vol.88, Issues 9-10, October 2013, p.2100-2104
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments 13 pages, 5 figures
Journal ref International Journal of Web & Semantic Technology (IJWesT) Vol.2, No.4, 2011, 67-79
专题命中 视觉定位与Grounding :grounding(title,abstract)
Comments expanded article
专题命中 视觉定位与Grounding :VLM(title,abstract)
Comments To appear in the April 10, 2003 issue of The Astrophysical Journal 30 pages, 17 figures
Journal ref Astrophys.J. 587 (2003) 407-422
专题命中 视觉定位与Grounding :grounding(title,comments);分类 cs.CV、cs.AI
Comments 3st Place in PIC Makeup Temporal Video Grounding (MTVG) Challenge in ACM-MM 2022
通过结构化奖励强化视频MLLMs的一致性
机构 * Rutgers University(罗格斯大学) ; University of Toronto(多伦多大学)
专题命中 视觉定位与Grounding :grounding(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 研究通过结构化奖励提升视频MLLMs的一致性,发现传统监督不足,提出结合事实和时间单元的奖励机制,提升视频理解准确性。
Comments Accepted by COLM 2026
增强视觉-语言能力的半监督医学图像分割基础模型
机构 * ECE, Northwestern University(电气工程与计算机科学系,西北大学) ; Stats, Northwestern University(统计学系,西北大学) ; Radiology, Northwestern University(放射学系,西北大学)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV
AI总结 本文提出VESSA模型,通过增强视觉-语言能力的半监督方法提升医学图像分割精度,实验表明其在有限标注条件下表现优于现有方法。
AmalthAI:面向文化遗产的开源计算机视觉平台
机构 * Democritus University of Thrace(德谟克利特色雷斯大学) ; Athena Research Center(雅典娜研究中心) ; National and Kapodistrian University of Athens(雅典国立卡波迪斯特里亚大学) ; University of Warsaw(华沙大学)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV
AI总结 AmalthAI是面向文化遗产领域专家的开源计算机视觉平台,通过集成Kubeflow、Katib等工具,支持数据集管理、模型训练与推理,可保障敏感考古数据安全,已在黏土织物印痕数据集上验证其功能。
Hunyuan3D-Buffalo 1.0:用于可扩展3D生成、理解与编辑的统一多模态模型
机构 * Tencent Hunyuan(腾讯混元)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);grounding(abstract_cn);分类 cs.CV
AI总结 Hunyuan3D-Buffalo 1.0是支持3D理解、文本到3D生成等功能的统一多模态模型,构建87M规模语料库训练,在相关基准上达SOTA或领先性能,验证了统一训练的有效性。
Comments Project Page: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0
作为辅助监督的生成:通过解耦嵌入预测以零推理开销增强视觉理解
机构 * ByteDance(字节跳动)
专题命中 视觉定位与Grounding :multimodal large language model(abstract,abstract_cn);grounding(abstract);分类 cs.CV
AI总结 本研究提出GAS框架,将视觉生成作为辅助监督,通过解耦MoT架构的NEP实现零推理开销,提升了多模态理解尤其是感知与空间理解能力。
ForgeryVCR: 通过高效的取证工具在MLLMs中实现视觉中心推理用于图像伪造检测与定位
机构 * Shenzhen University(深圳大学) ; Tencent Youtu Lab(腾讯优图实验室)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 ForgeryVCR通过高效的取证工具实现视觉中心推理,提升图像伪造检测与定位的性能。
RoboAtlas:上下文感知主动SLAM
机构 * Mitsubishi Electric Research Laboratories(三菱电机研究实验室) ; University of California, Los Angeles(加州大学洛杉矶分校)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV
AI总结 提出RoboAtlas框架,通过上下文感知的多臂赌博机平衡几何探索与语义推理,结合3D语义地图OpenRoboVox,在真实环境中实现100%任务成功率,并在GOAT-Bench上以90.6%成功率达到SOTA。
Comments Alexander Schperberg and Shivam K. Panda made equal contribution
TAU-Bench:从异常实例跟踪到细粒度视频异常理解
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV
AI总结 该研究引入TAU-Bench这一以跟踪为核心的基准,用于联合评估异常实例跟踪与细粒度视频异常理解,发现现有视觉-语言模型存在语义推理与视觉接地的差距,为构建更可靠的视频异常理解系统提供了关键评估方向。
SpatialCLI:先借助空间工具进行推理,再脱离工具进行推理
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.AI
AI总结 本研究提出SpatialCLI框架,通过三个阶段让视觉语言模型先借助专用空间工具推理,再内化感知能力,在MindCube上大幅提升了Qwen3-VL-8B-Instruct的性能,表现优于对比模型。
Tarot-SAM3: 无需训练的SAM3用于任意指称表达分割
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Nanyang Technological University(南洋理工大学)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 Tarot-SAM3通过引入推理辅助提示和自修复机制,实现对任意指称表达的高效分割,有效解决传统方法对长或隐含表达的处理难题。
Comments We need to make a huge revision
从像素到语义:一种多阶段AI框架用于卫星图像中的结构损坏检测
机构 * The Beacom College of Computer & Cyber Sciences, Dakota State University(达科他州立大学比康计算机与网络科学学院)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV
AI总结 本文提出一种多阶段AI框架,结合超分辨率、目标检测和视觉-语言模型,提升灾害后建筑损坏评估的语义解读能力,通过CLIPScore和多模型策略提高检测可靠性。
Journal ref IEEE/CVF Conference on Computer Vision & Pattern Recognition Workshop (CVPRW) 2026
DeceptionX: 基于多模态大语言模型的可解释欺骗检测
机构 * Great Bay University(大湾区大学) ; Hong Kong Polytechnic University(香港理工大学)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 提出DeceptionX框架,将欺骗检测从黑箱分类转变为可解释的观察-思考-总结推理过程,通过构建DeceptChain数据集和三阶段训练管道,在标准基准上超越现有方法,同时提供专家级可解释推理路径。