arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7302 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7302 篇

2410.13860 2024-10-18 cs.CV cs.RO 90%

VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Runsen Xu, Zhiwei Huang, Tai Wang, Yilun Chen, Jiangmiao Pang, Dahua Lin

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(title,abstract);vision-language model(abstract);分类 cs.CV

Comments CoRL 2024 Camera Ready. 25 pages. A novel zero-shot 3D visual grounding framework based solely on 2D images

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04933 2026-08-06 cs.RO 新提交 89%

Mimir: A Neuro-Symbolic Memory System with Dynamic Grounding for Embodied Agents in Interactive Environments

Mimir:面向交互环境中具身智能体的、具备动态 grounding 的神经符号记忆系统

Haoming Xu, Zhenlin He, Hengyi Wang, Jiafeng Xu, Hao Dong

专题命中 视觉定位与Grounding :grounding(title,title_cn)

AI总结 本文提出Mimir,一种分离世界与任务记忆并具备动态grounding的神经符号记忆系统,在EB-ALFRED、EB-Habitat等具身任务中显著提升了智能体的长程执行成功率。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28273 2026-06-29 cs.CL 新提交 89%

Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models

视觉默认,先验覆盖:视觉-语言模型中感知-知识冲突的因果机制

Niclas Lietzow, Danielle Bitterman, Carsten Eickhoff, William Rudman, Michal Golovanevsky

机构 * University of Tübingen(图宾根大学) Harvard University(哈佛大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(title,abstract);grounding(abstract)

AI总结 通过激活修补和消融实验,发现VLM中视觉默认激活,而先验知识依赖少量因果注意力头(2.5-4.8%),形成不对称因果结构。

Comments 14 pages, 11 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29585 2026-05-29 cs.CL 89%

World Models in Words: Auditing Physical State-Transition Commitments in Vision-Language Models

语言中的世界模型:审计视觉语言模型中的物理状态转换承诺

Emmanuelle Bourigault

机构 * University of Oxford(牛津大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(title,abstract);grounding(abstract)

AI总结 提出WMW框架,通过要求VLM输出结构化轨迹(初始状态、状态转换、结果状态和答案)并利用混合验证器检查模式有效性、状态基础、转换一致性和答案-轨迹兼容性,揭示仅评估最终答案所隐藏的物理推理失败。

Comments 8 pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15281 2026-05-21 cs.SE 89%

Semantic Grounding of Digital Twin Metamodels Using RDF Graphs

基于 RDF 图的数字孪生元模型语义 grounding

Faima Abbasi, Jean-Sébastien Sottet, Cedric Pruski

专题命中 视觉定位与Grounding :grounding(title,title_cn)

AI总结 本文提出了一种基于 RDF 图的数字孪生元模型语义 grounding 方法,通过设计多层数字孪生模型、将元模型提升为 RDF 图以及图基对齐方法 SSM-OM,实现了多层数字孪生的语义一致性与互操作性。

Comments Submitted to Conference, 15 pages excluding references, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03417 2026-05-15 cs.CL 89%

FactNet: A Billion-Scale Knowledge Graph for Multilingual Factual Grounding

FactNet:一个十亿级的知识图谱用于多语言事实 grounding

Yingli Shen, Wen Lai, Jie Zhou, Xueren Zhang, Yudong Wang, Kangyang Luo, Shuo Wang, Ge Gao, Alexander Fraser, Maosong Sun

机构 * Tsinghua University(清华大学) Technical University of Munich(慕尼黑技术大学) ModelBest Inc.(ModelBest公司) Minzu University of China(民族大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn)

AI总结 FactNet通过结合17亿条维基数据语句和301亿条证据指针,构建了一个十亿级多语言知识图谱,提供事实 grounding 的评估框架,并验证了跨语言结构的知识迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02791 2026-04-28 cs.CL 89%

Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension

使对话 grounding 数据丰富:一种三层数据合成框架用于通用指称表达理解

Juexi Shao, Siyou Li, Yujian Gan, Chris Madge, Vanja Karan, Massimo Poesio

机构 * Queen Mary University of London(伦敦玛丽女王大学) Queen's University Belfast(贝尔法斯特女王大学) University of Vienna(维也纳大学) Utrecht University(乌得勒支大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn)

AI总结 本文提出三层数据合成框架,通过平衡真实性和可控性,生成可扩展的对话条件 grounding 监督,提升指称表达理解的性能。

Journal ref ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2026, pp. 18142-18146

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12781 2026-08-14 cs.CV 新提交 89%

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

超越正确性:混合思维多模态大语言模型(MLLM)的响应行为基准测试与对齐

Xinming Wang, Weinong Wang, Hongming Yang, Yansong Lin, Zheng Ruan, Shangpin Peng, Qiming Peng, Nan Qiao, Fengyuan Lu, Guoqing Ma, Marito Li, Songyang Zhang, Saiyong Yang, Han Hu, Yonglong Tian, Xu-Yao Zhang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Large Language Model Department, Tencent(腾讯大语言模型部) University of Electronic Science and Technology of China(电子科技大学) Hong Kong University of Science and Technology(香港科技大学) Zhongguancun Academy(中关村学院)

专题命中 视觉定位与Grounding :MLLM(title_cn,summary_cn);grounding(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 该研究针对混合思维 MLLM 的思维与非思维模式响应错位问题,构建 PatternEval 基准并开发 PatternRL 方法,可减轻跨模式错位且任务性能损失极小。

Comments 8 tables and 6figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03631 2026-08-05 cs.CV 新提交 89%

SEER: A Self-Grounded Evidence Interface for Controlled Spatial Relation Classification

SEER:用于受控空间关系分类的自 grounding 证据接口

Feixiang Liu, Likun Wang, Qiang Qiu, Hui Xu, Huawei Shen, Xueqi Cheng

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);grounding(title_cn,abstract);分类 cs.CV

AI总结 本文提出针对冻结VLM的无训练推理时证据接口SEER,通过构建查询特定证据缓解空间关系分类错误,在多数据集上取得显著性能提升,证明该干预措施的有效性。

Comments 23 pages total, 2 figures. Code: https://github.com/SouthWinter/SEER

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18373 2026-04-14 cs.CV 89%

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models

MASS:面向视觉-语言模型中物理推理与理解的运动感知空间-时间 grounding

Xiyang Wu, Zongxia Li, Jihui Jin, Guangyao Shi, Gouthaman KV, Vishnu Raj, Nilotpal Sinha, Jingxi Chen, Fan Du, Dinesh Manocha

机构 * University of Maryland(马里兰大学) Dolby Laboratories(杜比实验室) University of Southern California(南加州大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(title);vision language model(abstract);VLM(abstract)

AI总结 MASS通过引入空间-时间信号提升视觉-语言模型对物理现象的理解能力,提出MASS-Bench基准测试集并验证模型在物理推理中的优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23067 2026-03-26 cs.CV 89%

MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding

MLLM-HWSI: 一种用于分层全滑动图像理解的多模态大语言模型

Basit Alawode, Arif Mahmood, Muaz Khalifa Al-Radi, Shahad Albastaki, Asim Khan, Muhammad Bilal, Moshira Ali Abdalla, Mohammed Bennamoun, Sajid Javed

机构 * Department of Computer Science, Khalifa University of Science and Technology(卡利法科技大学计算机科学系) Information Technology University(信息技术大学) KAU(卡乌大学) University of the Western Australia(西澳大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(title,abstract);grounding(abstract);分类 cs.CV

AI总结 本文提出MLLM-HWSI,一种分层全滑动图像级多模态大语言模型,通过四级尺度对齐视觉特征与病理语言,提升解释性证据接地推理能力,在六个CPath任务上取得新SOTA结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06965 2026-03-13 cs.CV 89%

MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images

MedMO:为医学图像构建和理解多模态大语言模型

Ankan Deria, Komal Kumar, Adinath Madhavrao Dukre, Eran Segal, Salman Khan, Imran Razzak

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 MedMO是一种基于通用MLLM架构构建的医学多模态基础模型,通过多阶段训练提升跨模态和任务的性能,超越现有开源基线,在医学图像识别和报告生成中取得显著提升。

Comments 21 pages, 6 figures and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02329 2026-03-04 cs.CV 89%

HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding

HAMMER: 通过跨模态整合利用大语言模型进行意图驱动的3D affordance grounding

Lei Yao, Yong Chen, Yuejiao Su, Yi Wang, Moyun Liu, Lap-Pui Chau

机构 * The Hong Kong Polytechnic University(香港理工大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 HAMMER通过跨模态整合多模态大语言模型,实现意图驱动的3D affordance grounding,提升3D表示的准确性和鲁棒性。

Comments Accepted by CVPR 2026. Project Page: https://rayyoh.github.io/Hammer

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20794 2026-02-25 cs.CV 89%

VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving

VGGDrive: 通过跨视角几何 grounding 为自动驾驶赋能 Vision-Language 模型

Jie Wang, Guang Li, Zhijian Huang, Chenxu Dang, Hangjun Ye, Yahong Han, Long Chen

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) Xiaomi EV(小米汽车)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract);分类 cs.CV

AI总结 VGGDrive通过引入跨视角几何 grounding 机制,提升Vision-Language模型在自动驾驶任务中的性能表现。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11904 2025-05-13 cs.CV 89%

GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding

Yue Zhou, Mengcheng Lan, Xiang Li, Litong Feng, Yiping Ke, Xue Jiang, Qingyun Li, Xue Yang, Wayne Zhang

机构 * Nanyang Technological University(南洋理工大学) University of Reading(阅读大学) Shanghai Jiao Tong University(上海交通大学) Harbin Institute of Technology(哈尔滨工业大学) SenseTime Research(商汤科技研究院)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract);分类 cs.CV

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13983 2025-04-14 cs.CV 89%

SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability

Jiankang Wang, Zhihan Zhang, Zhihang Liu, Yang Li, Jiannan Ge, Hongtao Xie, Yongdong Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19325 2025-03-14 cs.CV 89%

GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

Muhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Kartik Kuckreja, Fahad Shahbaz Khan, Paolo Fraccaro, Alexandre Lacoste, Salman Khan

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(title,abstract);LLaVA(abstract);分类 cs.CV

Comments This updated version includes revisions and additional analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10419 2025-03-14 cs.RO cs.AI 89%

HiFi-CS: Towards Open Vocabulary Visual Grounding For Robotic Grasping Using Vision-Language Models

Vineet Bhat, Prashanth Krishnamurthy, Ramesh Karri, Farshad Khorrami

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12694 2025-02-19 cs.CV cs.CL 89%

VividMed: Vision Language Model with Versatile Visual Grounding for Medicine

Lingxiao Luo, Bingda Tang, Xuanzhong Chen, Rong Han, Ting Chen

专题命中 视觉定位与Grounding :vision language model(title,abstract);grounding(title,abstract);visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10840 2024-12-17 cs.CV 89%

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning

Hai-Ming Xu, Qi Chen, Lei Wang, Lingqiao Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14901 2024-11-25 cs.CV cs.CL 89%

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos

Tanveer Hannan, Md Mohaiminul Islam, Jindong Gu, Thomas Seidl, Gedas Bertasius

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14492 2024-06-21 cs.CV cs.CL 89%

Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?

Gregor Geigle, Radu Timofte, Goran Glavaš

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12537 2023-08-25 cs.RO cs.CV 89%

HuBo-VLM: Unified Vision-Language Model designed for HUman roBOt interaction tasks

Zichao Dong, Weikun Zhang, Xufeng Huang, Hang Ji, Xin Zhan, Junbo Chen

专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(title);vision language model(abstract);grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.14824 2023-07-14 cs.CL cs.CV 89%

Kosmos-2: Grounding Multimodal Large Language Models to the World

Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, Furu Wei

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15517 2026-07-20 cs.CV cs.AI 新提交 89%

SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification

SLAPBench:用于四指SLAP指纹验证的多模态大语言模型基准测试

Bibesh Pyakurel, M. G. Sarwar Murshed

机构 * University of Wisconsin–Green Bay(威斯康星大学格林湾分校)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);multimodal large language model(title,abstract);分类 cs.CV、cs.AI

AI总结 研究四指SLAP指纹验证,介绍SLAPBench基准,评估多个MLLM在不同提示下的表现,发现提示控制崩溃,模型能力控制歧视,建立了特定于SLAP的MLLM基线,揭示了模型在指纹验证中的能力差距和公平性问题。

Comments 19 pages, 6 figures, 2 tables. Includes appendix with supporting figures and per-subgroup fairness detail. Code and data: https://github.com/bibeshpyakurel/SLAPBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03553 2026-06-04 cs.CV cs.AI 89%

Dynamic Content Moderation in Livestreams: Combining Supervised Classification with MLLM-Boosted Similarity Matching

直播中的动态内容审核:结合监督分类与MLLM增强的相似度匹配

Wei Chee Yew, Hailun Xu, Sanjay Saha, Xiaotian Fan, Hiok Hian Ong, David Yuchen Wang, Kanchan Sarkar, Zhenheng Yang, Danhui Guan

机构 * TikTok Singapore Singapore(TikTok新加坡) TikTok San Jose United States(TikTok旧金山美国) TikTok Shanghai China(TikTok上海中国)

专题命中 视觉定位与Grounding :MLLM(title,title_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 提出一种混合审核框架,结合监督分类和基于参考的相似度匹配,利用多模态大语言模型提升准确性,在保持轻量推理的同时实现大规模直播内容审核。

Comments To be published at KDD 2026 (ADS track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23797 2026-05-25 cs.LG cs.CV 89%

Debiased Negative Mining Improves Out-of-distribution Detection with Pre-trained Vision-Language Models

去偏负挖掘提升基于预训练视觉语言模型的分布外检测

Bo Peng, Jie Lu, Guangquan Zhang, Zhen Fang

机构 * University of Technology Sydney(悉尼科技大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(title,abstract);分类 cs.CV、cs.LG

AI总结 针对分布外检测中负标签的假阴性问题,提出通过间接近似负标签分布来校正采样偏差的理论框架,并转化为基于ID标签和未标注语料数据的蒙特卡洛采样方法,在多种OOD检测设置中达到新最优。

Comments KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09879 2026-01-16 cs.CV cs.AI 89%

MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation

MedVL-SAM2:一种统一的3D医学视觉-语言模型,用于多模态推理和基于提示的分割

Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, Kuang Gong

机构 * Department of Biomedical Engineering, University of Florida(佛罗里达大学生物医学工程系) Department of Radiology, University of Florida(佛罗里达大学放射学系) Research Computing, University of Florida(佛罗里达大学研究计算中心) Department of Medicine, University of Florida(佛罗里达大学医学系) Department of Radiology, UC San Francisco(旧金山大学放射学系)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);visual reasoning(abstract);visual question answering(abstract)

AI总结 MedVL-SAM2是一种统一的3D医学多模态模型,通过联合训练实现报告生成、VQA和多任务分割的高性能表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19130 2026-05-20 cs.LG cs.AI cs.CL cs.CV 89%

EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data

EgoBabyVLM:基于自然主义第一人称视频数据的跨模态学习基准测试

Dongyan Lin, Phillip Rust, Angel Villar Corrales, Alvin W. M. Tan, Mahi Luthra, Charles-Éric Saint-James, Rashel Moritz, Sheila Krogh-Jespersen, Vanessa Stark, Surya Parimi, Jiayi Shen, Youssef Benchekroun, Yosuke Higuchi, Martin Gleize, Tom Fizycki, Nicolas Hamilakis, Manel Khentout, Sho Tsuji, Balázs Kégl, Juan Pino, Michael C. Frank, Emmanuel Dupoux

机构 * Meta Superintelligence Labs(Meta超智能实验室) Stanford University(斯坦福大学) Meta Reality Labs(Meta现实实验室) The University of Tokyo(东京大学)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 研究探讨了儿童如何从有限的视觉-语言输入中获得语言 grounding 的鲁棒性,提出了 EgoBabyVLM 挑战,推动模型在自然主义数据中实现 grounded language learning。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14732 2026-04-15 cs.LG cs.AI cs.CV eess.IV 89%

INFORM-CT: INtegrating LLMs and VLMs FOR Incidental Findings Management in Abdominal CT

INFORM-CT:整合LLM和VLM用于腹部CT的偶发发现管理

Idan Tankel, Nir Mazor, Rafi Brada, Christina LeBedis, Guy ben-Yosef

机构 * GE Healthcare Technology and Innovation Center(GE医疗技术与创新中心) Boston Medical Center(波士顿医疗中心)

专题命中 视觉定位与Grounding :VLM(title_cn,summary_cn);vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出基于LLM和VLM的计划-执行框架,用于提高腹部CT偶发发现的检测、分类和报告效率与精度,通过自动化流程提升临床应用效果。

Comments Accepted for Spotlight presentation at MIDL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏