arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7302 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7302 篇

2508.09456 2026-06-02 cs.CV cs.CL cs.CR 95%

IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding

IAG: 基于输入感知的后门攻击针对VLM视觉定位

Junxian Li, Beining Xu, Simin Chen, Jiatong Li, Jingdi Lei, Haodong Zhao, Di Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学) Columbia University(哥伦比亚大学) Hong Kong Polytechnic University(香港理工大学) Nanyang Technological University(南洋理工大学)

专题命中 视觉定位与Grounding :VLM(title,title_cn);grounding(title,abstract);LLaVA(abstract,abstract_cn);InternVL(abstract,abstract_cn)

AI总结 提出IAG方法,通过文本条件UNet动态生成输入感知的触发器,实现首个多目标后门攻击VLM视觉定位,在多个模型和基准上达到最佳攻击成功率且不影响正常性能。

Comments Accepted by CVPR 2026; Code is at https://github.com/lijunxian111/IAG

Journal ref https://openaccess.thecvf.com/content/CVPR2026/papers/Li_IAG_Input-aware_Backdoor_Attack_on_VLM-based_Visual_Grounding_CVPR_2026_paper.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01535 2026-08-04 cs.CV cs.RO 新提交 94%

STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision

STAR-VLM:基于汽车雷达监督的时空视觉语言模型,用于运动与速度估计

Pou-Chun Kung, Aryaman Rao, Utkrisht Sahai, Hemanth Murali, Yi Liu, Rui-Yu Lin, Katherine A. Skinner

机构 * University of Michigan(密歇根大学)

专题命中 视觉定位与Grounding :VLM(title,title_cn);vision-language model(title,abstract);grounding(title);分类 cs.CV

AI总结 该研究提出STAR-VLM框架,利用汽车雷达监督提升时空视觉语言模型的运动推理与度量速度估计能力,在驾驶场景相关任务上实现了优于特定任务方法的最先进性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00518 2026-08-04 cs.CV 新提交 93%

GuideGround: VLM-guided Semantic Understanding and Viewpoint-aware Reasoning for 3D Visual Grounding

GuideGround:用于3D视觉定位的VLM引导语义理解与视点感知推理

Yiwen Wang, Yuyang Deng, Yihao Long, Xi Zhao

专题命中 视觉定位与Grounding :VLM(title,title_cn);grounding(title,abstract);vision-language model(abstract);分类 cs.CV

AI总结 GuideGround是一种VLM引导的3D视觉定位框架,用VLM生成的语义描述替代闭集分类、逐视图保留假设并经VLM验证,在ReferIt3D基准上性能优于现有SOTA。

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16461 2026-07-21 cs.CV 版本更新 93%

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models

GAP-MLLM:几何对齐预训练以激活多模态大语言模型中的3D空间感知

Jiaxin Zhang, Junjun Jiang, Haijie Li, Youyu Chen, Kui Jiang, Dave Zhenyu Chen

机构 * Harbin Institute of Technology(哈尔滨工业大学) School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院) Huawei(华为)

专题命中 视觉定位与Grounding :MLLM(title,title_cn);multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV

AI总结 本文提出GAP-MLLM,通过几何对齐预训练激活多模态大语言模型中的3D空间感知,改进了传统方法在3D空间感知上的不足。

Comments Accepted by ECCV 2026. Project page: https://gapmllm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20907 2026-07-02 cs.CV 版本更新 93%

PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding

PanoGrounder: 利用全景场景表示桥接2D和3D,实现基于VLM的3D视觉定位

Seongmin Jung, Seongho Choi, Gunwoo Jeon, Minsu Cho, Jongwoo Lim

机构 * Seoul National University(首尔大学) Robotics Lab, Hyundai Motor Company(现代汽车公司机器人实验室) Pohang University of Science and Technology (POSTECH)(浦项科技大学)

专题命中 视觉定位与Grounding :VLM(title,title_cn);grounding(title,abstract);vision-language model(abstract);分类 cs.CV

AI总结 提出PanoGrounder框架,通过多模态全景表示与预训练2D VLM结合,实现强泛化能力的3D视觉定位,在ScanRefer和Nr3D上取得最优结果。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26874 2026-06-26 cs.AI 新提交 93%

TAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation

TAVR-VLM:面向抗幻觉报告生成的风险条件因果基础

Zhixiang Lu, Xiwei Liu, Sifan Song, Changkai Ji, Anh Nguyen, Jionglong Su, Imran Razzak, Jinfeng Wang

机构 * Xi’an Jiaotong-Liverpool University(西交利物浦大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Shanghai Jiao Tong University(上海交通大学) University of Liverpool(利物浦大学) Kunming University of Science and Technology(昆明理工大学)

专题命中 视觉定位与Grounding :VLM(title,title_cn);grounding(title,abstract);multimodal large language model(abstract);分类 cs.AI

AI总结 提出TAVR-VLM框架,通过风险条件因果注意力(R-CGA)建立“风险→区域→词”的结构化基础路径,在TAVR规划中减少诊断幻觉,实现新最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18385 2026-06-18 cs.AI 新提交 93%

CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework

CaVe-VLM-CoT:一种可解释的视觉-语言模型框架

Sneha Rao, Shaina Raza, Dhanesh Ramachandram

机构 * Vector Institute(向量研究所)

专题命中 视觉定位与Grounding :VLM(title,title_cn);vision-language model(title,abstract);grounding(abstract);分类 cs.AI

AI总结 提出CaVe-VLM-CoT框架,通过五阶段闭环流水线(提取器、检索器、求解器、引用注入器、验证器)实现证据推理,并引入CaVeScore复合指标评估检索质量、引用忠实度和跨模态基础,在ScienceQA和MMMU上取得性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01850 2026-06-08 cs.CV cs.AI cs.LG cs.MM 版本更新 93%

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs

MoDA: 面向指令型多模态大语言模型的细粒度视觉定位的调制适配器

Wayner Barrios, Andrés Villa, Juan León Alcázar, SouYoung Jin, Bernard Ghanem

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);grounding(title,abstract);LLaVA(abstract,abstract_cn);visual question answering(abstract)

AI总结 提出MoDA调制适配器,通过指令引导的通道级乘法调制增强细粒度视觉定位,在12个基准上对三种MLLM架构取得一致提升,计算开销极小。

Comments Accepted at ICML 2026. Code is available at https://github.com/waybarrios/MoDA

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30506 2026-06-01 cs.RO cs.CV 92%

VLM-GLoc: Vision-Language Model Enhanced Monte Carlo Localization for Robust Semantic Global Localization in Cluttered Quasi-Static Environments

VLM-GLoc:视觉语言模型增强的蒙特卡洛定位,用于杂乱准静态环境中的鲁棒语义全局定位

Shivendra Agrawal, Bradley Hayes

机构 * University of Colorado Boulder(科罗拉多大学博尔德分校)

专题命中 视觉定位与Grounding :VLM(title,title_cn);vision-language model(title,abstract);分类 cs.CV

AI总结 提出VLM-GLoc方法,利用开放词汇视觉语言模型作为统一语义观测前端,通过逆语义提议机制和文本到地图检索,在几何模糊和语义歧义的准静态环境中实现鲁棒全局定位。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26104 2026-05-26 cs.CV 92%

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding

EVIDENT: 通过实体锚定的视觉证据路由MLLM适配用于跨域视频时间定位

Geo Ahn, Jiwook Han, Youngrae Kim, Joonseok Lee, Jinwoo Choi

机构 * Kyung Hee University(庆尚大学) University of Southern California(南加州大学) Seoul National University(首尔国立大学)

专题命中 视觉定位与Grounding :MLLM(title,title_cn);grounding(title,abstract);分类 cs.CV

AI总结 针对视频时间定位中域迁移导致性能下降的问题,提出EVIDENT框架,通过实体瓶颈适配器、实体绑定蒸馏损失和实体到证据门控机制,利用预训练MLLM的实体注意力实现参数高效的跨域鲁棒时间定位。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19859 2026-05-25 cs.CV 92%

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models

Eyes on VLM: 视觉语言模型中注视跟随与社会性注视预测的基准测试

Hengfei Wang, Anshul Gupta, Pierre Vuillecard, Jean-Marc Odobez

机构 * Idiap Research Institute(Idiap研究机构)

专题命中 视觉定位与Grounding :VLM(title,title_cn);vision language model(title);vision-language model(abstract);grounding(abstract)

AI总结 提出EyeVLM评估框架,通过零样本和微调方式测试视觉语言模型在注视跟随(几何视觉处理)和社会性注视预测(社交推理)两个核心任务上的能力,发现当前VLM缺乏精确的注视理解能力。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17070 2026-05-19 cs.CV 92%

EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models

EPIC-Bench: 一种以感知为中心的细粒度具身视觉 grounding 的基准

Haozhe Shan, Xiancong Ren, Han Dong, Haoyuan Shi, Yingji Zhang, Jiayu Hu, Yi Zhang, Yong Dai, Bin Shen, Lizhen Qu, Zenglin Xu, Xiaozhu Ju

机构 * X-Humanoid Fudan University(复旦大学) University of Science and Technology of China(中国科学技术大学) University of Manchester(曼彻斯特大学) Monash University(墨尔本大学) Celonis AI University of New South Wales(新南威尔士大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn);vision-language model(title,abstract);分类 cs.CV

AI总结 本文提出 EPIC-Bench,一种以感知为中心的细粒度具身视觉 grounding 基准,旨在系统评估 VLMs 在现实世界具身环境中的视觉感知能力。该基准包含 6.6k 个精心标注的元组(图像,文本,掩码),涵盖 23 个细粒度任务,涉及具身交互管道的三个核心阶段:目标定位、导航和操作。评估结果显示,尽管先进推理模型表现出潜力,但当前 VLMs 在复杂视觉-文本对齐方面普遍存在困难,特别是在多目标计数、部分-整体关系理解和 affordance 区域检测方面存在瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13530 2026-05-14 cs.CV cs.AI 92%

Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs

迈向统一的手术场景理解:通过多模态大语言模型弥合推理与 grounding

Jincai Huang, Shihao Zou, Yuchen Guo, Jingjing Li, Wei Ji, Kai Wang, Shanshan Wang, Weixin Si

机构 * Southern University of Science and Technology(南方科技大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Northwestern University(西北大学) University of Alberta(阿尔伯塔大学) Yale University(耶鲁大学) Nanfang Hospital(南华医院) Shenzhen University of Advanced Technology(深圳大学先进技术研究院)

专题命中 视觉定位与Grounding :grounding(title,title_cn);MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 本文提出SurgMLLM框架,通过统一推理与视觉 grounding 实现手术场景理解,提升三元组识别和分割精度,实验表明其在三元组识别指标AP_IVT上提升显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24396 2026-04-28 cs.CV cs.AI 92%

Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation

全局上下文还是局部细节?面向幻觉缓解的自适应视觉 grounding

Yubo Jiang, Xin Yang, Abudukelimu Wuerkaixi, Zheming Yuan, Xuxin Cheng, Fengying Xie, Zhiguo Jiang, Cao Liu, Ke Zeng, Haopeng Zhang

机构 * School of Astronautics, Beihang University(北航航天学院) Longcat Interaction Team, Meituan(美团Longcat交互团队) Tianmushan Laboratory, Beihang University(北航天门山实验室)

专题命中 视觉定位与Grounding :grounding(title,title_cn);VLM(abstract,abstract_cn);LLaVA(abstract,abstract_cn);InternVL(abstract,abstract_cn)

AI总结 本文提出PND框架,通过双路径对比在解码过程中增强视觉真实性,减少幻觉并提升描述细节,无需模型微调。

Comments 9 pages, 8 figures, Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26957 2026-07-03 cs.CY cs.AI 92%

Simulating Validity: Modal Decoupling in MLLM Generated Feedback on Science Drawings

模拟有效性:在MLLM生成的科学图表反馈中的模态解耦

Arne Bewersdorff, Nejla Yuruk, Xiaoming Zhai

机构 * University of Georgia, AI4STEM Education Center(佐治亚大学AI4STEM教育中心) Gazi University, Department of Mathematics and Science Education(加齐大学数学与科学教育系)

专题命中 视觉定位与Grounding :MLLM(title,title_cn);grounding(summary_cn,abstract);multimodal large language model(abstract);分类 cs.AI

AI总结 研究探讨了多模态大语言模型生成的科学图表反馈中模态解耦问题,发现反馈常存在 grounding 失败,表明需超越常规提示策略的 grounding 机制。

Comments Accepted as AIED Short Paper 2026, Seoul, South Korea. Submission #1147. This is the long paper version

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20306 2026-06-03 cs.CV cs.LG 92%

WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents

WildRoadBench: 面向视觉语言模型与自主智能体的野外航拍道路损伤定位基准

Bingnan Liu, Chenhang Cui, Rui Huang, Jiani Luo, Zhirong Shen, Tinghao Wang, Xiande Huang, Lingbei Meng, Fei Shen, An Zhang

机构 * University of Electronic Science and Technology of China(电子科技大学) National University of Singapore(新加坡国立大学) De Artificial Intelligence Lab(德人工智能实验室) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Science and Technology of China(中国科学技术大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(title,abstract);grounding(title,abstract);分类 cs.CV、cs.LG

AI总结 提出WildRoadBench基准,通过VLM直接定位和LLM驱动智能体自主研究两种协议,评估模型在航拍道路损伤定位上的性能,发现现有方法在野外场景下仍不可靠。

Comments Preprint. Under review. 4 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08315 2026-08-11 cs.CV cs.MM 新提交 92%

Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No

你的视觉语言模型(VLM)已经知道时间:通过提问“是/否”实现无需训练的时间定位

Ji Huang, Barry Devereux, Hui Wang

专题命中 视觉定位与Grounding :VLM(title,title_cn);grounding(title,abstract);分类 cs.CV

AI总结 针对VLM时间定位任务的自信错误问题,提出无需训练的FV-Action方法,通过从粗到细的二元问题扫描替代时间戳回归,在多个基准数据集上大幅提升了时间定位性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00060 2026-07-02 cs.CV 新提交 92%

Synergistic Perception-Reasoning Governance: Grounding Medical MLLMs with Verifiable Anatomical Evidence

协同感知-推理治理:用可验证解剖证据为医学多模态大语言模型提供基础

Rui Hao, Qiankun Li, Junyuan Mao, Linghao Meng, Dirui Xie, Dayu Tan, Zhigang Zeng

机构 * Huazhong University of Science and Technology(华中科技大学) Imperial Global Singapore, Imperial College London(帝国理工学院新加坡全球中心) Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学) Anhui University(安徽大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);LLaVA(abstract,abstract_cn);InternVL(abstract,abstract_cn);MLLM(summary_cn)

AI总结 提出一种无需训练的证据注入框架,通过ROI引导的视觉激活调制和解剖坐标语义标记,协同校准视觉感知与文本推理,动态路由任务特定干预,有效减少医学MLLM的幻觉。

Comments Accepted by MICCAI 2026 (Early Accept, Top 9%)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10468 2026-06-10 cs.CV 新提交 92%

Geometric Coastline Localization using Vision-Language Models

基于视觉语言模型的海岸线几何定位

Rafia Malik, Bernhard Pfahringer, Karin Bryan, Mark Dickson, Eibe Frank

机构 * The University of Waikato(怀卡托大学) The University of Auckland(奥克兰大学)

专题命中 视觉定位与Grounding :LLaVA(summary_cn,abstract);vision-language model(title,abstract);VLM(abstract,abstract_cn);grounding(abstract)

AI总结 提出将海岸线提取视为几何边界定位任务,基于GeoChat-7B/LLaVA-1.5架构构建CoastlineVLM-7B模型,直接预测折线而非分割掩码,在几何指标上优于传统分割方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19976 2026-05-20 cs.CV 92%

RECIPE: Procedural Planning via Grounding in Instructional Video

RECIPE: 通过指令视频中的 grounding 实现过程规划

Luigi Seminara, Antonino Furnari, Lorenzo Torresani

机构 * Khoury College of Computer Sciences, Northeastern University, Boston(东北大学北斯托顿学院计算机科学学院) Department of Mathematics and Computer Science, University of Catania, Italy(卡塔尼亚大学数学与计算机科学系)

专题命中 视觉定位与Grounding :grounding(title,title_cn);VLM(abstract,abstract_cn);分类 cs.CV

AI总结 该研究提出RECIPE方法,通过利用指令视频中的grounding信息来改进过程规划任务,通过利用预计算的文本嵌入实现大规模视频数据的验证,从而提升规划的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11616 2026-05-13 cs.CV 92%

Grounding by Remembering: Cross-Scene and In-Scene Memory for 3D Functional Affordances

通过记忆实现 grounding:跨场景和场景内的记忆用于3D功能 affordances

Qirui Wang, Jingyi He, Yining Pan, Xulei Yang, Shijie Li

机构 * TUM(慕尼黑工业大学) A*STAR(新加坡科技研究局)

专题命中 视觉定位与Grounding :grounding(title,title_cn);VLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出AFFORDMEM框架,通过跨场景和场景内的记忆实现3D功能 affordances的grounding,无需微调和标注,提升AP50性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05673 2025-07-09 cs.CV 92%

R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding

Joonhyung Park, Peng Tang, Sagnik Das, Srikar Appalaraju, Kunwar Yashraj Singh, R. Manmatha, Shabnam Ghadar

机构 * KAIST(韩国科学技术院) AWS AI Labs(亚马逊人工智能实验室)

专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(title,abstract);grounding(title,abstract);分类 cs.CV

Comments ACL 2025; 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02269 2026-07-03 cs.CV cs.AI 新提交 91%

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

AnyGroundBench: 视觉语言模型中视频定位的专业领域基准

Rintaro Otsubo, Ryo Fujii, Reina Ishikawa, Taiki Kanaya, Kanta Sawafuji, Hiroki Kajita, Shigeki Sakai, Hideo Saito, Ryo Hachiuma

机构 * Keio University(庆应大学) Keio AI Research Center(庆应人工智能研究中心) Keio University School of Medicine(庆应大学医学部) NVIDIA(英伟达)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);VLM(summary_cn,abstract_cn);分类 cs.CV、cs.AI

AI总结 提出AnyGroundBench基准,将STVG评估从零样本转向领域适应,涵盖五个专业领域,评估15个VLM的零样本和上下文学习能力,发现现有模型在专业领域表现不佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21134 2026-04-24 cs.CL 91%

Beyond Pixels: Introspective and Interactive Grounding for Visualization Agents

超越像素:可视化代理的 introspective 和 interactive 地基

Yiyang Lu, Woong Shin, Ahmad Maroof Karimi, Feiyi Wang, Jie Ren, Evgenia Smirni

机构 * William & Mary(威廉玛丽学院) Oak Ridge National Laboratory(橡树岭国家实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(summary_cn,abstract);vision-language model(abstract,abstract_cn)

AI总结 本文提出IVG框架,结合规范引导的反思与视图引导的交互,解决VLM在可视化任务中的误读与歧义问题,通过iPlotBench验证,提升问答准确度至0.81,并展示在自主探索与实时协作中的能力。

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27902 2026-08-04 cs.CV 版本更新 91%

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting

一个补丁就够:面向基于多模态大语言模型(MLLM)的场景文本定位的强化优化视觉标记接地

Rui Tang, Wentao Yang, Peirong Zhang, Yongxin Shi, Shun Zhang, Huiguo He, Lianwen Jin

机构 * South China University of Technology(华南理工大学) HiThink Research(海思思考研究院)

专题命中 视觉定位与Grounding :MLLM(title,title_cn);grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 针对现有多补丁场景文本定位范式的冗余噪声与定位歧义问题,提出以视觉为中心的单补丁文本定位框架 SPaTS,通过强化学习优化的单补丁选择等技术实现性能提升,显著优于前沿相关模型。

Comments 15 pages, 11 figures. Accepted to ACM Multimedia 2026

Journal ref Proceedings of the 34th ACM International Conference on Multimedia (MM '26), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26250 2026-04-30 cs.CV 91%

Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning

超越捷径:通过定性推理缓解冻结VLM中的视觉错觉

Hao Guo, Fei Wang, Junjie Chen, Yiqi Nie, Jiaqi Zhao, Qiankun Li, Subin Huang

机构 * Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(人工智能研究院,合肥国家科学中心) Anhui Polytechnic University(安徽理工大学) Hefei University of Technology(合肥工业大学) Anhui University(安徽大学) IGS, Imperial College London(帝国理工学院伦敦分校)

专题命中 视觉定位与Grounding :VLM(title_cn,summary_cn);grounding(summary_cn,abstract);vision-language model(abstract);分类 cs.CV

AI总结 本文提出SQI框架,通过定性约束提升冻结VLM的视觉 grounding,解决视觉错觉问题,实验显示在DataCV 2026挑战中表现优异,提升准确率并提供更好的可解释性。

Comments 4 pages, 2 figures, and 1 table. This is a methodology paper for the DataCV 2026 Challenge (CVPR Workshops), Task 1, where our method ranked 2nd

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07886 2026-08-11 cs.CV cs.AI cs.CL 新提交 91%

Vision-Language Grounding as Bidirectional Concept Correspondence

视觉-语言 Grounding 作为双向概念对应

Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer, Ranjay Krishna

机构 * University of Washington(华盛顿大学) Allen Institute for AI(艾伦人工智能研究所) FAIR at Meta(Meta FAIR实验室)

专题命中 视觉定位与Grounding :grounding(title,title_cn);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本研究将视觉-语言 Grounding 建模为双向概念对应,提出 ConCor-1 模型统一相关任务,在长文本数据集和零样本 LVIS 上对应 F1 分别提升 48%、29%,性能优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06154 2026-08-07 cs.RO cs.AI cs.CV 新提交 91%

Visual Grounding in Zero-Shot Vision-Language Control

零样本视觉语言控制中的视觉 grounding

J. de Curtò, Dayani Plasencia, Diego Sánchez, I. de Zarzà

专题命中 视觉定位与Grounding :grounding(title,title_cn);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 本文针对零样本视觉语言控制中VLMs决策是否基于视觉输入的问题,通过多组消融实验分析了多种VLMs的表现,提出对称共识守护者方法,验证了VLMs可作为有界的危险助手。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06891 2026-06-08 cs.CV 新提交 91%

Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors

Stream3D-VLM:基于增量几何先验的在线3D空间理解

Hanxun Yu, Xuan Qu, Lei Ke, Boqiang Zhang, Yuxin Wang, Jianke Zhu, Dong Yu

机构 * Zhejiang University(浙江大学) Tencent Hunyuan(腾讯文汇) HKUST(香港科技大学) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 视觉定位与Grounding :VLM(title,title_cn);vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 提出在线3D视觉语言模型Stream3D-VLM,通过自回归流控制、轻量视觉-空间特征融合模块和几何自适应体素压缩,实现从流式视频中实时理解3D空间,并构建超百万在线3D问答数据集,在多项任务上超越现有模型。

Comments Project Page: https://stream3d-vlm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17253 2026-07-15 cs.AR 版本更新 91%

PDAGENT-BENCH: Characterizing, Grounding, and Architecting LLM/VLM Agents for VLSI Physical Design

PDAGENT-BENCH: 用于VLSI物理设计的LLM代理的特征化、基础化与架构化

Qiufeng Li, Rongqian Chen, Quan Cheng, Chengxuan Wang, Sizhe Tang, Chia-Tung Ho, David Z. Pan, Tian Lan, Weidong Cao

专题命中 视觉定位与Grounding :VLM(title,summary_cn);grounding(title);vision-language model(abstract)

AI总结 提出PDAGENT-BENCH基准,用于评估LLM/VLM代理在VLSI物理设计中的能力,涵盖任务级和工作流级评估,揭示模型在工具执行和长程推理上的局限,并验证人类技能增强工作流的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏