arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7314 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7314 篇

2509.10345 2025-09-16 cs.CV cs.AI 88%

Towards Understanding Visual Grounding in Visual Language Models

Georgios Pantazopoulos, Eda B. Özyiğit

机构 * The Alan Turing Institute(艾伦·图灵研究所) Heriot-Watt University(赫瑞-沃德大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual language model(title);vision language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20758 2025-08-29 cs.CV cs.AI 88%

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding

Jiawen Lin, Shiran Bian, Yihang Zhu, Wenbin Tan, Yachao Zhang, Yuan Xie, Yanyun Qu

机构 * School of Informatics, Xiamen University(厦门大学信息学院) School of Computer Science, Nanjing University(南京大学计算机科学学院) School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(教育部多媒体可信感知与高效计算重点实验室,厦门大学)

专题命中 视觉定位与Grounding :VLM(title,abstract);grounding(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13377 2025-07-01 cs.CV cs.AI cs.CL 88%

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Ye Wang, Ziheng Wang, Boshen Xu, Yang Du, Kejun Lin, Zihan Xiao, Zihao Yue, Jianzhong Ju, Liang Zhang, Dingyi Yang, Xiangnan Fang, Zewen He, Zhenbo Luo, Wenxuan Wang, Junqi Lin, Jian Luan, Qin Jin

机构 * AIM3 Lab, Renmin University of China(中国人民大学AIM3实验室) MiLM Plus, Xiaomi Inc.(小米公司MiLM Plus)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision language model(title);vision-language model(abstract);分类 cs.CV、cs.AI

Comments Project Page: https://xuboshen.github.io/Time-R1/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06287 2025-03-11 cs.CV cs.AI 88%

Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding

Seil Kang, Jinyeong Kim, Junhyeok Kim, Seong Jae Hwang

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05767 2025-02-19 cs.CL cs.AI cs.CV 88%

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

You Li, Heyu Huang, Chi Chen, Kaiyu Huang, Chao Huang, Zonghao Guo, Zhiyuan Liu, Jinan Xu, Yuhua Li, Ruixuan Li, Maosong Sun

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);分类 cs.CV、cs.AI

Comments 21 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18038 2024-11-28 cs.CV cs.AI 88%

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis

Donggoo Kang, Dasol Jeong, Hyunmin Lee, Sangwoo Park, Hasil Park, Sunkyu Kwon, Yeongjoon Kim, Joonki Paik

专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(title,abstract);分类 cs.CV、cs.AI

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23822 2024-11-01 cs.CV cs.AI 88%

Parameter-Efficient Fine-Tuning Medical Multimodal Large Language Models for Medical Visual Grounding

Jinlong He, Pengfei Li, Gang Liu, Shenjun Zhong

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06157 2024-07-09 cs.CV cs.AI 88%

Temporal Grounding of Activities using Multimodal Large Language Models

Young Chol Song

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19128 2024-05-01 cs.CV cs.CL cs.LG 88%

Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM

Navid Rajabi, Jana Kosecka

专题命中 视觉定位与Grounding :grounding(title,abstract);vision language model(title);VLM(abstract);分类 cs.CV、cs.LG

Comments Accepted to CVPR 2024, Second Workshop on Foundation Models (WFM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13435 2023-12-14 cs.CV cs.AI 88%

PG-Video-LLaVA: Pixel Grounding Large Video-Language Models

Shehan Munasinghe, Rusiru Thushara, Muhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Mubarak Shah, Fahad Khan

专题命中 视觉定位与Grounding :LLaVA(title,abstract);grounding(title,abstract);分类 cs.CV、cs.AI

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04635 2026-07-31 cs.RO 版本更新 88%

Relational Scene Graphs for Object Grounding of Natural Language Commands

基于关系场景图的对象接地自然语言指令

Julia Kuhn, Francesco Verdoja, Tsvetomila Mihaylova, Ville Kyrki

机构 * School of Electrical Engineering, Aalto University(艾尔沃大学电气工程学院) School of Science, Aalto University(艾尔沃大学科学学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(summary_cn,abstract_cn);vision language model(abstract)

AI总结 本文提出基于关系场景图的方法,通过结合LLM和VLM增强自然语言指令中对象接地的准确性。

Comments Accepted to the 35th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00020 2026-07-02 cs.RO 新提交 88%

EmbodimentSemantic: A Spatial Scene-Graph Dataset and Benchmark for Vision-Language Models on Embodied Manipulation Trajectories

EmbodimentSemantic:面向具身操作轨迹的空间场景图数据集与视觉语言模型基准

Hassan Jaber, Refinath S N, Luca Cagliero, Christopher E. Mower, Haitham Bou-Ammar

机构 * Politecnico di Torino(都灵理工大学) University College London(伦敦大学学院)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(title);grounding(abstract)

AI总结 提出EmbodimentSemantic数据集与基准,通过场景图三元组评估VLM在具身操作中的空间关系理解,发现现有模型在深度感知和视角依赖关系上存在不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19374 2026-05-20 cs.CV cs.AI cs.LG 88%

Concept-Guided Noisy Negative Suppression for Zero-Shot Classification and Grounding of Chest X-Ray Findings

基于概念的噪声负样本抑制用于零样本分类和胸片发现的 grounding

Chenyu Lian, Hong-Yu Zhou, Chun-Ka Wong, Jing Qin

机构 * The Center for Smart Health, School of Nursing, the Hong Kong Polytechnic University, Hong Kong, China(香港理工大学智能健康中心,护理学院,中国香港) Research Institute for Smart Ageing, the Hong Kong Polytechnic University, Hong Kong, China(香港理工大学智能老龄化研究 institute,中国香港) School of Biomedical Engineering, Tsinghua Medicine, Tsinghua University, Beijing, China(清华大学生物医学工程学院,清华大学,北京,中国) Queen Mary Hospital, LKS Faculty of Medicine, The University of Hong Kong, Hong Kong, China(香港大学李嘉诚医学院Queen Mary医院,中国香港)

专题命中 视觉定位与Grounding :grounding(title,title_cn);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出了一种基于概念的噪声负样本抑制框架CoNNS,通过构建层次化概念本体,解决不同患者间相似发现导致的噪声负样本问题,提升零样本理解任务的性能。

Comments Early accepted by MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09060 2026-05-12 cs.CL 88%

Language-Conditioned Visual Grounding with CLIP Multilingual

基于CLIP多语言的语言条件视觉 grounding

J. de Curtò, Mauro Liz, I. de Zarzà

机构 * 1 Department of Computer Applications in Science \& Engineering, BARCELONA Supercomputing Center, Barcelona, Spain 3 Department of Electrical Computer Engineering, Boston University, Boston, MA 02215, USA 4 Human centered AI, Data \& Software, LUXEMBOURG Institute of Science

专题命中 视觉定位与Grounding :grounding(title,title_cn);vision-language model(abstract)

AI总结 研究通过密集多语言CLIP探针分析多语言视觉语言模型的性能差异,发现低资源语言在文本分支存在结构性惩罚,编码器扩展能缓解部分失败案例,同时保持跨语言相似性,但空间对齐问题导致集群掩码IoU下降。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26521 2026-04-30 cs.AI cs.CV cs.LG cs.LO 88%

Grounding vs. Compositionality: On the Non-Complementarity of Reasoning in Neuro-Symbolic Systems

基础性与组成性:神经符号系统中推理非互补性研究

Mahnoor Shahid, Hannes Rothe

机构 * Place-Beyond-Bytes

专题命中 视觉定位与Grounding :grounding(title,summary_cn);分类 cs.CV、cs.AI、cs.LG

AI总结 研究指出神经符号系统中组成性推理并非符号 grounding 的副产品,通过 iLTN 实验证明仅依赖 grounding 无法实现泛化,联合训练 grounding 与推理可提升零样本准确性。

Comments Accepted at AAAI MAKE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15732 2026-08-17 cs.CV 版本更新 88%

IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models

IoUPD:用于多模态大语言模型视觉定位的IoU感知特权蒸馏

Xiuyuan Zhu, Ke Lu, Hao Wu, Siwen Jiao, Zijin Du, Dongming Zhang, Jian Xue

机构 * University of Chinese Academy of Sciences(中国科学院大学) State Key Laboratory of Communication Content Cognition(通信内容认知技术国家重点实验室) Peng Cheng Laboratory(鹏城实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);分类 cs.CV

AI总结 研究多模态大语言模型视觉定位中训练与评估不匹配问题,提出IoUPD方法,利用真实框作特权指导,训练时学生模型接收原始信息,教师模型接收增强提示,经监督微调与特权蒸馏损失训练,推理时无需额外模块,实验显示该方法有改进。

Comments 16 pages, 7 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03647 2026-07-07 cs.CV 新提交 88%

Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs

医学视觉语言模型真的能“看”吗?用于视觉依赖型医学视觉语言模型的反事实基础框架和硬负对比训练

Anas Zafar, Leema Krishna Murali, Siddhant Bharadwaj, Ashish Vashist, Jia Wu

机构 * The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心) Eisai Inc.(卫材株式会社) IISc, Bangalore(印度科学研究所班加罗尔分校) Cohere Labs Community(Cohere实验室社区)

专题命中 视觉定位与Grounding :vision language model(title,abstract);grounding(title,abstract);分类 cs.CV

AI总结 探讨医学视觉语言模型是依据视觉证据推理还是利用文本捷径,引入反事实评估框架和对比检索增强学习方法,提升模型视觉依赖能力并揭示跨域诊断差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17555 2026-05-14 cs.CV 88%

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking

GraphThinker: 通过事件图思维强化时间感知的视频推理

Zixu Cheng, Da Li, Jian Hu, Yuhang Zang, Ziquan Liu, Shaogang Gong, Wei Li

机构 * Queen Mary University of London(伦敦玛丽女王大学) Samsung AI Centre Cambridge(剑桥三星人工智能中心) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Nanyang Technological University(南洋理工大学)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出GraphThinker,通过构建结构化事件表示并强化视觉 grounding,减少视频推理中的时间幻觉。在RexTime和VidHalluc数据集上取得显著提升。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19689 2026-04-22 cs.AI 88%

A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding

A-MAR:基于代理的多模态艺术检索用于细粒度艺术作品理解

Shuai Wang, Hongyi Zhu, Jia-Hong Huang, Yixian Shen, Chengxi Zeng, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, Marcel Worring

机构 * University of Amsterdam(阿姆斯特丹大学) University of Bristol(布里斯托大学) Amazon AGI(亚马逊人工智慧) College of Business and Economics(商学院和经济学学院) University of Johannesburg(约翰内斯堡大学)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.AI

AI总结 本文提出A-MAR框架,通过结构化推理计划显式指导多模态艺术检索,提升解释质量和证据 grounding 能力,在艺术领域验证了代理式多模态推理的有效性。

Journal ref ICMR 2026, ACM International Conference on Multimedia Retrieval

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29798 2026-04-01 cs.CV 88%

SceneTeract: Agentic Functional Affordances and VLM Grounding in 3D Scenes

SceneTeract:代理功能可能性与视觉语言模型在3D场景中的 grounded 性

Léopold Maillard, Francis Engelmann, Tom Durand, Boxiao Pan, Yang You, Or Litany, Leonidas Guibas, Maks Ovsjanikov

机构 * École Polytechnique(巴黎综合理工学院) Dassault Systèmes(达索系统) Stanford University(斯坦福大学) USI Lugano(卢加诺大学) Technion(以色列理工学院) NVIDIA(英伟达)

专题命中 视觉定位与Grounding :VLM(title,abstract);grounding(title);vision-language model(abstract);分类 cs.CV

AI总结 SceneTeract 通过结合高层语义推理与低层几何检查,验证3D场景功能,评估合成环境中的功能失败及VLM对功能可能性的预测能力,揭示语义信心与物理可行性之间的系统性不匹配。

Comments Project page: https://sceneteract.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21778 2026-03-24 cs.CV 88%

Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models

Scene-VLM:通过视觉-语言模型进行多模态视频场景分割

Nimrod Berman, Adam Botach, Emanuel Ben-Baruch, Shunit Haviv Hakimi, Asaf Gendler, Ilan Naiman, Erez Yosef, Igor Kviatkovsky

机构 * Ben-Gurion University(本·古里安大学) Amazon Prime Video(亚马逊Prime视频) Tel-Aviv University(特拉维夫大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(title,abstract);分类 cs.CV

AI总结 本文提出Scene-VLM,首个基于视觉-语言模型的视频场景分割框架,通过融合视觉与文本信息实现多模态推理,提升场景分割的准确性和可解释性。

Comments Accepted for publication at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21509 2026-03-17 cs.CV 88%

Curing Semantic Drift: A Dynamic Approach to Grounding Generation in Large Vision-Language Models

消除语义漂移:一种动态方法用于在大视觉-语言模型中进行生成基础

Jiahe Chen, Jiaying He, Qiyuan Chen, Qian Shao, Jiahe Ying, Hongxia Xu, Jintai Chen, Jianwei Zheng, Jian Wu

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);分类 cs.CV

AI总结 本文提出DLC方法,通过动态校准logits来减少大视觉-语言模型中的语义漂移,提升生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09826 2026-03-11 cs.CV 88%

VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models

基于视觉-语言模型的点云地图定位

Shuhao Kang, Youqi Liao, Peijie Wang, Wenlong Liao, Qilin Zhang, Benjamin Busam, Xieyuanli Chen, Yun Liu

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(title,abstract);分类 cs.CV

AI总结 VLM-Loc通过视觉-语言模型的空间推理能力,实现了更准确的文本到点云定位,提升了复杂环境中的定位鲁棒性。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09638 2026-02-11 cs.CV 88%

VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model

VideoAfford: 通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding

Hanqing Wang, Mingyu Liu, Xiaoyu Chen, Chengwei MA, Yiming Zhong, Wenti Yin, Yuhao Liu, Zhiqing Cui, Jiahao Yuan, Lu Dai, Zhiyuan Ma, Hui Xiong

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);分类 cs.CV

AI总结 VideoAfford通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding,结合动态交互先验和空间感知损失函数,提升机器人操作的可操作区域识别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00998 2026-01-06 cs.CV 88%

DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models

DVGBench: 面向无人机影像的隐式到显式视觉 grounding 评估基准

Yue Zhou, Jue Chen, Zilun Zhang, Penghui Huang, Ran Ding, Zhentao Zou, PengFei Gao, Yuchen Wei, Ke Li, Xue Yang, Xue Jiang, Hongxin Yang, Jonathan Li

机构 * Hinton STAI Institute(Hinton STAI研究所) Key Laboratory of Geographic Information Science (Ministry of Education), East China Normal University(地理信息科学重点实验室(教育部)) School of Geospatial Artificial Intelligence, East China Normal University(地理空间人工智能学院) Zhejiang University(浙江大学) Shanghai Jiao Tong University(上海交通大学) Information Engineering University(信息工程大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);分类 cs.CV

AI总结 DVGBench 是一个面向无人机影像的隐式到显式视觉 grounding 评估基准,通过设计 DroneVG-R1 模型提升 LVLM 的推理能力。

Comments 20 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10011 2025-10-14 cs.CV 88%

MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output

Yanyuan Chen, Dexuan Xu, Yu Huang, Songkun Zhan, Hanpin Wang, Dongxue Chen, Xueping Wang, Meikang Qiu, Hang Li

机构 * School of Software & Microelectronics, Peking University(软件与微电子学院,北京大学) School of Computer Science, Peking University(计算机学院,北京大学) National Engineering Research Center for Software Engineering, Peking University(软件工程国家工程研究中心,北京大学) Peking University Sixth Hospital(北京大学第六医院) Augusta University(奥古斯塔大学) Peking University First Hospital(北京大学第一医院)

专题命中 视觉定位与Grounding :vision language model(title,abstract);grounding(title,abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10922 2025-08-18 cs.CV 88%

A Survey on Video Temporal Grounding with Multimodal Large Language Model

Jianlong Wu, Wei Liu, Ye Liu, Meng Liu, Liqiang Nie, Zhouchen Lin, Chang Wen Chen

机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) School of Computer Science and Technology, Shandong Jianzhu University(山东建筑大学计算机科学与技术学院) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);分类 cs.CV

Comments 20 pages,6 figures,survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22429 2025-05-29 cs.CV cs.RO 88%

Zero-Shot 3D Visual Grounding from Vision-Language Models

Rong Li, Shijie Li, Lingdong Kong, Xulei Yang, Junwei Liang

机构 * HKUST(GZ)(香港科技大学(广州)) I 2 R, A*STAR(I2R, A*STAR) NUS(国立大学) CSE, HKUST(计算机科学与工程系,香港科技大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);分类 cs.CV

Comments 3D-LLM/VLA @ CVPR 2025; Project Page at https://seeground.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06510 2025-05-27 cs.CV 88%

Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves?

Yuan-Hong Liao, Rafid Mahmood, Sanja Fidler, David Acuna

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所) NVIDIA(英伟达) University of Ottawa(渥太华大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);分类 cs.CV

Comments Accepted at CVPR 2025. 22 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19281 2025-03-26 cs.RO cs.AI 88%

CubeRobot: Grounding Language in Rubik's Cube Manipulation via Vision-Language Model

Feiyang Wang, Xiaomin Yu, Wangyu Wu

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title);VLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏