arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7314 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7314 篇

2508.01008 2025-08-05 cs.CV 85%

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation

Cihang Peng, Qiming Hou, Zhong Ren, Kun Zhou

机构 * State Key Lab of CAD&CG(计算机辅助设计与图形学国家重点实验室)

专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract);grounding(abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21316 2025-07-17 cs.CV 85%

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents

Badri Vishal Kasuba, Parag Chaudhuri, Ganesh Ramakrishnan

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);visual question answering(abstract);分类 cs.CV

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10596 2025-07-16 cs.CV 85%

GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding

Rui Hu, Lianghui Zhu, Yuxuan Zhang, Tianheng Cheng, Lei Liu, Heng Liu, Longjin Ran, Xiaoxin Chen, Wenyu Liu, Xinggang Wang

机构 * School of EIC, Huazhong University of Science & Technology(华中科技大学电子信息学院) vivo AI Lab(vivo人工智能实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments To appear at ICCV 2025. Code: https://github.com/hustvl/GroundingSuite

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19498 2025-06-25 cs.RO cs.AI 85%

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models

Yiteng Chen, Wenbo Li, Shiyi Wang, Huiping Zhuang, Qingyao Wu

机构 * School of Software Engineering, South China University of Technology(软件工程学院,华南理工大学) School of Future Technology, South China University of Technology(未来技术学院,华南理工大学) Shien-Ming Wu School of Intelligent Engineering, South China University of Technology(智能工程学院,华南理工大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.AI

Comments submitted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17901 2025-06-24 cs.CV 85%

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs

Yixuan Wu, Yang Zhang, Jian Wu, Philip Torr, Jindong Gu

机构 * University of Oxford(牛津大学) Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20104 2025-06-16 cs.CV 85%

New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM Collaboration

Xuzheng Yang, Junzhuo Liu, Peng Wang, Guoqing Wang, Yang Yang, Heng Tao Shen

机构 * University of Electronic Science and Technology of China(电子科技大学) Tongji University(同济大学)

专题命中 视觉定位与Grounding :MLLM(title,abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV

Comments Accepted by TPAMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01977 2025-06-10 cs.CV 85%

AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs

Hongxin Li, Jingfan Chen, Jingran Su, Yuntao Chen, Qing Li, Zhaoxiang Zhang

机构 * University of Chinese Academy of Sciences (UCAS)(中国科学院大学) New Laboratory of Pattern Recognition (NLPR), CASIA(中国科学院模式识别新技术实验室) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA(中国科学院多模态人工智能系统国家重点实验室) Hong Kong Institute of Science & Innovation, CASIA(香港科学与创新研究院) The Hong Kong Polytechnic University(香港理工大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments Accepted to ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22222 2025-05-29 cs.CV cs.CL 85%

Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation

Yunsoo Kim, Jinge Wu, Su-Hwan Kim, Pardeep Vasudev, Jiashu Shen, Honghan Wu

机构 * UCL(伦敦大学学院) Technical University of Munich(慕尼黑技术大学) University of Oxford(牛津大学) University of Glasgow(格拉斯哥大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);LLaVA(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19552 2025-05-23 cs.CV 85%

GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing

Hosam Elgendy, Ahmed Sharshar, Ahmed Aboeitta, Yasser Ashraf, Mohsen Guizani

机构 * Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(莫扎德·本·扎耶德人工智能大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV

Comments 14 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14231 2025-05-21 cs.CV 85%

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

Sule Bai, Mingxing Li, Yong Liu, Jing Tang, Haoji Zhang, Lei Sun, Xiangxiang Chu, Yansong Tang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) AMAP, Alibaba Group(阿里云研究院)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11852 2025-05-20 cs.CV 85%

MedSG-Bench: A Benchmark for Medical Image Sequences Grounding

Jingkun Yue, Siqi Zhang, Zinan Jia, Huihuan Xu, Zongbo Han, Xiaohong Liu, Guangyu Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tianjin University(天津大学) South China Hospital, Medical School, Shenzhen University(深圳大学医学院南方医院)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02278 2025-05-06 cs.CV 85%

Compositional Image-Text Matching and Retrieval by Grounding Entities

Madhukar Reddy Vongala, Saurabh Srivastava, Jana Košecká

机构 * Department of Computer Science, George Mason University(计算机科学系,乔治·马歇尔大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);visual question answering(abstract);分类 cs.CV

Comments Accepted at CVPR-W

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01578 2025-05-06 cs.CV 85%

Grounding Task Assistance with Multimodal Cues from a Single Demonstration

Gabriel Sarch, Balasaravanan Thoravi Kumaravel, Sahithya Ravi, Vibhav Vineet, Andrew D. Wilson

机构 * Microsoft Research(微软研究院)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision language model(abstract);visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17365 2025-04-30 cs.CV cs.CL 85%

TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation

Ling You, Wenxuan Huang, Xinni Xie, Xiangyi Wei, Bangyan Li, Shaohui Lin, Yang Li, Changbo Wang

机构 * East China Normal University(东华大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11914 2025-04-17 cs.CV 85%

AnomalyR1: A GRPO-based End-to-end MLLM for Industrial Anomaly Detection

Yuhao Chao, Jie Liu, Jie Tang, Gangshan Wu

专题命中 视觉定位与Grounding :MLLM(title,abstract);VLM(abstract);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10491 2025-03-21 cs.CV 85%

TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning

Aritra Bhowmik, Mohammad Mahdi Derakhshani, Dennis Koelma, Yuki M. Asano, Martin R. Oswald, Cees G. M. Snoek

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16024 2024-11-27 cs.AI 85%

From Goal-Conditioned to Language-Conditioned Agents via Vision-Language Models

Theo Cachet, Christopher R. Dance, Olivier Sigaud

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06622 2024-08-14 cs.CV 85%

ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding

Yubin Wang, Xinyang Jiang, De Cheng, Dongsheng Li, Cairong Zhao

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract);分类 cs.CV

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17672 2024-08-06 cs.CV cs.GR 85%

BlenderAlchemy: Editing 3D Graphics with Vision-Language Models

Ian Huang, Guandao Yang, Leonidas Guibas

专题命中 视觉定位与Grounding :vision-language model(title,abstract);visual reasoning(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21465 2024-08-01 cs.CV 85%

MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection

Kuo Wang, Lechao Cheng, Weikai Chen, Pingping Zhang, Liang Lin, Fan Zhou, Guanbin Li

专题命中 视觉定位与Grounding :vision-language model(title,abstract);vision language model(abstract);VLM(abstract);分类 cs.CV

Comments Codes are available at https://github.com/wkfdb/MarvelOVD

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03118 2024-06-26 cs.CV 85%

LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models

Gabriela Ben Melech Stan, Estelle Aflalo, Raanan Yehezkel Rohekar, Anahita Bhiwandiwalla, Shao-Yen Tseng, Matthew Lyle Olson, Yaniv Gurwicz, Chenfei Wu, Nan Duan, Vasudev Lal

专题命中 视觉定位与Grounding :vision-language model(title,abstract);LLaVA(abstract);grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13894 2024-06-21 cs.CV cs.CY 85%

Using Multimodal Large Language Models for Automated Detection of Traffic Safety Critical Events

Mohammad Abu Tami, Huthaifa I. Ashqar, Mohammed Elhenawy

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);LLaVA(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08730 2024-04-04 cs.CL cs.CV 85%

Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization

Renjie Pi, Tianyang Han, Wei Xiong, Jipeng Zhang, Runtao Liu, Rui Pan, Tong Zhang

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12507 2023-06-16 cs.AI 85%

Distilling Internet-Scale Vision-Language Models into Embodied Agents

Theodore Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus, Ishita Dasgupta

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.AI

Comments 9 pages, 7 figures. Presented at ICML 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.15517 2023-02-08 cs.CV 85%

Medical Image Understanding with Pretrained Vision Language Models: A Comprehensive Study

Ziyuan Qin, Huahui Yi, Qicheng Lao, Kang Li

专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract);grounding(abstract);分类 cs.CV

Comments Accepted to ICLR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01554 2025-09-03 cs.CV cs.AI cs.LG 85%

Unified Supervision For Vision-Language Modeling in 3D Computed Tomography

Hao-Chih Lee, Zelong Liu, Hamza Ahmed, Spencer Kim, Sean Huver, Vishwesh Nath, Zahi A. Fayad, Timothy Deyer, Xueyan Mei

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,comments);分类 cs.CV、cs.AI、cs.LG

Comments ICCV 2025 VLM 3d Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21794 2025-06-19 cs.CV cs.AI cs.LG 85%

Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey

Atsuyuki Miyai, Jingkang Yang, Jingyang Zhang, Yifei Ming, Yueqian Lin, Qing Yu, Go Irie, Shafiq Joty, Yixuan Li, Hai Li, Ziwei Liu, Toshihiko Yamasaki, Kiyoharu Aizawa

机构 * The University of Tokyo(东京大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室) Duke University(杜克大学) Salesforce AI Research(Salesforce人工智能研究) Nanyang Technological University(南洋理工大学) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Tokyo University of Science(东京科学大学)

专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract,comments);分类 cs.CV、cs.AI、cs.LG

Comments Accepted at TMLR2025. Survey paper. We welcome questions, issues, and paper requests via https://github.com/AtsuMiyai/Awesome-OOD-VLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20149 2024-10-29 cs.CV cs.AI cs.LG 85%

AdaNeg: Adaptive Negative Proxy Guided OOD Detection with Vision-Language Models

Yabin Zhang, Lei Zhang

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,comments);分类 cs.CV、cs.AI、cs.LG

Comments NIPS 2024 Camera Ready, Codes are available at \url{https://github.com/YBZh/OpenOOD-VLM}

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13690 2026-08-17 cs.CV cs.AI 新提交 85%

MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation

MedPlex:用于临床基础医学分割的深度视觉-语言协同适配

Rafi Ibn Sultan, Hui Zhu, Chengyin Li, Dongxiao Zhu

机构 * Wayne State University(韦恩州立大学) Henry Ford Health(亨利福特医疗集团) Institute for AI and Data Science Wayne State University(韦恩州立大学人工智能与数据科学研究院)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 MedPlex是一种端到端VLM框架,通过双向融合和两级概念对齐,实现医学图像分割的视觉-语言协同适配,在CT、MR的多类医学分割任务中达到最优性能。

Comments Accepted By BMVC-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24353 2026-07-31 cs.CV cs.LG 版本更新 85%

Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints

开放词汇BEV分割:基于3D感知几何约束的方法

Hojun Choi, Seulbin Hwang, Daejung Kim, Kisung Kim, Hyunjung Shim, Jinhan Lee

机构 * KAIST AI(韩国科学技术院人工智能) NAVER LABS(NAVER实验室)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);分类 cs.CV、cs.LG

AI总结 针对开放词汇BEV分割中2D VLM语义到BEV的3D几何不一致问题,提出OVBEVSeg框架,通过三阶段几何约束(伪标签、场景优化、知识蒸馏)实现高效在线推理,在nuScenes上未见类别mIoU超越闭集方法15.3,推理速度提升2.5倍。

Comments This paper has been accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏