arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7360 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7360 篇

2510.17384 2025-10-21 cs.CV 79%

Closed-Loop Transfer for Weakly-supervised Affordance Grounding

Jiajin Tang, Zhengxuan Wei, Ge Zheng, Sibei Yang

机构 * ShanghaiTech University(上海科技大学) School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17023 2025-10-21 cs.CV cs.MM 79%

Enrich and Detect: Video Temporal Grounding with Multimodal LLMs

Shraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa, Lorenzo Torresani, Triantafyllos Afouras

机构 * FAIR, Meta(FAIR、Meta) Johns Hopkins University(约翰霍普金斯大学) Northeastern University(东北大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICCV 2025 (Highlights)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17007 2025-10-21 cs.CV 79%

An empirical study of the effect of video encoders on Temporal Video Grounding

Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Felipe Bravo-Marquez

机构 * Department of Computer Science, University of Chile(计算机科学系,智利大学) CENIA and IMFD(CENIA和IMFD) Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) National Institute of Advanced Industrial Science and Technology(国家先进工业科学与技术研究所)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16989 2025-10-21 cs.CV 79%

Training-free Online Video Step Grounding

Luca Zanella, Massimiliano Mancini, Yiming Wang, Alessio Tonioni, Elisa Ricci

机构 * University of Trento(特伦托大学) Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会) Google(谷歌)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments NeurIPS 2025. Project website at https://lucazanella.github.io/baglm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11375 2025-10-21 cs.CV cs.MM 79%

Text-controlled Motion Mamba: Text-Instructed Temporal Grounding of Human Motion

Xinghan Wang, Zixi Kang, Yadong Mu

机构 * Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Image Processing (TIP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16080 2025-10-21 q-bio.QM cs.AI 79%

TriAgent: Automated Biomarker Discovery with Deep Research Grounding for Triage in Acute Care by LLM-Based Multi-Agent Collaboration

Kerem Delikoyun, Qianyu Chen, Win Sen Kuan, John Tshon Yit Soong, Matthew Edward Cove, Oliver Hayden

机构 * Technical University of Munich(慕尼黑技术大学) National University of Singapore(新加坡国立大学) National University Hospital(新加坡国立大学医院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14965 2025-10-17 cs.CV 79%

ChangingGrounding: 3D Visual Grounding in Changing Scenes

Miao Hu, Zhiwei Huang, Tai Wang, Jiangmiao Pang, Dahua Lin, Nanning Zheng, Runsen Xu

机构 * Xi’an Jiaotong University(西安交通大学) Zhejiang University(浙江大学) The Chinese University of Hong Kong(香港中文大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13800 2025-10-17 cs.CV 79%

Reasoning in Space via Grounding in the World

Yiming Chen, Zekun Qi, Wenyao Zhang, Xin Jin, Li Zhang, Peidong Liu

机构 * Westlake University(西湖大学) Shanghai Innovation Institute(上海创新研究院) Zhejiang University(浙江大学) Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) Eastern Institute of Technology(技术研究院) Fudan University(复旦大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02477 2025-10-16 cs.RO cs.CV 79%

Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision

Xiaofeng Han, Shunpeng Chen, Zenghuang Fu, Zhe Feng, Lue Fan, Dong An, Changwei Wang, Li Guo, Weiliang Meng, Xiaopeng Zhang, Rongtao Xu, Shibiao Xu

机构 * aThe State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, China [1ex] bSchool of Artificial Intelligence, University of Chinese Academy of Sciences, China [1ex] cSchool of Artificial Intelligence, Beijing University of Posts Telecommunications, China [1ex] dKey Laboratory of Computing Power Network Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences), China [1ex] e Shandong Provincial Key Laboratory of Computing Power Internet Service Computing, Shandong Fundamental Research Center for Computer Science, China

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments 27 pages, 11 figures. Accepted to Information Fusion. Final journal version: volume 126 (Part B), February 2026

Journal ref Information Fusion, 126 (Part B), February 2026, 103652

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10587 2025-10-14 cs.CV 79%

A Simple and Better Baseline for Visual Grounding

Jingchao Wang, Wenlong Zhang, Dingjiang Huang, Hong Wang, Yefeng Zheng

机构 * School of Data Science(数据科学学院) Engineering East China Normal University Shanghai, China(工程 东华师范大学 上海中国) OpenScience Lab Shanghai AI Laboratory Shanghai, China(OpenScience Lab 上海AI实验室 上海中国) School of Life Science(生命科学学院) Technology Xi'an Jiaotong University Xi'an, China(技术 西安交通大学 西安中国) Medical Artificial Intelligence Laboratory Westlake University Hangzhou, China(医学人工智能实验室 西湖大学 杭州中国)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05153 2025-10-08 cs.AI cs.IT math.IT 79%

An Algorithmic Information-Theoretic Perspective on the Symbol Grounding Problem

Zhangchi Liu

机构 * Zhangchi Liu(刘志强)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 7 pages, 1 table (in appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03840 2025-10-07 cs.CV 79%

Mirage: Unveiling Hidden Artifacts in Synthetic Images with Large Vision-Language Models

Pranav Sharma, Shivank Garg, Durga Toshniwal

机构 * Indian Institute of Technology Roorkee(印度理工学院罗奥里分校)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments ACM MM'25, MALLM Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03376 2025-10-07 cs.CV eess.IV 79%

Visual Language Model as a Judge for Object Detection in Industrial Diagrams

Sanjukta Ghosh

专题命中 视觉定位与Grounding :visual language model(title,abstract);分类 cs.CV

Comments Pre-review version submitted to IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00907 2025-10-06 cs.AI 79%

Grounding Multimodal LLMs to Embodied Agents that Ask for Help with Reinforcement Learning

Ram Ramrakhya, Matthew Chang, Xavier Puig, Ruta Desai, Zsolt Kira, Roozbeh Mottaghi

机构 * Georgia Institute of Technology(佐治亚理工学院) Meta FAIR

专题命中 视觉定位与Grounding :grounding(title);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11955 2025-09-30 cs.CV 79%

Temporal Grounding as a Learning Signal for Referring Video Object Segmentation

Seunghun Lee, Jiwan Seo, Jeonghoon Kim, Sungho Moon, Siwon Kim, Haeun Yun, Hyogyeong Jeon, Wonhyeok Choi, Jaehoon Jeong, Zane Durante, Sang Hyun Park, Sunghoon Im

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Project page: https://seung-hun-lee.github.io/projects/TGL/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04511 2025-09-30 cs.CV 79%

FA: Forced Prompt Learning of Vision-Language Models for Out-of-Distribution Detection

Xinhua Lu, Runhe Lai, Yanqi Wu, Kanghao Chen, Wei-Shi Zheng, Ruixuan Wang

机构 * Sun Yat-sen University(中山大学) Peng Cheng Laboratory(鹏城实验室) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Key Laboratory of Machine Intelligence and Advanced Computing(人工智能与先进计算重点实验室)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments 12 pages, 4 figures, Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23768 2025-09-29 cs.CL cs.CV 79%

Texture or Semantics? Vision-Language Models Get Lost in Font Recognition

Zhecheng Li, Guoxian Song, Yujun Cai, Zhen Xiong, Junsong Yuan, Yiwei Wang

机构 * University of California, San Diego(加州大学圣地亚哥分校) ByteDance(字节跳动) The University of Queensland(昆士兰大学) University of Southern California(南加州大学) University at Buffalo(布法罗大学) University of California, Merced(加州大学默塞德分校)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19096 2025-09-26 cs.CV cs.SE 79%

Investigating Traffic Accident Detection Using Multimodal Large Language Models

Ilhan Skender, Kailin Tong, Selim Solmaz, Daniel Watzenig

机构 * Embedded Systems Group (Dept.-E)(嵌入式系统组) Virtual Vehicle Research GmbH(虚拟车辆研究公司) Control Systems Group (Dept.-E)(控制系统组) Institute of Visual Computing(视觉计算研究所) Graz University of Technology(格拉茨技术大学)

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

Comments Accepted for presentation at the 2025 IEEE International Automated Vehicle Validation Conference (IAVVC 2025). Final version to appear in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24372 2025-09-26 cs.CV 79%

Beyond Quantity: Distribution-Aware Labeling for Visual Grounding

Yichi Zhang, Gongwei Chen, Jun Zhu, Jia Wan, Liqiang Nie

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 18pages, 8figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21188 2025-09-23 cs.CV 79%

GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding

Zijun Lin, Shuting He, Cheston Tan, Bihan Wen

机构 * Nanyang Technological University(南洋理工大学) Centre for Frontier AI Research, A*STAR(前沿人工智能研究中心,A*STAR) MoE Key Laboratory of Interdisciplinary Research of Computation and Economics, Shanghai University of Finance and Economics(计算与经济学交叉研究关键实验室,上海财经大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16806 2025-09-23 cs.CV 79%

FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation

Fan Yang, Yousong Zhu, Xin Li, Yufei Zhan, Hongyin Zhao, Shurong Zheng, Yaowei Wang, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Science(中国科学院大学人工智能学院) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室) Wuhan AI Research, Wuhan, China(武汉人工智能研究所)

专题命中 视觉定位与Grounding :vision-language model(title);vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16534 2025-09-23 cs.CL cs.AI 79%

InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding

Cheng Jiayang, Qianqian Zhuang, Haoran Li, Chunkit Chan, Xin Liu, Lin Qiu, Yangqiu Song

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Shanghai Jiaotong University(上海交通大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15243 2025-09-22 cs.CV 79%

Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models

Muhammad Imran, Yugyung Lee

机构 * Computer Science, School of Science and Engineering, University of Missouri - Kansas City(计算机科学系,科学与工程学院,密苏里大学-堪萨斯城分校)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments 8 pages, 6 figures, 3 tables

Journal ref Non-Archival track - The First Workshop on Multimodal Knowledge and Language Modeling IJCAI 2025 Workshop, August 16, 2025 IJCAI 2025 Workshop, August 16, 2025 Room 516B, Palais des congrès, Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13836 2025-09-18 cs.CV cs.CL 79%

Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models

Weihang Wang, Xinhao Li, Ziyue Wang, Yan Pang, Jielei Zhang, Peiyi Li, Qiang Zhang, Longwen Gao

机构 * Bilibili(哔哩哔哩) UESTC University of Virginia(弗吉尼亚大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted by EMNLP2025 Finding

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13747 2025-09-18 cs.CV 79%

Improving Generalized Visual Grounding with Instance-aware Joint Learning

Ming Dai, Wenxuan Cheng, Jiang-Jiang Liu, Lingfeng Yang, Zhenhua Feng, Wankou Yang, Jingdong Wang

机构 * School of Automation, Southeast University(东南大学自动化学院) Baidu Inc.(百度公司) JiangNan University(江南大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) in September 2025

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11866 2025-09-16 cs.CV 79%

Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding

Meng Luo, Shengqiong Wu, Liqiang Jing, Tianjie Ju, Li Zheng, Jinxiang Lai, Tianlong Wu, Xinya Du, Jian Li, Siyuan Yan, Jiebo Luo, William Yang Wang, Hao Fei, Mong-Li Lee, Wynne Hsu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 25 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10278 2025-09-15 cs.CV 79%

Detecting Text Manipulation in Images using Vision Language Models

Vidit Vidit, Pavel Korshunov, Amir Mohammadi, Christophe Ecabert, Ketan Kotwal, Sébastien Marcel

机构 * IDIAP Research Institute(IDIAP研究 institute)

专题命中 视觉定位与Grounding :vision language model(title,abstract);分类 cs.CV

Comments Accepted in Synthetic Realities and Biometric Security Workshop BMVC-2025. For paper page see https://www.idiap.ch/paper/textvlmdet/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09584 2025-09-12 cs.CV cs.RO 79%

Visual Grounding from Event Cameras

Lingdong Kong, Dongyue Lu, Ao Liang, Rong Li, Yuhao Dong, Tianshuai Hu, Lai Xing Ng, Wei Tsang Ooi, Benoit R. Cottereau

机构 * NUS(新加坡国立大学) HKUST(GZ)(香港科技大学(广州)) NTU(南洋理工大学) HKUST(香港科技大学) I 2 R, A*STAR(新加坡科技研究局) IPAL, CNRS(法国国家科学研究中心IPAL) CerCo, CNRS(法国国家科学研究中心CerCo)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Abstract Paper (Non-Archival) @ ICCV 2025 NeVi Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06291 2025-09-09 cs.CV 79%

Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding

Jiangnan Xie, Xiaolong Zheng, Liang Zheng

机构 * College of Electronics and Information, Hangzhou Dianzi University(电子信息学院,杭州电子大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06233 2025-09-09 cs.RO cs.CV 79%

O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation

Tongxuan Tian, Xuhui Kang, Yen-Ling Kuo

机构 * University of Virginia(弗吉尼亚大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Conference on Robot Learning (CoRL) 2025. Project website: https://o3afford.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏