arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26117 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 4478 篇

2510.22798 2025-10-28 cs.CL cs.LG 79%

VEHME: A Vision-Language Model For Evaluating Handwritten Mathematics Expressions

Thu Phuong Nguyen, Duc M. Nguyen, Hyotaek Jeon, Hyunwook Lee, Hyunmin Song, Sungahn Ko, Taehwan Kim

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.LG

Comments EMNLP 2025. Project Website: https://vehme.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22455 2025-10-28 cs.SD cs.AI eess.AS 79%

Evaluating Multimodal Large Language Models on Core Music Perception Tasks

Brandon James Carone, Iran R. Roman, Pablo Ripollés

机构 * Department of Psychology, Music and Audio Research Laboratory(心理学系、音乐与音频研究实验室) Department of Electronic Engineering and Computer Science(电子工程与计算机科学系)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.AI

Comments Accepted to the NeurIPS 2025 Workshop on AI for Music (AI4Music), 16 pages, 1 figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02745 2025-10-28 cs.CV 79%

Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval

Lanyun Zhu, Deyi Ji, Tianrun Chen, Haiyang Wu, Shiqi Wang

机构 * City University of Hong Kong(香港城市大学) Tencent(腾讯) Zhejiang University(浙江大学)

专题命中 视觉推理 :MLLM(title,abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11261 2025-10-24 cs.AI cs.CL 79%

Sycophancy in Vision-Language Models: A Systematic Analysis and an Inference-Time Mitigation Framework

Yunpu Zhao, Rui Zhang, Junbin Xiao, Changxin Ke, Ruibo Hou, Yifan Hao, Ling Li

机构 * School of Computer Science and Technology, University of Science and Technology of China(计算机科学与技术学院,中国科学技术大学) State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences(处理器国家重点实验室,中国科学院计算技术研究所) Department of Computer Science, National University of Singapore(计算机科学系,新加坡国立大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences(软件智能研究中心,中国科学院软件研究所)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI

Journal ref Neurocomputing, Volume 659, 2026, 131217

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19451 2025-10-23 cs.CV cs.MM 79%

Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis

Xueqi Ma, Yanbei Jiang, Sarah Erfani, James Bailey, Weifeng Liu, Krista A. Ehinger, Jey Han Lau

机构 * The University of Melbourne(墨尔本大学) China University of Petroleum (East China)(中国石油大学(华东))

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18303 2025-10-22 cs.CV 79%

Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language Models

Lehan Wang, Yi Qin, Honglong Yang, Xiaomeng Li

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17274 2025-10-21 cs.CV 79%

Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models

Katie Luo, Jingwei Ji, Tong He, Runsheng Xu, Yichen Xie, Dragomir Anguelov, Mingxing Tan

机构 * Computer and Information Sciences Department, Cornell University(康奈尔大学计算机与信息科学系) Waymo LLC(Waymo公司) UC Berkeley(伯克利大学)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.CV

Comments In proceedings of IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25373 2025-10-17 cs.AI 79%

From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models

Chenyue Zhou, Mingxuan Wang, Yanbiao Ma, Chenxu Wu, Wanyi Chen, Zhe Qian, Xinyu Liu, Yiwei Zhang, Junhao Wang, Hengbo Xu, Fei Luo, Xiaohua Chen, Xiaoshuai Hao, Hehan Li, Andi Zhang, Wenxuan Wang, Kaiyan Zhang, Guoli Jia, Lingling Li, Zhiwu Lu, Yang Lu, Yike Guo

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) Xiamen University(厦门大学) The Hong Kong University of Science and Technology(香港理工大学) Nanyang Technological University(南洋理工大学)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12190 2025-10-15 cs.CV 79%

Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos

Shingo Yokoi, Kento Sasaki, Yu Yamaguchi

机构 * Turing Inc.(图灵公司)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV

Comments 2nd Place Winner, ICCV 2025 2COOOL Competition

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10117 2025-10-14 cs.AI 79%

DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay

Yunxiang Mo, Tianshi Zheng, Qing Zong, Jiayu Liu, Baixuan Xu, Yauwai Yim, Chunkit Chan, Jiaxin Bai, Yangqiu Song

机构 * Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China(计算机科学与工程系,香港科技大学,香港特别行政区,中国)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI

Comments EMNLP 2025 Wordplay (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09358 2025-10-13 cs.CV 79%

Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models

Qihang Ma, Shengyu Li, Jie Tang, Dingkang Yang, Shaodong Chen, Yingyi Zhang, Chao Feng, Jiao Ran

机构 * ByteDance Douyin Content Group(字节跳动抖音内容团队)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV

Comments EMNLP2025. Code is avaible at https://github.com/bytedance/DynamicCoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10610 2025-10-07 cs.CV cs.CL 79%

MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly

Zhaowei Wang, Wenhao Yu, Xiyu Ren, Jipeng Zhang, Yu Zhao, Rohit Saxena, Liang Cheng, Ginny Wong, Simon See, Pasquale Minervini, Yangqiu Song, Mark Steedman

机构 * CSE Department, HKUST(香港科技大学计算机科学与工程系) Tencent AI Seattle Lab(腾讯AI西雅图实验室) University of Edinburgh(爱丁堡大学) NVIDIA AI Technology Center (NVAITC), NVIDIA, Santa Clara, USA(英伟达圣克拉拉人工智能技术中心)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV

Comments Accepted as a spotlight at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16162 2025-10-03 cs.CV cs.CL 79%

Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning

Yihong Tang, Ao Qu, Zhaokai Wang, Dingyi Zhuang, Zhaofeng Wu, Wei Ma, Shenhao Wang, Yunhan Zheng, Zhan Zhao, Jinhua Zhao

机构 * McGill University(麦吉尔大学) Massachusetts Institute of Technology(麻省理工学院) Shanghai Jiao Tong University(上海交通大学) The Hong Kong Polytechnic University(香港理工大学) University of Florida(佛罗里达大学) The University of Hong Kong(香港大学)

专题命中 视觉推理 :vision language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00690 2025-10-02 cs.AI 79%

ACPO: Adaptive Curriculum Policy Optimization for Aligning Vision-Language Models in Complex Reasoning

Yunhao Wang, Ziting Li, Shuai Chen, Tao Liu, Chao Song, Junjie Jiang, Jian Zhu, Peng Gao, Bin Qin

机构 * Xiaomi Inc.(小米公司)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19778 2025-10-02 cs.AI 79%

Multimodal Large Language Models for Bioimage Analysis

Shanghang Zhang, Gaole Dai, Tiejun Huang, Jianxu Chen

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19825 2025-09-30 cs.AI cs.CL 79%

Evaluating Compliance with Visualization Guidelines in Diagrams for Scientific Publications Using Large Vision Language Models

Johannes Rückert, Louise Bloch, Christoph M. Friedrich

机构 * Department of Computer Science, University of Applied Sciences and Arts Dortmund(应用科学与艺术大学多特蒙德计算机科学系) Institute for Medical Informatics, Biometry and Epidemiology (IMIBE)(医学信息学、生物统计与流行病学研究所) Institute for Artificial Intelligence in Medicine (IKIM), University Hospital Essen(医学人工智能研究所,埃森大学医院)

专题命中 视觉推理 :vision language model(title,abstract);分类 cs.AI

Comments Accepted at ICDAR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15298 2025-09-30 cs.RO cs.CL cs.CV 79%

AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving

Kangan Qian, Sicong Jiang, Yang Zhong, Ziang Luo, Zilin Huang, Tianze Zhu, Kun Jiang, Mengmeng Yang, Zheng Fu, Jinyu Miao, Yining Shi, He Zhe Lim, Li Liu, Tianbao Zhou, Huang Yu, Yifei Hu, Guang Li, Guang Chen, Hao Ye, Lijun Sun, Diange Yang

机构 * School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动学院) McGill University(麦吉尔大学) Automotive and Robotics, Xiaomi Corporation(小米公司汽车与机器人部门) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV

Comments 19 pages, 8 figures

Journal ref EMNLP2025 Fundings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21874 2025-09-29 cs.LG 79%

Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models

Yifei Peng, Yaoli Liu, Enbo Xia, Yu Jin, Wang-Zhou Dai, Zhong Ren, Yao-Xiang Ding, Kun Zhou

机构 * State Key Laboratory of CAD&CG(计算机辅助设计与图形学国家重点实验室) National Key Laboratory for Novel Software Technology(新型软件技术国家实验室)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21451 2025-09-29 cs.CV cs.CL 79%

VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding

Abdul Waheed, Zhen Wu, Dareen Alharthi, Seungone Kim, Bhiksha Raj

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉推理 :MLLM(title,abstract);分类 cs.CV

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17046 2025-09-29 cs.CL cs.LG 79%

MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models

Xiaolong Wang, Zhaolu Kang, Wangyuxuan Zhai, Xinyue Lou, Yunghwei Lai, Ziyue Wang, Yawen Wang, Kaiyu Huang, Yile Wang, Peng Li, Yang Liu

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19003 2025-09-24 cs.CV 79%

Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards

Honghao Chen, Xingzhou Lou, Xiaokun Feng, Kaiqi Huang, Xinlong Wang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18425 2025-09-24 cs.CV 79%

Losing the Plot: How VLM responses degrade on imperfect charts

Philip Wootaek Shin, Jack Sampson, Vijaykrishnan Narayanan, Andres Marquez, Mahantesh Halappanavar

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Pacific Northwest National Laboratory(太平洋西北国家实验室)

专题命中 视觉推理 :VLM(title);vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05255 2025-09-23 cs.CV cs.CL 79%

Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning

Yana Wei, Liang Zhao, Jianjian Sun, Kangheng Lin, Jisheng Yin, Jingcheng Hu, Yinmin Zhang, En Yu, Haoran Lv, Zejia Weng, Jia Wang, Chunrui Han, Yuang Peng, Qi Han, Zheng Ge, Xiangyu Zhang, Daxin Jiang, Vishal M. Patel

机构 * Johns Hopkins University(约翰霍普金斯大学) StepFun BUPT(北京邮电大学) UCAS(中国科学院大学) THU(清华大学) HUST(华中科技大学)

专题命中 视觉推理 :visual reasoning(title,abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13317 2025-09-17 cs.CV 79%

3D Aware Region Prompted Vision Language Model

An-Chieh Cheng, Yang Fu, Yukang Chen, Zhijian Liu, Xiaolong Li, Subhashree Radhakrishnan, Song Han, Yao Lu, Jan Kautz, Pavlo Molchanov, Hongxu Yin, Xiaolong Wang, Sifei Liu

机构 * MIT(麻省理工学院) NVIDIA(英伟达)

专题命中 视觉推理 :vision language model(title);vision-language model(abstract);分类 cs.CV

Comments Project Website: https://www.anjiecheng.me/sr3d

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12883 2025-09-17 cs.CV 79%

Lego-Edit: A General Image Editing Framework with Model-Level Bricks and MLLM Builder

Qifei Jia, Yu Liu, Yajie Chai, Xintong Yao, Qiming Lu, Yasen Zhang, Runyu Shi, Ying Huang, Guoquan Zhang

机构 * Xiaomi Corporation Beijing, China(小米公司北京)

专题命中 视觉推理 :MLLM(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03837 2025-09-05 cs.LG cs.IT math.IT 79%

Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models

Kimia Ehsani, Walid Saad

机构 * Bradley Department of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 视觉推理 :multimodal large language model(title,abstract);分类 cs.LG

Comments Accepted at IEEE GLOBECOM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09698 2025-09-03 cs.RO cs.AI 79%

ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation

Enyu Zhao, Vedant Raval, Hejia Zhang, Jiageng Mao, Zeyu Shangguan, Stefanos Nikolaidis, Yue Wang, Daniel Seita

机构 * Department of Computer Science, University of Southern California(计算机科学系,南加州大学)

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.AI

Comments Conference on Robot Learning (CoRL) 2025. 50 pages and 30 figures. v2 is the camera-ready and includes a few more new experiments compared to v1

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20851 2025-08-29 cs.CV 79%

PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis

Ye Zhang, Yu Zhou, Jingwen Qi, Yongbing Zhang, Simon Puettmann, Finn Wichmann, Larissa Pereira Ferreira, Lara Sichward, Julius Keyl, Sylvia Hartmann, Shuo Zhao, Hongxiao Wang, Xiaowei Xu, Jianxu Chen

机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) Leibniz-Institut für Analytische Wissenschaften – ISAS – e.V.(莱比锡分析科学研究所(ISAS)) Department of Pathology, The Sixth Affiliated Hospital, Sun Yat-sen University(中山大学第六附属医院病理科部) Institute of Pathology, University Hospital Essen(埃森大学医院病理科研究所) Academy for Multidisciplinary Studies, Capital Normal University(首都师范大学多学科研究学院)

专题命中 视觉推理 :visual reasoning(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19967 2025-08-28 cs.CV 79%

Assessing the Geolocation Capabilities, Limitations and Societal Risks of Generative Vision-Language Models

Oliver Grainge, Sania Waheed, Jack Stilgoe, Michael Milford, Shoaib Ehsan

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV

Comments Accepted to AAAI Fall Symposium 2025 on AI Trustworthiness and Risk Assessment for Challenging Contexts (ATRACC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01777 2025-08-28 cs.CL cs.CV 79%

NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision

Xiang Li, Wenyue Hua, Kaijie Zhu, Lingyao Li, Haoyang Ling, Jinkui Chi, Qi Dou, Jindong Wang, Yongfeng Zhang, Xin Ma, Lizhou Fan

机构 * Shandong University, School of Control Science Rutgers University, Department of Computer Science 110 Frelinghuysen Road Piscataway New Brunswick NJ USA 08854 University of California, Santa Barbara 1210 Cheadle Hall Santa Barbara California USA 93106 University of Michigan, School of Information 500 S State St Ann Arbor Michigan USA 48109 The Chinese University of Hong Kong, Department of Computer Science The College of William \& Mary, School of Computing, Data Sciences \& Physics 200 Stadium Dr Williamsburg Virginia USA 23185 The Chinese University of Hong Kong, Department of Psychiatry 9 Chuen On Rd Tai Po New Territories Hong Kong SAR, China 999077 Rutgers University, Department of Computer Science University of California, Santa Barbara University of Michigan, School of Information The College of William \& Mary, School of Computing, Data Sciences \& Physics The Chinese University of Hong Kong, Department of Psychiatry

专题命中 视觉推理 :vision-language model(title,abstract);分类 cs.CV

Comments 25 pages, 9 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏