arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 3122 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3122 篇

2511.14567 2025-11-24 cs.HC cs.AI 79%

SweeperBot: Making 3D Browsing Accessible through View Analysis and Visual Question Answering

SweeperBot:通过视图分析和视觉问答实现3D浏览的可访问性

Chen Chen, Cuong Nguyen, Alexa Siu, Dingzeyu Li, Nadir Weibel

机构 * Florida International University(佛罗里达国际大学) Adobe Research(Adobe研究) University of California San Diego(加州大学圣地亚哥分校)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.AI

AI总结 SweeperBot通过视图分析和视觉问答技术,帮助屏幕阅读器用户更有效地探索和比较3D模型。

Comments 28 pages, 16 figures, this is an original manuscript of an article published by Taylor & Francis in the International Journal of Human-Computer Interaction (IJHCI), available online: https://doi.org/10.1080/10447318.2025.2594750

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11198 2025-11-17 cs.CV 79%

Geospatial Chain of Thought Reasoning for Enhanced Visual Question Answering on Satellite Imagery

Shambhavi Shanker, Manikandan Padmanaban, Jagabondhu Hazra

机构 * Indian Institute of Technology Bombay(印度理工学院班加罗尔) IBM Research India(IBM印度研究)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09058 2025-11-13 cs.CV 79%

VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering

Hai-Dang Nguyen, Minh-Anh Dang, Minh-Tan Le, Minh-Tuan Le

机构 * Faculty Of Information Technology, VNU University of Engineering and Technology(信息科技学院,越南工程与技术大学) IT-BT Convergence Technology Division, Vietnam-Korea Institute of Science and Technology(IT-BT融合技术部,越南-韩国科学技术院) TADI Global Lab, TADI Global Company Limited(TADI全球实验室,TADI全球公司) Faculty of Finance, Banking Academy of Vietnam(金融学院,越南银行学院)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

Comments 7 pages, 3 figures, 3 tables, FAIR 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03617 2025-11-06 cs.GR cs.AI 79%

Visualization Biases MLLM's Decision Making in Network Data Tasks

Timo Brand, Henry Förster, Stephen G. Kobourov, Jacob Miller

机构 * Technical University of Munich, Heilbronn, Germany(慕尼黑技术大学)

专题命中 视觉问答 :MLLM(title,abstract);分类 cs.AI

Comments This manuscript was presented at VIS x GenAI, a workshop co-located with IEEE VIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00389 2025-11-04 cs.CV 79%

Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond

Fan Zhang, Haoxuan Li, Shengju Qian, Xin Wang, Zheng Lian, Hao Wu, Zhihong Zhu, Yuan Gao, Qiankun Li, Yefeng Zheng, Zhouchen Lin, Pheng-Ann Heng

机构 * The Chinese University of Hong Kong(香港中文大学) Peking University(北京大学) Tencent(腾讯) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学) Westlake University(西湖大学)

专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22803 2025-10-28 cs.CV 79%

MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering

Hai-Dang Nguyen, Minh-Anh Dang, Minh-Tan Le, Minh-Tuan Le

机构 * Faculty Of Information Technology VNU University of Engineering(信息科技学院越南工程大学) IT-BT Convergence Technology Division Vietnam-Korea Institute of Science(IT-BT融合技术部门越南-韩国科学技术院) TADI Global Lab TADI Global Company Limited(TADI全球实验室TADI全球有限公司) Faculty of Finance Banking Academy of Vietnam(金融学院越南银行学院)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

Comments 10 pages, 4 figures, IEEE conference format

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08974 2025-10-27 cs.CV 79%

Text-conditioned State Space Model For Domain-generalized Change Detection Visual Question Answering

Elman Ghazaei, Erchan Aptoula

机构 * Faculty of Engineering and Natural Sciences (VPALab)(工程与自然科学学院(VPALab)) Sabanci University(萨班奇大学)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07447 2025-10-15 cs.CV 79%

In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting

Taiying Peng, Jiacheng Hua, Miao Liu, Feng Lu

机构 * State Key Laboratory of VR Technology and Systems, School of CSE, Beihang University(虚拟现实技术与系统国家重点实验室,北京航空航天大学计算机科学与工程学院) College of AI, Tsinghua University(清华大学人工智能学院)

专题命中 视觉问答 :MLLM(title);multimodal large language model(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11645 2025-10-14 cs.AI 79%

Adapting and Evaluating Multimodal Large Language Models for Adolescent Idiopathic Scoliosis Self-Management: A Divide and Conquer Framework

Zhaolong Wu, Pu Luo, Nan Meng, Jason Pui Yin Cheung, Teng Zhang

机构 * Department of Orthopaedics and Traumatology, The University of Hong Kong(骨科与创伤外科部,香港大学)

专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.AI

Comments Accepted by MICCAI 2025 MLLMCP Workshop

Journal ref Lecture Notes in Computer Science 16147 (2026) 280-289

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08791 2025-10-13 cs.CV 79%

Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering

Yuanhao Zou, Zhaozheng Yin

机构 * University of Michigan(密歇根大学) Stony Brook University(石溪大学)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

Comments CVPR2025 Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03903 2025-10-07 cs.CV 79%

Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models

Md. Atabuzzaman, Andrew Zhang, Chris Thomas

机构 * Department of Computer Science(计算机科学系) Virginia Tech(弗吉尼亚理工大学)

专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03232 2025-10-06 cs.CV 79%

LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models

Ci-Siang Lin, Min-Hung Chen, Yu-Yang Sheng, Yu-Chiang Frank Wang

机构 * Graduate Institute of Communication Engineering, National Taiwan University, Taiwan(台湾国立台湾大学通信工程研究所) NVIDIA

专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23899 2025-10-06 cs.CV 79%

Q-FSRU: Quantum-Augmented Frequency-Spectral For Medical Visual Question Answering

Rakesh Thakur, Yusra Tariq, Rakesh Chandra Joshi

机构 * Amity University(阿米蒂大学)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

Comments 12 pages (9 main + 2 references/appendix), 2 figures, conference paper submitted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17974 2025-09-23 cs.CL cs.CV 79%

Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts

Xuyang Wu, Yuan Wang, Hsin-Tai Wu, Zhiqiang Tao, Yi Fang

机构 * Santa Clara University(圣克拉拉大学) DOCOMO Innovations, Inc.(DOCOMO创新公司) Rochester Institute of Technology(罗切斯特理工学院)

专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV

Comments EMNLP Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14886 2025-09-19 cs.CL cs.AI 79%

A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation

Ye Shen, Junying Wang, Farong Wen, Yijin Guo, Qi Jia, Zicheng Zhang, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学)

专题命中 视觉问答 :MLLM(title,abstract);分类 cs.AI

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06010 2025-09-09 cs.CV 79%

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users

Wanyin Cheng, Zanxi Ruan

机构 * School of Cyber Science and Engineering, Qufu Normal University(网络科学与工程学院,曲阜师范大学) Department of Computer Science, University of Verona(计算机科学系,威尼斯大学)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04502 2025-09-08 cs.CL cs.AI 79%

VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples

Qixin Sun, Ziqin Wang, Hengyuan Zhao, Yilin Li, Kaiyou Song, Linjiang Huang, Xiaolin Hu, Qingpei Guo, Si Liu

专题命中 视觉问答 :multimodal large language model(title);visual question answering(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19887 2025-08-28 cs.CL cs.CV 79%

Bangla-Bayanno: A 52K-Pair Bengali Visual Question Answering Dataset with LLM-Assisted Translation Refinement

Mohammed Rakibul Hasan, Rafi Majid, Ahanaf Tahmid

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12404 2025-08-19 cs.CV 79%

LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving

Nan Song, Bozhou Zhang, Xiatian Zhu, Jiankang Deng, Li Zhang

专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV

Comments 7 pages, 4 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05262 2025-08-15 cs.CV 79%

Debiasing Multimodal Large Language Models via Penalization of Language Priors

YiFan Zhang, Yang Shi, Weichen Yu, Qingsong Wen, Xue Wang, Wenjing Yang, Zhang Zhang, Liang Wang, Rong Jin

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Peking University(北京大学) Carnegie Mellon University(卡内基梅隆大学) Alibaba Group(阿里巴巴集团) National University of Defense Technology(国防科技大学) Meta

专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV

Comments 10 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00549 2025-08-11 cs.CV 79%

Your other Left! Vision-Language Models Fail to Identify Relative Positions in Medical Images

Daniel Wolf, Heiko Hillenhagen, Billurvan Taskin, Alex Bäuerle, Meinrad Beer, Michael Götz, Timo Ropinski

机构 * Visual Computing Group, Institute of Media Informatics, Ulm University, Germany(媒体信息研究所视觉计算组,乌尔姆大学,德国) Diagnostic and Interventional Radiology, Ulm University Medical Center, Germany(乌尔姆大学医学中心诊断与介入放射学) Axiom Bio, USA(Axiom Bio公司,美国)

专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV

Comments Accepted at the International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18351 2025-08-08 cs.CL cs.AI 79%

Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering

Zhongjian Hu, Peng Yang, Bing Li, Zhenqi Wang

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.AI

Comments We would like to withdraw this submission due to ongoing internal review and coordination among the author team. Upon the supervisor's recommendation, we have decided to delay public dissemination until the manuscript undergoes further refinement and aligns with our intended academic trajectory

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16936 2025-08-08 cs.CL cs.AI 79%

Rationale-guided Prompting for Knowledge-based Visual Question Answering

Zhongjian Hu, Peng Yang, Bing Li, Fengyuan Liu

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.AI

Comments We would like to withdraw this submission due to ongoing internal review and coordination among the author team. Upon the supervisor's recommendation, we have decided to delay public dissemination until the manuscript undergoes further refinement and aligns with our intended academic trajectory

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17050 2025-07-24 cs.CV 79%

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models

Tz-Ying Wu, Tahani Trigui, Sharath Nittur Sridhar, Anand Bodas, Subarna Tripathi

机构 * Intel(英特尔)

专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV

Comments Accepted to CVAM Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11200 2025-07-21 cs.CV 79%

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study

Che Liu, Jiazhen Pan, Weixiang Shen, Wenjia Bai, Daniel Rueckert, Rossella Arcucci

机构 * Imperial College London, UK(伦敦帝国学院) Technical University of Munich, Germany(慕尼黑技术大学) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)

专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08036 2025-07-15 cs.CL cs.CV 79%

Barriers in Integrating Medical Visual Question Answering into Radiology Workflows: A Scoping Review and Clinicians' Insights

Deepali Mishra, Chaklam Silpasuwanchai, Ashutosh Modi, Madhumita Sushil, Sorayouth Chumnanvej

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

Comments 29 pages, 5 figures (1 in supplementary), 3 tables (1 in main text, 2 in supplementary). Scoping review and clinician survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04333 2025-07-08 cs.CV cs.CL 79%

Computed Tomography Visual Question Answering with Cross-modal Feature Graphing

Yuanhe Tian, Chen Su, Junwen Duan, Yan Song

机构 * University of Washington(华盛顿大学) University of Science and Technology of China(中国科学技术大学) Central South University(中南大学)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10798 2025-07-01 cs.CV 79%

Object Retrieval for Visual Question Answering with Outside Knowledge

Shichao Kan, Yuhai Deng, Jiale Fu, Lihui Cen, Zhe Qu, Linna Zhang, Yixiong Liang, Yigang Cen

机构 * School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院) School of Automation, Central South University(中南大学自动化学院) Dundee International Institute, Central South University(中南大学邓迪国际学院) College of Mechanical Engineering, Guizhou University(贵州大学机械工程学院) Institute of Information Science, School of Computer and Information Technology, Beijing Jiaotong University(北京交通大学计算机与信息学院信息科学研究所) Beijing Key Laboratory of Advanced Information Science and Network Technology(北京高级信息科学与网络技术重点实验室)

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20362 2025-06-30 cs.CV 79%

Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding

Joao Pereira, Vasco Lopes, David Semedo, Joao Neves

机构 * DeepNeuronic NOVA LINCS, University of Beira Interior(NOVA LINCS,贝拉内里大学)

专题命中 视觉问答 :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19367 2025-06-30 cs.CV 79%

VGAT: A Cancer Survival Analysis Framework Transitioning from Generative Visual Question Answering to Genomic Reconstruction

Zizhi Chen, Minghao Han, Xukun Zhang, Shuwei Ma, Tao Liu, Xing Wei, Lihua Zhang

机构 * 1 Academy for Engineering Technology, Fudan University, Shanghai, China 2 Institute of Metaverse \& Intelligent Medicine, Fudan University, Shanghai, China 3 Engineering Research Center of AI 4 Jilin Provincial Key Laboratory of Intelligence Science

专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏