arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 3129 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3129 篇

2506.15635 2025-09-30 cs.CV cs.RO 57%

FindingDory: A Benchmark to Evaluate Memory in Embodied Agents

Karmesh Yadav, Yusuf Ali, Gunshi Gupta, Yarin Gal, Zsolt Kira

机构 * Georgia Tech(佐治亚理工学院) University of Oxford(牛津大学)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments Our dataset and code can be found at: https://findingdory-benchmark.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20858 2025-09-26 cs.GR cs.CV cs.MM 57%

ArchGPT: Understanding the World's Architectures with Large Multimodal Models

Yuze Wang, Luo Yang, Junyi Wang, Yue Qi

机构 * State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室) School of Computer Science and Engineering(计算机科学与工程学院) Beihang University(北京航空航天大学) School of Computer Science and Technology(计算机科学与技术学院) Shandong University(山东大学)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04931 2025-09-25 cs.CV 57%

Long Video Understanding with Learnable Retrieval in Video-Language Models

Jiaqi Xu, Cuiling Lan, Wenxuan Xie, Xuejin Chen, Yan Lu

机构 * School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 视觉问答 :VLM(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Multimedia (TMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14267 2025-09-19 cs.CL cs.AI cs.IT math.IT 57%

Graph-Enhanced Retrieval-Augmented Question Answering for E-Commerce Customer Support

Piyushkumar Patel

机构 * Microsoft(微软)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14227 2025-09-18 cs.CV 57%

Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark

Nisarg A. Shah, Amir Ziai, Chaitanya Ekanadham, Vishal M. Patel

机构 * Netflix, Inc.(Netflix公司) Johns Hopkins University(约翰霍普金斯大学)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments 11 pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11589 2025-09-16 cs.CV 57%

MVQA-68K: A Multi-dimensional and Causally-annotated Dataset with Quality Interpretability for Video Assessment

Yanyun Pu, Kehan Li, Zeyi Huang, Zhijie Zhong, Kaixiang Yang

机构 * Huawei Technologies Co.(华为技术有限公司) South China University of Technology(南方科技大学)

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05479 2025-09-16 cs.CV 57%

LATTE: Learning to Think with Vision Specialists

Zixian Ma, Jianguo Zhang, Zhiwei Liu, Jieyu Zhang, Juntao Tan, Manli Shu, Juan Carlos Niebles, Shelby Heinecke, Huan Wang, Caiming Xiong, Ranjay Krishna, Silvio Savarese

机构 * University of Washington(华盛顿大学) Salesforce Research(Salesforce研究)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07846 2025-09-10 cs.AI 57%

Aligning LLMs for the Classroom with Knowledge-Based Retrieval -- A Comparative RAG Study

Amay Jain, Liu Cui, Si Chen

机构 * Student(学生) Downingtown STEM Academy Department of Computer Science(计算机科学系) West Chester University of Pennsylvania(宾夕法尼亚州韦斯特切斯特大学)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02864 2025-09-04 cs.CL cs.AI 57%

A-SEA3L-QA: A Fully Automated Self-Evolving, Adversarial Workflow for Arabic Long-Context Question-Answer Generation

Kesen Wang, Daulet Toibazar, Pedro J. Moreno

机构 * Humain

专题命中 视觉问答 :vision language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19724 2025-08-29 cs.CL cs.AI 57%

NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks

Aritra Dutta, Swapnanil Mukherjee, Deepanway Ghosal, Somak Aditya

机构 * IIT Kharagpur(印度理工学院Kharagpur分校) Ashoka University(阿什oka大学)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16081 2025-08-25 cs.CL cs.LG 57%

CEQuest: Benchmarking Large Language Models for Construction Estimation

Yanzhao Wu, Lufan Wang, Rui Liu

机构 * Florida International University(佛罗里达国际大学) University of Florida(佛罗里达大学)

专题命中 视觉问答 :LLaVA(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12663 2025-08-19 cs.CV 57%

Stable Diffusion-Based Approach for Human De-Occlusion

Seung Young Noh, Ju Yong Chang

机构 * Kwangwoon University(韩国成均馆大学)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12108 2025-08-19 cs.CV 57%

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine

Ziyang Zhang, Yang Yu, Xulei Yang, Si Yong Yeo

机构 * MedVisAI Lab Department of ECE Northwestern University(MedVisAI实验室 电子工程系 西北大学) Institute for Infocomm Research (I 2 R) A*STAR, Singapore(信息与通信研究所(I 2 R)A*STAR,新加坡) MedVisAI Lab Lee Kong Chian School of Medicine, Nanyang Technological University(MedVisAI实验室 李科田医学院,南洋理工大学)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04630 2025-08-19 cs.CV 57%

Learn 3D VQA Better with Active Selection and Reannotation

Shengli Zhou, Yang Liu, Feng Zheng

机构 * Southern University of Science and Technology(南方科技大学) Wangxuan Institute of Computer Technology, Peking University(北京大学王瑄计算机技术研究所)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 13 pages, 16 figures, accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.01955 2025-08-19 cs.CV 57%

Adaptively Clustering Neighbor Elements for Image-Text Generation

Zihua Wang, Xu Yang, Hanwang Zhang, Haiyang Xu, Ming Yan, Fei Huang, Yu Zhang

机构 * School of Computer Science and Engineering, Nanyang Technological University(计算机科学与工程学院,南洋理工大学) DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments This work has been accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16636 2025-08-18 cs.CL cs.CV 57%

Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries

Yin Wu, Quanyu Long, Jing Li, Jianfei Yu, Wenya Wang

专题命中 视觉问答 :grounding(abstract);分类 cs.CV

Comments 21 pages, 6 figures, 17 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08224 2025-08-14 cs.CL cs.AI 57%

Capabilities of GPT-5 on Multimodal Medical Reasoning

Shansong Wang, Mingzhe Hu, Qiang Li, Mojtaba Safari, Xiaofeng Yang

机构 * Department of Radiation Oncology, Winship Cancer Institute, Emory University School of Medicine(放射肿瘤科、Winship癌症研究所、埃默里大学医学院)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI

Comments Corrected some typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01422 2025-08-12 cs.CV 57%

DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes

Zhende Song, Chenchen Wang, Jiamu Sheng, Chi Zhang, Shengji Tang, Jiayuan Fan, Tao Chen

机构 * Fudan University(复旦大学) Westlake University, Tencent PCG(西湖大学、腾讯PCG)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments Accepted by ACMM'MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06152 2025-08-11 cs.CV 57%

VISTAR:A User-Centric and Role-Driven Benchmark for Text-to-Image Evaluation

Kaiyuan Jiang, Ruoxi Sun, Ying Cao, Yuqi Xu, Xinran Zhang, Junyan Guo, ChengSheng Deng

机构 * Peking University(北京大学) LinkSure University of Glasgow(格拉斯哥大学) Boston University(波士顿大学)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments 17 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08131 2025-08-11 cs.CV 57%

SAR Strikes Back: A New Hope for RSVQA

Lucrezia Tosato, Flora Weissgerber, Laurent Wendling, Sylvain Lobry

机构 * LIPADE, Université Paris Cité(巴黎Cité大学LIPADE研究所) GENCI-IDRIS French Aerospace Lab, ONERA(法国航空航天实验室ONERA)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted at IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 13 pages, 6 figures

Journal ref 10.1109/JSTARS.2025.3596678

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03488 2025-08-06 cs.AI cs.SE 57%

VQA support to Arabic Language Learning Educational Tool

Khaled Bachir Delassi, Lakhdar Zeggane, Hadda Cherroun, Abdelhamid Haouhat, Kaoutar Bouzouad

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20466 2025-08-06 cs.CV 57%

LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMs

Woo Yi Yang, Jiarui Wang, Sijing Wu, Huiyu Duan, Yuxin Zhu, Liu Yang, Kang Fu, Guangtao Zhai, Xiongkuo Min

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02082 2025-08-05 cs.CV 57%

S-RRG-Bench: Structured Radiology Report Generation with Fine-Grained Evaluation Framework

Yingshu Li, Yunyi Liu, Zhanyu Wang, Xinyu Liang, Lingqiao Liu, Lei Wang, Luping Zhou

机构 * University of Sydney(悉尼大学) University of Wollongong(沃林戈大学) University of Adelaide(阿德莱德大学) First Clinical Medical College, Guangzhou University of Chinese Medicine(广州中医药大学第一临床学院)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22938 2025-08-01 cs.CL cs.AI 57%

A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents

Sumit Soman, H. G. Ranjani, Sujoy Roychowdhury, Venkata Dharma Surya Narayana Sastry, Akshat Jain, Pranav Gangrade, Ayaaz Khan

机构 * Ericsson R&D Bangalore Karnataka India(爱立信研发部班加罗尔卡纳塔克邦印度)

专题命中 视觉问答 :VLM(abstract);分类 cs.AI

Comments Accepted for publication at the KDD 2025 Workshop on Structured Knowledge for Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02201 2025-08-01 cs.CV 57%

Acknowledging Focus Ambiguity in Visual Questions

Chongyan Chen, Yu-Yun Tseng, Zhuoheng Li, Anush Venkatesh, Danna Gurari

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Colorado Boulder(科罗拉多大学博尔德分校)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21165 2025-07-30 eess.IV cs.CV 57%

Querying GI Endoscopy Images: A VQA Approach

Gaurav Parajuli

机构 * Johannes Kepler University Linz(约翰·凯撒大学林茨)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20939 2025-07-29 cs.CV 57%

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Yuying Ge, Yixiao Ge, Chen Li, Teng Wang, Junfu Pu, Yizhuo Li, Lu Qiu, Jin Ma, Lisheng Duan, Xinyu Zuo, Jinwen Luo, Weibo Gu, Zexuan Li, Xiaojing Zhang, Yangyu Tao, Han Hu, Di Wang, Ying Shan

机构 * ARC Lab, Tencent PCG(腾讯PCG ARC实验室) Search Application Department, Tencent CSIG(腾讯CSIG搜索应用部门) Tencent Hunyuan(腾讯文生视频) Big Data Platform Department, Tencent PCG(腾讯PCG大数据平台部门)

专题命中 视觉问答 :grounding(abstract);分类 cs.CV

Comments Project Page: https://tencentarc.github.io/posts/arc-video-announcement/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19882 2025-07-29 cs.AI 57%

Causality-aligned Prompt Learning via Diffusion-based Counterfactual Generation

Xinshu Li, Ruoyu Wang, Erdun Gao, Mingming Gong, Lina Yao

机构 * The University of New South Wales(新南威尔士大学) The University of Adelaide(阿德莱德大学) The University of Melbourne(墨尔本大学) CSIRO’s Data 61(CSIRO数据61)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18910 2025-07-28 cs.CL cs.LG 57%

A Systematic Review of Key Retrieval-Augmented Generation (RAG) Systems: Progress, Gaps, and Future Directions

Agada Joseph Oche, Ademola Glory Folashade, Tirthankar Ghosal, Arpan Biswas

机构 * Bredesen Center for Interdisciplinary Research, University of Tennessee, Knoxville, USA(跨学科研究中心,田纳西大学,肯塔基州 Knoxville) National Center for Computational Sciences, Oak Ridge National Laboratory, Oak Ridge, USA(计算科学国家中心,橡树岭国家实验室,橡树岭) University of Tennessee-Oak Ridge Innovation Institute, University of Tennessee, Knoxville, USA(田纳西大学橡树岭创新研究所,田纳西大学,肯塔基州 Knoxville)

专题命中 视觉问答 :grounding(abstract);分类 cs.LG

Comments 33 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06303 2025-07-28 cs.CL cs.CV 57%

Long-Form Answers to Visual Questions from Blind and Low Vision People

Mina Huh, Fangyuan Xu, Yi-Hao Peng, Chongyan Chen, Hansika Murugu, Danna Gurari, Eunsol Choi, Amy Pavel

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Carnegie Mellon University(卡内基梅隆大学) Hong Kong University of Science and Technology(香港科学与技术大学) University of Colorado Boulder(科罗拉多大学博尔德分校)

专题命中 视觉问答 :vision language model(abstract);分类 cs.CV

Comments COLM 2024 Oral Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏