arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 3126 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3126 篇

2508.03009 2025-08-06 cs.CV cs.AI 62%

Enhancing Long Video Question Answering with Scene-Localized Frame Grouping

Xuyi Yang, Wenhao Zhang, Hongbo Jin, Lin Liu, Hongbo Xu, Yongwei Nie, Fei Yu, Fei Ma

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21124 2025-07-30 cs.HC cs.AI cs.GR cs.LG 62%

VizGenie: Toward Self-Refining, Domain-Aware Workflows for Next-Generation Scientific Visualization

Ayan Biswas, Terece L. Turton, Nishath Rajiv Ranasinghe, Shawn Jones, Bradley Love, William Jones, Aric Hagberg, Han-Wei Shen, Nathan DeBardeleben, Earl Lawrence

机构 * Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室) Coastal Carolina University(海岸卡罗来纳大学) Ohio State University(俄亥俄州立大学)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09081 2025-07-30 cs.CV cs.AI cs.CL 62%

FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation

Zheqi He, Yesheng Liu, Jing-shu Zheng, Xuejing Li, Jin-Ge Yao, Bowen Qin, Richeng Xuan, Xi Yang

机构 * BAAI FlagEval Team(BAAI 评测团队)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACL 2025 Demo

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10503 2025-07-29 cs.CV cs.CL cs.LG 62%

Everything is a Video: Unifying Modalities through Next-Frame Prediction

G. Thomas Hudson, Dean Slack, Thomas Winterbottom, Jamie Sterling, Chenghao Xiao, Junjie Shentu, Noura Al Moubayed

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments 10 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14544 2025-07-22 cs.CV cs.AI 62%

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025

Sujata Gaihre, Amir Thapa Magar, Prasuna Pokharel, Laxmi Tiwari

机构 * NCIT(尼泊尔信息技术研究所) Fusemachine(Fusemachine公司) Logictronix Technologies(Logictronix Technologies公司)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments accepted to ImageCLEF 2025, to be published in the lab proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07274 2025-07-11 cs.CV cs.AI cs.CL 62%

LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation

Ananya Raval, Aravind Narayanan, Vahid Reza Khazaie, Shaina Raza

机构 * Vector Institute for AI(向量人工智能研究所)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments Accepted at ASONAM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00525 2025-07-02 cs.CV cs.AI 62%

Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving

Djamahl Etchegaray, Yuxia Fu, Zi Huang, Yadan Luo

机构 * The University of Queensland(昆士兰大学)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22567 2025-07-01 cs.CV cs.AI 62%

Unifying Biomedical Vision-Language Expertise: Towards a Generalist Foundation Model via Multi-CLIP Knowledge Distillation

Shansong Wang, Zhecheng Jin, Mingzhe Hu, Mojtaba Safari, Feng Zhao, Chih-Wei Chang, Richard LJ Qiu, Justin Roper, David S. Yu, Xiaofeng Yang

机构 * Department of Radiation Oncology, Winship Cancer Institute, Emory University School of Medicine(放射肿瘤科,Winship癌症研究所,埃默里大学医学院) Department of Biomedical Engineering, College of Engineering, Georgia Institute of Technology(生物医学工程系,工程学院,佐治亚理工学院) Department of Computer Science and Mathematics, Laney Graduate School, Emory University(计算机科学与数学系,兰尼研究生院,埃默里大学) School of Electrical and Computer Engineering, College of Engineering, Georgia Institute of Technology(电气与计算机工程系,工程学院,佐治亚理工学院)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18071 2025-06-30 cs.CV cs.AI 62%

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering

Jisheng Dang, Huilin Song, Junbin Xiao, Bimei Wang, Han Peng, Haoxuan Li, Xun Yang, Meng Wang, Tat-Seng Chua

专题命中 视觉问答 :grounding(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17734 2025-06-23 cs.AI cs.CL cs.CV 62%

Cost-effective Instruction Learning for Pathology Vision and Language Analysis

Kaitao Chen, Mianxin Liu, Fang Yan, Lei Ma, Xiaoming Shi, Lilong Wang, Xiaosong Wang, Lifeng Zhu, Zhe Wang, Mu Zhou, Shaoting Zhang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) School of Computer Science, Fudan University(复旦大学计算机学院) National Biomedical Imaging Center, College of Future Technology, Peking University(北京大学未来技术学院生物医学成像中心) Ruijin Hospital, Shanghai Jiaotong University School of Medicine(上海交通大学医学院瑞金医院) Department of Pathology, State Key Laboratory of Cancer Biology, Xijing Hospital(西京医院病理科,癌症生物学国家重点实验室) Department of Computer Science, Rutgers University(罗格斯大学计算机科学系)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15594 2025-06-19 cs.CL cs.AI cs.LG 62%

WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts

Negar Foroutan, Angelika Romanou, Matin Ansaripour, Julian Martin Eisenschlos, Karl Aberer, Rémi Lebret

机构 * EPFL(苏黎世联邦理工学院) Google DeepMind(谷歌DeepMind) Universidad Nacional de Córdoba(国家科隆大学)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.AI、cs.LG

Comments ACL 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05317 2025-06-17 cs.IR cs.AI cs.CL cs.LG 62%

On Synthesizing Data for Context Attribution in Question Answering

Gorjan Radevski, Kiril Gashteovski, Shahbaz Syed, Christopher Malon, Sebastien Nicolas, Chia-Chien Hung, Timo Sztyler, Verena Heußer, Wiem Ben Rim, Masafumi Enomoto, Kunihiro Takeoka, Masafumi Oyamada, Goran Glavaš, Carolin Lawrence

专题命中 视觉问答 :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04223 2025-06-17 cs.CV cs.AI 62%

VideoQA in the Era of LLMs: An Empirical Study

Junbin Xiao, Nanxin Huang, Hangyu Qin, Dongyang Li, Yicong Li, Fengbin Zhu, Zhulin Tao, Jianxing Yu, Liang Lin, Tat-Seng Chua, Angela Yao

机构 * National University of Singapore(新加坡国立大学) Communication University of China(中国传媒大学) Sun Yat-Sen University(中山大学)

专题命中 视觉问答 :grounding(abstract);分类 cs.CV、cs.AI

Comments IJCV'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12202 2025-06-17 cs.PL cs.AI cs.CR cs.LG 62%

A Fast, Reliable, and Secure Programming Language for LLM Agents with Code Actions

Stephen Mell, Botong Zhang, David Mell, Shuo Li, Ramya Ramalingam, Nathan Yu, Steve Zdancewic, Osbert Bastani

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11394 2025-06-16 cs.CV cs.AI 62%

Dynamic Double Space Tower

Weikai Sun, Shijie Song, Han Wang

机构 * Huawei Cloud Computing(华为云计算)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09958 2025-06-12 cs.CV cs.LG 62%

Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy

Sushant Gautam, Michael A. Riegler, Pål Halvorsen

机构 * Simula Metropolitan Center for Digital Engineering (SimulaMet)(Simula Metropolitan Center for Digital Engineering) Oslo Metropolitan University (OsloMet)(Oslo Metropolitan University) Simula Research Laboratory

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09953 2025-06-12 cs.CV cs.AI cs.CL 62%

Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos

Benjamin Reichman, Constantin Patsch, Jack Truxal, Atishay Jain, Larry Heck

机构 * Georgia Institute of Technology(佐治亚理工学院) Technical University of Munich(慕尼黑技术大学)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09566 2025-06-12 cs.CL cs.AI cs.LG 62%

From Symbolic to Neural and Back: Exploring Knowledge Graph-Large Language Model Synergies

Blaž Škrlj, Boshko Koloski, Senja Pollak, Nada Lavrač

机构 * Jožef Stefan Institute(乔泽夫·斯塔芬研究所)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI、cs.LG

Comments To-appear as a book chapter

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03097 2025-06-04 cs.CV cs.AI 62%

EgoVLM: Policy Optimization for Egocentric Video Understanding

Ashwin Vinod, Shrey Pandit, Aditya Vavre, Linshen Liu

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments Our Code can be found at https://github.com/adityavavre/VidEgoVLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02167 2025-06-04 cs.CV cs.AI 62%

Fire360: A Benchmark for Robust Perception and Episodic Memory in Degraded 360-Degree Firefighting Videos

Aditi Tiwari, Farzaneh Masoud, Dac Trong Nguyen, Jill Kraft, Heng Ji, Klara Nahrstedt

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Illinois Fire Service Institute(伊利诺伊州消防服务研究所)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments 20 pages, 9 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03730 2025-06-04 cs.LG cs.CR cs.CV 62%

NeurIPS 2023 Competition: Privacy Preserving Federated Learning Document VQA

Marlon Tobaben, Mohamed Ali Souibgui, Rubèn Tito, Khanh Nguyen, Raouf Kerkouche, Kangsoo Jung, Joonas Jälkö, Lei Kang, Andrey Barsky, Vincent Poulain d'Andecy, Aurélie Joseph, Aashiq Muhamed, Kevin Kuo, Virginia Smith, Yusuke Yamasaki, Takumi Fukami, Kenta Niwa, Iifan Tyou, Hiro Ishii, Rio Yokota, Ragul N, Rintu Kutum, Josep Llados, Ernest Valveny, Antti Honkela, Mario Fritz, Dimosthenis Karatzas

机构 * University of Helsinki(赫尔辛基大学) Computer Vision Center, Universitat Autònoma de Barcelona(巴塞罗那自治大学计算机视觉中心) CISPA Helmholtz Center for Information Security(信息安全赫尔姆霍茨中心) INRIA(法国国家信息与自动化技术研究所) Yooz Carnegie Mellon University(卡内基梅隆大学) NTT(日本NTT公司) Institute of Science Tokyo(东京科学研究所) Department of Computer Science(计算机科学系) Mphasis AI & Applied Tech Lab at Ashoka, Ashoka University(阿什oka大学人工智能与应用技术实验室) Koita Centre for Digital Health - Ashoka (KCDH-A)(阿什oka数字健康中心(KCDH-A)) Trivedi School of Biosciences, Ashoka University(阿什oka大学Trivedi生物科学学院)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments 33 pages, 7 figures; published in TMLR 06/2025 https://openreview.net/forum?id=3HKNwejEEq

Journal ref Transactions on Machine Learning Research, ISSN 2835-8856, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17794 2025-06-02 cs.CV cs.AI 62%

Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models

Ketan Suhaas Saichandran, Xavier Thomas, Prakhar Kaushik, Deepti Ghadiyaram

机构 * Boston University(波士顿大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI

Comments Accepted at CVPR 2025 workshops (AI4CC (oral) & GMCV (poster))

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01645 2025-05-14 cs.CV cs.AI 62%

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding

Heqing Zou, Tianze Luo, Guiyang Xie, Victor Xiao Jie Zhang, Fengmao Lv, Guangcong Wang, Junyang Chen, Zhuochen Wang, Hansheng Zhang, Huaijian Zhang

机构 * ByteDance(字节跳动) Nanyang Technological University(南洋理工大学) Southwest Jiaotong University(西南交通大学) Great Bay University(大湾大学) Shenzhen University(深圳大学)

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments Accepted to ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06814 2025-05-13 cs.CV cs.AI cs.CL 62%

Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge

Bin Li, Shenxi Liu, Yixuan Weng, Yue Du, Yuhang Tian, Shoujun Zhou

机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究所,中国科学院) School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) School of Engineering, Westlake University(工程学院,西湖大学)

专题命中 视觉问答 :grounding(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04147 2025-05-08 cs.CV cs.AI 62%

R^3-VQA: "Read the Room" by Video Social Reasoning

Lixing Niu, Jiapeng Li, Xingping Yu, Shu Wang, Ruining Feng, Bo Wu, Ping Wei, Yisen Wang, Lifeng Fan

机构 * School of Intelligence Science and Technology, Peking University(北京理工大学智能科学与技术学院) College of Artificial Intelligence, Xi’an Jiaotong University(西安交通大学人工智能学院) State Key Laboratory of General Artificial Intelligence, Beijing Institute for General Artificial Intelligence(一般人工智能国家重点实验室) Yuanpei College, Peking University(北京大学元培学院) Tsinghua University(清华大学) University of California, Los Angeles(加州大学洛杉矶分校) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03883 2025-04-24 cs.CL cs.AI cs.LG 62%

MEG: Medical Knowledge-Augmented Large Language Models for Question Answering

Laura Cabello, Carmen Martin-Turrero, Uchenna Akujuobi, Anders Søgaard, Carlos Bobed

机构 * University of Copenhagen(哥本哈根大学) Sony AI(索尼人工智能) University of Zaragoza(阿拉维萨大学)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00622 2025-04-24 cs.CV cs.AI 62%

Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering

Xingrui Wang, Wufei Ma, Angtian Wang, Shuo Chen, Adam Kortylewski, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学) Tsinghua University(清华大学) Max Planck Institute for Informatics(马克斯·普朗克研究所(信息学)) University of Freiburg(弗赖堡大学)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV、cs.AI

Comments ICLR 2025 accepted paper. Project url: https://xingruiwang.github.io/projects/DynSuperCLEVR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11363 2025-04-21 cs.AI cs.CE cs.LG q-bio.BM 62%

ProteinGPT: Multimodal LLM for Protein Property Prediction and Structure Understanding

Yijia Xiao, Edward Sun, Yiqiao Jin, Qifan Wang, Wei Wang

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.AI、cs.LG

Comments Spotlight, Machine Learning for Genomics Explorations @ ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00584 2025-04-18 cs.CV cs.LG 62%

Online Video Understanding: OVBench and VideoChat-Online

Zhenpeng Huang, Xinhao Li, Jiaqi Li, Jing Wang, Xiangyu Zeng, Cheng Liang, Tao Wu, Xi Chen, Liang Li, Limin Wang

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV、cs.LG

Comments CVPR 2025 Camera Ready Version. Project Page: https://videochat-online.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11777 2025-04-17 cs.CV cs.LG 62%

Bridging the Semantic Gaps: Improving Medical VQA Consistency with LLM-Augmented Question Sets

Yongpei Ma, Pengyu Wang, Adam Dunn, Usman Naseem, Jinman Kim

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.LG

Comments The first two listed authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏