arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 3129 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3129 篇

2506.13589 2025-11-25 cs.CV 57%

AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding

AdaVideoRAG:多情境自适应检索增强高效长视频理解

Zhucun Xue, Jiangning Zhang, Xurong Xie, Yuxuan Cai, Yong Liu, Xiangtai Li, Dacheng Tao

机构 * Zhejiang University(浙江大学) Youtu Lab, Tencent(腾讯优图实验室) Huazhong University of Science and Technolog(华中科技大学) Nanyang Technological University(南洋理工大学)

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV

AI总结 AdaVideoRAG通过自适应检索增强框架提升长视频理解效率与准确性,支持多层级知识检索与深度语义分析。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17943 2025-11-25 cs.CV 57%

SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System

SciEducator: 基于Deming循环多智能体系统的科学视频理解与教育

Zhiyu Xu, Weilong Yan, Yufei Shi, Xin Meng, Tao He, Huiping Zhuang, Ming Li, Hehe Fan

机构 * Jinan University(济南大学) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) Peking University(北京大学) University of Electronic Science and Technology of China(电子科技大学) South China University of Technology(华南理工大学) Guangming Laboratory(光明实验室) Zhejiang University(浙江大学)

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV

AI总结 SciEducator通过Deming循环多智能体系统实现科学视频的自演化理解与教育,生成多模态教学内容并超越现有大语言模型和视频智能体。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12530 2025-11-18 cs.CV 57%

ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding

Yuan Zhou, Litao Hua, Shilong Jin, Wentao Huang, Haoran Duan

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments Accepted to AAAI 2026. Code is available at: https://github.com/robin-hlt/AAAI26-ReaSon

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10810 2025-11-18 cs.CV 57%

SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding

Zhenyu Yang, Yuhang Hu, Zemin Du, Dizhan Xue, Shengsheng Qian, Jiahong Wu, Fan Yang, Weiming Dong, Changsheng Xu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) Kuaishou Technology(快手科技) Zhengzhou University(郑州大学) ShanghaiTech University(上海科技大学) Peng Cheng Laboratory(鹏城实验室)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments ICLR 2025 Accepted (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11597 2025-11-18 cs.AI cs.CL 57%

CLINB: A Climate Intelligence Benchmark for Foundational Models

Michelle Chen Huebscher, Katharine Mach, Aleksandar Stanić, Markus Leippold, Ben Gaiarin, Zeke Hausfather, Elisa Rawat, Erich Fischer, Massimiliano Ciaramita, Joeri Rogelj, Christian Buck, Lierni Sestorain Saralegui, Reto Knutti

机构 * University of Miami(迈阿密大学) University of Zurich(苏黎世大学) Stripe(Stripe公司) ETH Zurich(苏黎世联邦理工学院) Imperial College London(伦敦帝国理工学院)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI

Comments Questions, system prompt and model judge prompts available here: https://www.kaggle.com/datasets/deepmind/clinb-questions

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09868 2025-11-14 cs.CV 57%

Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience Strategies

Peng Gao, Yujian Lee, Xiaofeng Zhang, Zailong Chen, Hui Zhang

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments Accepted in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04369 2025-11-14 cs.CV 57%

TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding

Canhui Tang, Zifan Han, Hongbo Sun, Sanping Zhou, Xuchong Zhang, Xin Wei, Ye Yuan, Huayu Zhang, Jinglin Xu, Hao Sun

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09474 2025-11-13 cs.CV 57%

Surgical AI Copilot: Energy-Based Fourier Gradient Low-Rank Adaptation for Surgical LLM Agent Reasoning and Planning

Jiayuan Huang, Runlong He, Danyal Zaman Khan, Evangelos B. Mazomenos, Danail Stoyanov, Hani Marcus, Linzhe Jiang, Matthew J Clarkson, Mobarak I. Hoque

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07290 2025-11-11 eess.IV cs.CV cs.MM 57%

CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video

Xinyi Wang, Angeliki Katsenou, Junxiao Shen, David Bull

机构 * School of Computer Science, University of Bristol(布里斯托大学计算机科学学院)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06441 2025-11-11 cs.CL cs.LG 57%

Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models

Mayank Saini, Arit Kumar Bishwas

机构 * PwC US(普华永道美国)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.LG

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05393 2025-11-10 cs.CV 57%

PreResQ-R1: Towards Fine-Grained Rank-and-Score Reinforcement Learning for Visual Quality Assessment via Preference-Response Disentangled Policy Optimization

Zehui Feng, Tian Qiu, Tong Wu, Junxuan Li, Huayuan Xu, Ting Han

机构 * Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学)

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV

Comments 27 pages, 14 figures, under review as a conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04502 2025-11-07 cs.CL cs.AI 57%

RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG

Joshua Gao, Quoc Huy Pham, Subin Varghese, Silwal Saurav, Vedhus Hoskere

机构 * University of Houston(德克萨斯大学休斯敦分校)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03325 2025-11-07 cs.CV 57%

SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding

Mauro Orazio Drago, Luca Carlini, Pelinsu Celebi Balyemez, Dennis Pierantozzi, Chiara Lena, Cesare Hassan, Danail Stoyanov, Elena De Momi, Sophia Bano, Mobarak I. Hoque

机构 * Dipartimento di Elettronica, Informazione e Bioingegneria (DEIB)(电子、信息与生物工程系) Politecnico di Milano(米兰理工大学) IRCCS Humanitas Research Hospital(IRCCS人类itas研究医院) UCL Hawkes Institute and Department of Computer Science(UCL Hawkes研究所和计算机科学系) University College London(伦敦大学学院) University of Manchester(曼彻斯特大学)

专题命中 视觉问答 :visual reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03178 2025-11-06 cs.CV 57%

SurgAnt-ViVQA: Learning to Anticipate Surgical Events through GRU-Driven Temporal Cross-Attention

Shreyas C. Dhake, Jiayuan Huang, Runlong He, Danyal Z. Khan, Evangelos B. Mazomenos, Sophia Bano, Hani J. Marcus, Danail Stoyanov, Matthew J. Clarkson, Mobarak I. Hoque

机构 * UCL Hawkes Institute(UCL哈维斯研究所) University College London(伦敦大学学院) Dept of Medical Physics & Biomedical Engineering(医学物理与生物医学工程系) UCL(伦敦大学学院) Dept of Computer Science(计算机科学系) National Hospital for Neurology and Neurosurgery(神经病学与神经外科国家医院) Division of Informatics, Imaging and Data Science(信息学、成像与数据科学 division)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12926 2025-11-05 eess.SP cs.CV 57%

Task-Oriented Feature Compression for Multimodal Understanding via Device-Edge Co-Inference

Cheng Yuan, Zhening Liu, Jiashu Lv, Jiawei Shao, Yufei Jiang, Jun Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究所) School of Electronic and Information Engineering, Harbin Institute of Technology(哈尔滨工业大学电子与信息工程学院) Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(香港科技大学电子与计算机工程系) School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Mobile Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08564 2025-11-05 cs.CV cs.CL 57%

Visual Program Distillation with Template-Based Augmentation

Michal Shlapentokh-Rothman, Yu-Xiong Wang, Derek Hoiem

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments EMNLP Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07769 2025-11-04 cs.CV 57%

BiMediX2: Bio-Medical EXpert LMM for Diverse Medical Modalities

Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Sara Pieri, Saeed Yahya Alseiari, Shanavas Cholakkal, Khaled Aldahmani, Fahad Khan, Rao Anwer, Salman Khan, Timothy Baldwin, Hisham Cholakkal

机构 * Mohamed Bin Zayed University of Artificial Intelligence(Mohamed Bin Zayed大学人工智能学院) Linköping University(林肯大学) Shaikh Tahnoon bin Mohammed Medical City(Shaikh Tahnoon bin Mohammed医疗城) Tawam Hospital(Tawam医院) Sheikh Shakhbout Medical City(Sheikh Shakhbout医疗城) Govt Medical College Kozhikode(科钦政府医学院)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments Accepted to EMNLP 2025 (Findings)

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, pages 14051-14071

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26027 2025-10-31 cs.CV 57%

Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders

Ali Rasekh, Erfan Bagheri Soula, Omid Daliran, Simon Gottschalk, Mohsen Fayyaz

机构 * Leibniz University Hannover(莱比锡大学汉诺威分校) L3S Research Center(L3S研究中心) Microsoft(微软公司)

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23464 2025-10-29 cs.AI 57%

The Confidence Paradox: Can LLM Know When It's Wrong

Sahil Tripathi, Md Tabrez Nafis, Imran Hussain, Jiechao Gao

机构 * Jamia Hamdard(贾迈亚哈姆达德大学) Center for SDGC, Stanford University(SDGC中心,斯坦福大学)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI

Comments Accepted at the 14th IJCNLP & 4th AACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22694 2025-10-28 cs.CV cs.CL cs.IR 57%

Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation

Shu Zhao, Tianyi Shen, Nilesh Ahuja, Omesh Tickoo, Vijaykrishnan Narayanan

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Intel(英特尔)

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV

Comments Accepted at NeurIPS 2025 UniReps Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11520 2025-10-24 cs.CV 57%

mmWalk: Towards Multi-modal Multi-view Walking Assistance

Kedi Ying, Ruiping Liu, Chongyan Chen, Mingzhe Tao, Hao Shi, Kailun Yang, Jiaming Zhang, Rainer Stiefelhagen

机构 * CV:HCI, KIT(KIT计算机视觉与人机交互中心) Hunan University(湖南大学) ETH Zurich(苏黎世联邦理工学院) University of Texas at Austin(德克萨斯大学奥斯汀分校) Zhejiang University(浙江大学)

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025 Datasets and Benchmarks Track. Data and Code: https://github.com/KediYing/mmWalk

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18411 2025-10-22 cs.CL cs.LG 57%

DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding

Yue Jiang, Jichu Li, Yang Liu, Dingkang Yang, Feng Zhou, Quyu Kong

机构 * Fudan University(复旦大学) Center for Applied Statistics and School of Statistics, Renmin University of China(应用统计中心和中国人民大学统计学院) Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(北京未来区块链与隐私计算高级创新中心)

专题命中 视觉问答 :visual reasoning(abstract);分类 cs.LG

Comments Accepted by Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12801 2025-10-15 cs.CV cs.IR 57%

DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search

Kartik Narayan, Yang Xu, Tian Cao, Kavya Nerella, Vishal M. Patel, Navid Shiee, Peter Grasch, Chao Jia, Yinfei Yang, Zhe Gan

机构 * Johns Hopkins University(约翰霍普金斯大学) Apple(苹果公司)

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11907 2025-10-15 cs.CV 57%

Task-Specific Dual-Model Framework for Comprehensive Traffic Safety Video Description and Analysis

Blessing Agyei Kyem, Neema Jakisa Owor, Andrews Danyo, Joshua Kofi Asamoah, Eugene Denteh, Tanner Muturi, Anthony Dontoh, Yaw Adu-Gyamfi, Armstrong Aboah

机构 * North Dakota State University(北达科塔州立大学) University of Missouri–Columbia(密苏里大学哥伦比亚分校) University of Memphis(孟菲斯大学)

专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV

Comments This paper was accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01674 2025-10-10 cs.CV 57%

MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs

Yipeng Du, Tiehan Fan, Kepan Nan, Rui Xie, Penghao Zhou, Xiang Li, Jian Yang, Zhenheng Yang, Ying Tai

机构 * Nanjing University(南京大学) ByteDance(字节跳动) Nankai University(南开大学)

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06240 2025-10-09 cs.CL cs.AI cs.DB 57%

Knowledge Graph-Guided Multi-Agent Distillation for Reliable Industrial Question Answering with Datasets

Jiqun Pan, Zhenke Duan, Jiani Tu, Anzhi Cheng, Yanqing Wang

机构 * Zhongnan University of Economics and Law(中南财经政法大学)

专题命中 视觉问答 :grounding(abstract);分类 cs.AI

Comments 41 pages, 12 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04022 2025-10-09 cs.CV 57%

Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning

Chendong Wang, Donglin Bai, Yifan Yang, Xiao Jin, Anlan Zhang, Rui Wang, Shiqi Jiang, Yuqing Yang, Hao Wu, Qi Dai, Chong Luo, Ting Cao, Lili Qiu, Suman Banerjee

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Microsoft Research(微软研究院) Columbia University(哥伦比亚大学) University of Southern California(南加州大学) Fudan University(复旦大学)

专题命中 视觉问答 :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11334 2025-10-07 cs.AI 57%

Program Synthesis Benchmark for Visual Programming in XLogoOnline Environment

Chao Wen, Jacqueline Staub, Adish Singla

专题命中 视觉问答 :visual question answering(abstract);分类 cs.AI

Comments ACL'25 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02543 2025-10-06 cs.CV 57%

Exploring OCR-augmented Generation for Bilingual VQA

JoonHo Lee, Sunho Park

机构 * KL-Net, South Korea(韩国KL-Net)

专题命中 视觉问答 :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24445 2025-09-30 cs.CV cs.CL 57%

Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA

Jianxin Liang, Tan Yue, Yuxuan Wang, Yueqian Wang, Zhihan Yin, Huishuai Zhang, Dongyan Zhao

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所)

专题命中 视觉问答 :grounding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏