arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 6732 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1368 篇

2605.12571 2026-05-14 cs.CV cs.AI 86%

VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority

VideoSEAL: 通过解耦答案权威性来缓解代理长视频理解中的证据错位

Chenhao Qiu, Yechao Zhang, Xin Luo, Shien Song, Xusheng Liu

机构 * Nanyang Technological University, Singapore(南洋理工大学,新加坡)

专题命中 视频理解 :long video(title,abstract);video understanding(title);分类 cs.CV

AI总结 本文提出VideoSEAL框架,通过解耦规划与答案权威性,解决长视频理解中证据错位问题,提升回答准确性和证据对齐度,实验结果显示在多个基准上表现优异。

Comments Accepted to ICML 2026. 33 pages, 13 figures. Code and models are available at https://github.com/Echochef/VideoSEAL

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05079 2026-04-08 cs.CV 86%

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration

SVAgent:通过跨模态多智能体协作实现故事线引导的长视频理解

Zhongyu Yang, Zuhao Yang, Shuo Zhan, Tan Yue, Wei Pang, Yingfang Yuan

机构 * BCML, Heriot-Watt University(赫瑞瓦特大学BCML) Nanyang Technological University(南洋理工大学) Peking University(北京大学)

专题命中 视频理解 :video understanding(title,abstract);long video(title);分类 cs.CV

AI总结 SVAgent通过跨模态多智能体协作,利用故事线引导实现视频问答的高效理解,提升推理鲁棒性和答案一致性。

Comments Published in CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24271 2026-01-01 cs.CV cs.AI 86%

Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation

驯服幻觉:通过反事实视频生成提升MLLMs的视频理解

Zhe Huang, Hao Wen, Aiming Hao, Bingze Song, Meiqi Wu, Jiahong Wu, Xiangxiang Chu, Sheng Lu, Haoqian Wang

机构 * Tsinghua University(清华大学) Beihang University(北航) AMAP, Alibaba Group(阿里妈妈实验室,阿里巴巴集团)

专题命中 视频理解 :video understanding(title,abstract);video generation(title);分类 cs.CV

AI总结 通过反事实视频生成和对比训练,提升MLLMs对视频的理解能力并减少幻觉。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03500 2025-12-04 cs.CV 86%

EEA: Exploration-Exploitation Agent for Long Video Understanding

EEA:长视频理解的探索-利用代理

Te Yang, Xiangyu Zhu, Bo Wang, Quan Chen, Peng Jiang, Zhen Lei

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Kuaishou Technology(快手科技) Centre for Artificial Intelligence and Robotics, HKISI, Chinese Academy of Sciences(人工智能与机器人中心,HKISI,中国科学院)

专题命中 视频理解 :video understanding(title,abstract);long video(title);分类 cs.CV

AI总结 EEA通过语义引导的分层树搜索实现长视频理解中的探索与利用平衡,结合视觉语言模型和语义先验提升视频分析的效率与精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20362 2025-06-30 cs.CV 86%

Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding

Joao Pereira, Vasco Lopes, David Semedo, Joao Neves

机构 * DeepNeuronic NOVA LINCS, University of Beira Interior(NOVA LINCS,贝拉内里大学)

专题命中 视频理解 :video understanding(title,abstract);long video(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21184 2025-06-27 cs.CV cs.AI 86%

Task-Aware KV Compression For Cost-Effective Long Video Understanding

Minghao Qin, Yan Shu, Peitian Zhang, Kun Lun, Huaying Yuan, Juenjie Zhou, Shitao Xiao, Bo Zhao, Zheng Liu

机构 * Beijing Academy of Artificial Intelligence(北京人工智能研究院) Shanghai Jiao Tong University(上海交通大学) University of Trento(特伦特大学) Renmin University of China(中国人民大学) Beijing University of Posts and Telecommunications(北京邮电大学) Hong Kong Polytechnic University(香港理工大学) Institute of Automation CAS Beijing China(中国科学院自动化研究所)

专题命中 视频理解 :video understanding(title,abstract);long video(title);分类 cs.CV

Comments 14 pages, 3 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20384 2025-04-30 cs.CV 86%

FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding

Yanan Guo, Wenhui Dong, Jun Song, Shiding Zhu, Xuan Zhang, Hanqing Yang, Yingbo Wang, Yang Du, Xianing Chen, Bo Zheng

机构 * University of Science and Technology of China(中国科学技术大学) Alibaba Group(阿里巴巴集团) ZheJiang University(浙江大学)

专题命中 视频理解 :video understanding(title,abstract);long video(title);分类 cs.CV

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14622 2024-12-23 cs.CV 86%

Language Repository for Long Video Understanding

Kumara Kahatapitiya, Kanchana Ranasinghe, Jongwoo Park, Michael S. Ryoo

专题命中 视频理解 :video understanding(title,abstract);long video(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05352 2024-06-11 cs.CV 86%

1st Place Winner of the 2024 Pixel-level Video Understanding in the Wild (CVPR'24 PVUW) Challenge in Video Panoptic Segmentation and Best Long Video Consistency of Video Semantic Segmentation

Qingfeng Liu, Mostafa El-Khamy, Kee-Bong Song

专题命中 视频理解 :video understanding(title,abstract);long video(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13891 2025-10-17 cs.LG cs.AI 86%

K-frames: Scene-Driven Any-k Keyframe Selection for long video understanding

Yifeng Yao, Yike Yun, Jing Wang, Huishuai Zhang, Dongyan Zhao, Ke Tian, Zhihao Wang, Minghui Qiu, Tao Wang

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) Bytedance(字节跳动)

专题命中 视频理解 :video understanding(title,abstract);long video(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24869 2026-04-16 cs.CV 85%

SiLVR: A Simple Language-based Video Reasoning Framework

SiLVR:一种简单的基于语言的视频推理框架

Ce Zhang, Yan-Bo Lin, Ziyang Wang, Mohit Bansal, Gedas Bertasius

机构 * Department of Computer Science(计算机科学系) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 视频理解 :video reasoning(title,abstract);video understanding(abstract);video-language(abstract);分类 cs.CV

AI总结 SiLVR通过将复杂视频理解分解为两个阶段,利用多感官输入生成语言表示,并通过强大推理LLM解决复杂视频-语言任务,实现了在多个视频评估任务上的最佳表现。

Comments Accepted by TMLR (01/2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16210 2026-02-24 cs.CV cs.AI 85%

PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation

PyraTok: 用于视频理解和生成的语言对齐金字塔分词器

Onkar Susladkar, Tushar Prakash, Adheesh Juvekar, Kiet A. Nguyen, Dong-Hwan Jang, Inderjit S Dhillon, Ismini Lourentzou

专题命中 视频理解 :video understanding(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

AI总结 PyraTok通过多尺度文本引导量化和全局自回归目标,实现了视频理解和生成中语言与视觉的紧密对齐,提升了视频重建和零样本迁移性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17873 2026-02-04 cs.CV cs.AI 85%

SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model

SurgVidLM:迈向多粒度外科视频理解的大型语言模型

Guankun Wang, Junyi Wang, Wenjin Mo, Long Bai, Kun Yuan, Ming Hu, Jinlin Wu, Junjun He, Yiming Huang, Nicolas Padoy, Zhen Lei, Hongbin Liu, Nassir Navab, Hongliang Ren

机构 * The Chinese University of Hong Kong(香港中文大学) Sun Yat-sen University(中山大学) University of Strasbourg(斯特拉斯堡大学) Technical University of Munich(慕尼黑技术大学) Monash University(墨尔本大学) Centre for Artificial Intelligence and Robotics, HKISI-CAS(人工智能与机器人中心,HKISI-CAS) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视频理解 :video understanding(title,abstract);video language model(abstract);video reasoning(abstract);分类 cs.CV

AI总结 SurgVidLM通过多粒度分析提升外科视频理解能力,结合全局与局部机制实现更精确的手术流程解析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06509 2025-10-13 cs.CV 85%

From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding

Shih-Yao Lin, Sibendu Paul, Caren Chen

机构 * Amazon Prime Video(亚马逊Prime视频)

专题命中 视频理解 :video-language(title,abstract);video understanding(abstract);long video(abstract);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06973 2025-10-09 cs.CV 85%

Addressing the ID-Matching Challenge in Long Video Captioning

Zhantao Yang, Huangji Wang, Ruili Feng, Han Zhang, Yuting Hu, Shangwen Zhu, Junyan Li, Yu Liu, Fan Cheng

机构 * Shanghai Jiao Tong University(上海交通大学) Alibaba group(阿里巴巴集团)

专题命中 视频理解 :long video(title,abstract);video generation(abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00624 2025-08-12 cs.CV 85%

VideoSAVi: Self-Aligned Video Language Models without Human Supervision

Yogesh Kulkarni, Pooyan Fazli

机构 * Arizona State University(亚利桑那州立大学)

专题命中 视频理解 :video language model(title,abstract);video understanding(abstract);video-language(abstract);分类 cs.CV

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13250 2024-05-20 cs.CV 85%

Video ReCap: Recursive Captioning of Hour-Long Videos

Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan, Lorenzo Torresani, Gedas Bertasius

专题命中 视频理解 :long video(title,abstract);video understanding(abstract);video-language(abstract);分类 cs.CV

Comments Accepted by CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.00436 2024-03-04 cs.CV cs.AI 85%

Abductive Ego-View Accident Video Understanding for Safe Driving Perception

Jianwu Fang, Lei-lei Li, Junfei Zhou, Junbin Xiao, Hongkai Yu, Chen Lv, Jianru Xue, Tat-Seng Chua

专题命中 视频理解 :video understanding(title,abstract);video generation(abstract);video diffusion(abstract);分类 cs.CV

Comments Accepted by CVPR2024. This is not the camera-ready version. The Project page: http://www.lotvsmmau.net

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.12681 2022-04-19 cs.CV 85%

VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin, William Yang Wang, Lijuan Wang, Zicheng Liu

专题命中 视频理解 :video-language(title,abstract);video understanding(abstract);text-to-video(abstract);分类 cs.CV

Comments Code is available at https://github.com/tsujuifu/pytorch_violet

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09146 2025-09-03 cs.CV cs.MM 85%

Generative Frame Sampler for Long Video Understanding

Linli Yao, Haoning Wu, Kun Ouyang, Yuanxing Zhang, Caiming Xiong, Bei Chen, Xu Sun, Junnan Li

机构 * National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(国家多媒体信息处理重点实验室,计算机学院,北京大学) Peking University(北京大学) Salesforce Research(Salesforce 研究) Independent Researcher(独立研究者)

专题命中 视频理解 :video understanding(title);long video(title);分类 cs.CV、cs.MM

Comments ACL 2025 Findings. Code: https://github.com/yaolinli/GenS

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02438 2025-09-11 cs.CL cs.AI 85%

Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Chuanqi Cheng, Jian Guan, Wei Wu, Rui Yan

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 视频理解 :video-language(title,abstract);video understanding(abstract);long video(abstract)

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12833 2025-05-29 cs.CV 84%

FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering

Zheng Cheng, Rendong Wang, Zhicheng Wang

专题命中 视频理解 :video understanding(title);long video(title);分类 cs.CV

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.09996 2021-10-04 cs.CV cs.CL 84%

VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Hu Xu, Gargi Ghosh, Po-Yao Huang, Prahal Arora, Masoumeh Aminzadeh, Christoph Feichtenhofer, Florian Metze, Luke Zettlemoyer

专题命中 视频理解 :video understanding(title);video-language(title);分类 cs.CV

Comments 9 pages, ACL Findings 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20579 2026-03-20 cs.CV cs.AI cs.MM 84%

Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence

Open-o3-Video: 基于显式时空证据的视频推理

Jiahao Meng, Xiangtai Li, Haochen Wang, Yue Tan, Tao Zhang, Lingdong Kong, Yunhai Tong, Anran Wang, Zhiyang Teng, Yujing Wang, Zhuochen Wang

机构 * PKU(北京大学) ByteDance(字节跳动) CASIA(CASIA研究院) WHU(武汉大学) NUS(新加坡国立大学)

专题命中 视频理解 :video reasoning(title,abstract);video understanding(abstract);分类 cs.CV、cs.MM

AI总结 本文提出Open-o3-Video,通过显式时空证据提升视频推理的可追溯性与可验证性,基于STGR数据集和冷启动强化学习策略,在V-STAR基准上取得SOTA性能,并支持置信度感知的测试时扩展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00126 2026-03-03 cs.CV cs.AI cs.IR cs.MM cs.PF cs.SY eess.SY 84%

QuickGrasp: Responsive Video-Language Querying Service via Accelerated Tokenization and Edge-Augmented Inference

QuickGrasp: 通过加速分词和边缘增强推理实现响应式视频-语言查询服务

Miao Zhang, Ruixiao Zhang, Jianxin Shi, Hengzhi Wang, Hao Fang, Jiangchuan Liu

机构 * School of Computing Science, Simon Fraser University(计算科学学院,西蒙弗雷泽大学) Computer Science Department, University of Illinois Urbana-Champaign(计算机科学系,伊利诺伊大学厄巴纳-香槟分校) College of Cryptology and Cyber Science, Nankai University(密码学与网络科学学院,南开大学) College of Computer Science and Software Engineering, Shenzhen University(计算机科学与软件工程学院,深圳大学)

专题命中 视频理解 :video-language(title,abstract);video understanding(abstract);分类 cs.CV、cs.MM

AI总结 QuickGrasp通过加速分词和边缘增强推理,实现响应式视频-语言查询服务,兼顾准确性和响应速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24097 2026-01-01 cs.CV cs.AI cs.CL cs.MM 84%

Factorized Learning for Temporally Grounded Video-Language Models

分解学习用于时间感知的视频-语言模型

Wenzheng Zeng, Difei Gao, Mike Zheng Shou, Hwee Tou Ng

机构 * National University of Singapore(新加坡国立大学)

专题命中 视频理解 :video-language(title,abstract);video understanding(abstract);分类 cs.CV、cs.MM

AI总结 本文提出D$^2$VLM框架,通过分解学习方法提升视频-语言模型在时间定位和文本响应任务中的性能,引入证据标记和FPO算法以优化学习过程。

Comments ICCV 2025 paper. This arXiv version updates Figure 1 to include the concurrent work Qwen2.5-VL to ensure consistency with Table 1

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09068 2025-07-24 cs.CV cs.AI cs.IR cs.LG cs.MM 84%

Infinite Video Understanding

Dell Zhang, Xiangyu Chen, Jixiang Luo, Mengxi Jia, Changzhi Sun, Ruilong Ren, Jingren Liu, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) Peking University(北京大学) Tianjin University(天津大学)

专题命中 视频理解 :video understanding(title,abstract);long video(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17729 2024-05-29 cs.CV cs.MM 84%

Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions

Rui Zhang, Shuailong Li, Junxiao Xue, Feng Lin, Qing Zhang, Xiao Ma, Xiaoran Yan

专题命中 视频理解 :video-language(title,abstract);video understanding(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13054 2025-11-18 cs.CV 84%

ViSS-R1: Self-Supervised Reinforcement Video Reasoning

Bo Fang, Yuxin Song, Qiangqiang Wu, Haoyuan Sun, Wenhao Wu, Antoni B. Chan

机构 * City University of Hong Kong(香港城市大学) Baidu Inc.(百度公司) Tsinghua University(清华大学) The University of Sydney(悉尼大学)

专题命中 视频理解 :video reasoning(title,abstract);video understanding(abstract);分类 cs.CV

Comments Our paper was initially titled "Video-SSR1: Self-Supervised Reinforcement Video Reasoning." Upon noticing its close resemblance to the title of a recently released paper, we have decided to rename our work as "ViSS-R1."

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02259 2025-08-26 cs.CV 84%

T*: Re-thinking Temporal Search for Long-Form Video Understanding

Jinhui Ye, Zihan Wang, Haosen Sun, Keshigeyan Chandrasegaran, Zane Durante, Cristobal Eyzaguirre, Yonatan Bisk, Juan Carlos Niebles, Ehsan Adeli, Li Fei-Fei, Jiajun Wu, Manling Li

专题命中 视频理解 :video understanding(title,abstract);long video(abstract,comments);分类 cs.CV

Comments Accepted by CVPR 2025; A real-world long video needle-in-haystack benchmark; long-video QA with human ref frames

详情

展开后加载摘要…

URL PDF HTML 收藏