arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 1369 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1369 篇

2103.15662 2021-03-30 cs.CV 79%

Unified Graph Structured Models for Video Understanding

Anurag Arnab, Chen Sun, Cordelia Schmid

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.11451 2020-12-16 cs.CV 79%

Large Scale Holistic Video Understanding

Ali Diba, Mohsen Fayyaz, Vivek Sharma, Manohar Paluri, Jurgen Gall, Rainer Stiefelhagen, Luc Van Gool

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments ECCV 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.00375 2020-11-10 cs.CV 79%

Towards Visually Explaining Video Understanding Networks with Perturbation

Zhenqiang Li, Weimin Wang, Zuoyue Li, Yifei Huang, Yoichi Sato

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments Accepted by WACV2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.14104 2020-10-28 cs.CV cs.AI cs.CL 79%

Co-attentional Transformers for Story-Based Video Understanding

Björn Bebensee, Byoung-Tak Zhang

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments 10 pages, 2 figures, submitted to ICASSP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.07385 2020-10-01 cs.CV 79%

Representation Learning on Visual-Symbolic Graphs for Video Understanding

Effrosyni Mavroudi, Benjamín Béjar Haro, René Vidal

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments ECCV 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.07203 2020-09-21 cs.CV 79%

Video Understanding as Machine Translation

Bruno Korbar, Fabio Petroni, Rohit Girdhar, Lorenzo Torresani

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments The authors have temporarily withdrawn this paper to reassess some of the experimental results

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.04465 2020-09-09 cs.CV 79%

SoccerDB: A Large-Scale Database for Comprehensive Video Understanding

Yudong Jiang, Kaixu Cui, Leilei Chen, Canjin Wang, Changliang Xu

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments accepted by MM2020 sports workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.09933 2020-07-21 cs.CV 79%

MotionSqueeze: Neural Motion Feature Learning for Video Understanding

Heeseung Kwon, Manjin Kim, Suha Kwak, Minsu Cho

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments Accepted to ECCV 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.08383 2019-08-23 cs.CV 79%

TSM: Temporal Shift Module for Efficient Video Understanding

Ji Lin, Chuang Gan, Song Han

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.10654 2019-05-28 cs.CV 79%

Exploring Temporal Information for Improved Video Understanding

Yi Zhu

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments PhD dissertation. Posting here for easier access and update. Chapter 4, 6 and 7 are collaborative work

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.06739 2018-11-12 cs.CV 79%

Constrained-size Tensorflow Models for YouTube-8M Video Understanding Challenge

Tianqi Liu, Bo Liu

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments Accepted paper at 2018 ECCV Youtube8M workshop: https://research.google.com/youtube8m/workshop2018/

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.03316 2018-09-11 cs.CV cs.LG stat.ML 79%

Hierarchical Video Understanding

Farzaneh Mahdisoltani, Roland Memisevic, David Fleet

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.06330 2018-03-22 cs.CV 79%

Attend and Interact: Higher-Order Object Interactions for Video Understanding

Chih-Yao Ma, Asim Kadav, Iain Melvin, Zsolt Kira, Ghassan AlRegib, Hans Peter Graf

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments CVPR 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.04555 2017-07-17 cs.CV 79%

Temporal Modeling Approaches for Large-scale Youtube-8M Video Understanding

Fu Li, Chuang Gan, Xiao Liu, Yunlong Bian, Xiang Long, Yandong Li, Zhichao Li, Jie Zhou, Shilei Wen

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments To appear on CVPR 2017 YouTube-8M Workshop(Rank 3rd out of 650 teams)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07959 2026-08-11 cs.AI 新提交 78%

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning

SCOUT:面向超长时长自我中心视频推理的自检查与恢复感知工具思维智能体

Keyang Zhong, Kuo Wang, Peng Liu, Quanlong Zheng, Junlin Xie, Zhijia Liang, Yanhao Zhang, Guanbin Li

机构 * Sun Yat-sen University(中山大学) Shenzhen Loop Area Institute(深圳河套学院) Guangdong OPPO Mobile Telecommunications Corp., Ltd.(广东欧珀移动通信有限公司) OPPO AI Center(OPPO人工智能中心) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 视频理解 :video reasoning(title);video understanding(abstract)

AI总结 本研究提出SCOUT智能体框架,结合自适应策略与UPS-GRPO方法,解决超长自我中心视频推理的错误传播及信用分配问题,在对应基准上达最优性能且短时长场景仍具竞争力。

Comments Accepted by ACM Multimedia 2026 (MM '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08342 2026-08-03 cs.LG 版本更新 78%

EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment

EgoEverything:一个受人类行为启发的长上下文第一人称视频理解AR环境基准

Qiance Tang, Ziqi Wang, Jieyu Lin, Ziyun Li, Barbara De Salvo, Sai Qian Zhang

机构 * New York University(纽约大学) Meta Reality Labs

专题命中 视频理解 :video understanding(title,abstract)

AI总结 本文提出EgoEverything基准,通过利用人类注意力信号生成问题,更真实地捕捉自然人类行为,用于评估AR环境中长上下文第一人称视频理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20937 2026-06-23 cs.LG 版本更新 78%

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs

面向高效视频大语言模型的sink-token感知剪枝:用于细粒度视频理解

Kibum Kim, Jiwan Kim, Kyle Min, Yueqi Wang, Jinyoung Moon, Julian McAuley, Chanyoung Park

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国高级科学技术研究院) Oracle University of California, San Diego(加州大学圣地亚哥分校) Electronics and Telecommunications Research Institute (ETRI)(电子电信研究院)

专题命中 视频理解 :video understanding(title,abstract)

AI总结 本文提出Sink-Token-aware Pruning方法,通过识别并抑制semantically uninformative tokens,提升细粒度视频理解性能,在多种基准测试中表现优异。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04588 2026-06-04 cs.CL 78%

VCIFBench: Evaluating Complex Instruction Following for Video Understanding

VCIFBench:评估视频理解中的复杂指令遵循能力

Huangchen Xu, Yuan Wu, Yi Chang

机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, Jilin University(知识驱动人机智能工程研究中心,吉林大学) International Center of Future Science, Jilin University(未来科学国际中心,吉林大学)

专题命中 视频理解 :video understanding(title,abstract)

AI总结 提出VCIFBench基准,通过混合验证流水线评估多模态大模型在视频理解中遵循内容、格式、风格和结构约束的复杂指令能力,实验表明联合约束满足仍具挑战,DPO训练可提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21444 2026-04-24 cs.AI 78%

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration

HiCrew:通过问题感知多智能体协作实现长视频理解的层次推理

Yuehan Zhu, Jingqi Zhao, Jiawen Zhao, Xudong Mao, Baoquan Zhao

机构 * School of Artificial Intelligence, Sun Yat-sen University, China(中山大学人工智能学院)

专题命中 视频理解 :video understanding(title,abstract)

AI总结 本文提出HiCrew框架,通过混合树结构、问题感知描述生成和动态规划层,解决长视频中时空冗余和叙事依赖问题,提升时间与因果推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23990 2025-11-13 cs.AI 78%

Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding

Mingyang Mao, Mariela M. Perez-Cabarcas, Utteja Kallakuri, Nicholas R. Waytowich, Xiaomin Lin, Tinoosh Mohsenin

机构 * Johns Hopkins Whiting School of Engineering(约翰霍普金斯大学惠廷工程学院) DEVCOM Army Research Laboratory(国防部陆军研究实验室)

专题命中 视频理解 :video understanding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16810 2025-09-23 cs.AI 78%

Automated Procedural Analysis via Video-Language Models for AI-assisted Nursing Skills Assessment

Shen Chang, Dennis Liu, Renran Tian, Kristen L. Swartzell, Stacie L. Klingler, Amy M. Nagle, Nan Kong

机构 * Weldon School of Biomedical Engineering, Purdue University(普渡大学生物医学工程学院) Department of Industrial and Operations Engineering, University of Michigan(密歇根大学工业与运作工程系) Edward P. Fitts Department of Industrial and Systems Engineering, North Carolina State University(北卡罗来纳州立大学工业与系统工程系) School of Nursing, Purdue University(普渡大学护理学院)

专题命中 视频理解 :video-language(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14395 2025-08-21 cs.HC cs.AI 78%

NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding

Running Zhao, Zhihan Jiang, Xinchen Zhang, Chirui Chang, Handi Chen, Weipeng Deng, Luyao Jin, Xiaojuan Qi, Xun Qian, Edith C. H. Ngai

机构 * The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学) Google(谷歌)

专题命中 视频理解 :video understanding(title);long video(abstract)

Comments Accepted to UIST 2025. Project website: https://zhaorunning.github.io/NoteIt/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02001 2025-07-04 cs.LG 78%

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

Anurag Arnab, Ahmet Iscen, Mathilde Caron, Alireza Fathi, Cordelia Schmid

机构 * Google DeepMind(谷歌DeepMind)

专题命中 视频理解 :video understanding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23870 2025-05-14 cs.LO 78%

A SAT-centered XAI method for Deep Learning based Video Understanding

Hojer Key

专题命中 视频理解 :video understanding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.06807 2018-10-17 cs.LG cs.AR cs.NE stat.ML 78%

Morph: Flexible Acceleration for 3D CNN-based Video Understanding

Kartik Hegde, Rohit Agrawal, Yulun Yao, Christopher W. Fletcher

专题命中 视频理解 :video understanding(title,abstract)

Comments Appears in the proceedings of the 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.08222 2017-06-27 stat.ML 78%

YouTube-8M Video Understanding Challenge Approach and Applications

Edward Chen

专题命中 视频理解 :video understanding(title,abstract)

Comments YouTube-8M Workshop submission, 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02327 2026-05-04 cs.CV 77%

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

PPLLaVA:通过提示引导的视频序列理解

Shangkun Sun, Ruyang Liu, Haoran Tang, Yixiao Ge, Haibo Lu, Wei Gao, Jiankun Yang, Chen Li

机构 * Peng Cheng Laboratory(鹏城实验室) Peking University(北京大学) XPeng Inc.(小鹏汽车有限公司) Ministry of Industry and Information Technology of the People’s Republic of China, China Electronics Standardization Institute, Beijing, China(中华人民共和国工业和信息化部,中国电子标准化研究院,北京,中国)

专题命中 视频理解 :video understanding(abstract);video reasoning(abstract);long video(abstract);分类 cs.CV

AI总结 PPLLaVA通过提示引导的池化策略实现视频序列高效理解,减少18倍token,提升推理效率并在多种视频理解任务中取得最佳性能。

Comments Accepted to ICLR' 26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04202 2026-02-05 cs.CV 77%

VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents

VTok: 一种具有解耦空间-时间潜在变量的统一视频分词器

Feng Wang, Yichun Shi, Ceyuan Yang, Qiushan Guo, Jingxiang Sun, Alan Yuille, Peng Wang

机构 * Johns Hopkins University(约翰斯·霍普金斯大学)

专题命中 视频理解 :video generation(abstract);video understanding(abstract);text-to-video(abstract);分类 cs.CV

AI总结 VTok通过解耦空间和时间潜在变量,实现紧凑且高效的视频分词,提升视频理解和生成任务的性能与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20470 2025-11-21 cs.CV 77%

Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence

Conan:基于多尺度视觉证据的逐步学习以像侦探一样推理

Kun Ouyang, Yuanxin Liu, Linli Yao, Yishuo Cai, Hao Zhou, Jie Zhou, Fandong Meng, Xu Sun

机构 * State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) WeChat AI, Tencent Inc., China(微信AI,腾讯公司,中国)

专题命中 视频理解 :video understanding(abstract);video reasoning(abstract);long video(abstract);分类 cs.CV

AI总结 Conan通过多阶段渐进冷启动策略和AIR RLVR框架,实现证据基础的多步视频推理,超越基线模型,达到最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04369 2025-11-14 cs.CV 77%

TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding

Canhui Tang, Zifan Han, Hongbo Sun, Sanping Zhou, Xuchong Zhang, Xin Wei, Ye Yuan, Huayu Zhang, Jinglin Xu, Hao Sun

专题命中 视频理解 :video understanding(abstract);video-language(abstract);long video(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏