arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 6732 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1368 篇

2604.14149 2026-04-17 cs.CV 88%

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding

每个高度选择性帧一个标记:朝着长视频理解的极端压缩

Zheyu Zhang, Ziqi Pang, Shixing Chen, Xiang Hao, Vimal Bhat, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon Prime Video(亚马逊Prime视频)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

AI总结 本文提出XComp模型,通过极端视频标记压缩提升长视频理解性能,结合标记级和帧级压缩,实现更高的压缩比和更密集的帧采样。

Comments Appear in the proceedings of NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20913 2026-04-16 cs.CV 88%

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding

LongVideo-R1: 低成本长视频理解的智能导航

Jihao Qiu, Lingxi Xie, Xinyue Huo, Qi Tian, Qixiang Ye

机构 * University of Chinese Academy of Sciences(中国科学院大学) Huawei Consumer Business Group(华为消费者业务集团)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

AI总结 本文提出LongVideo-R1,一种基于多模态大语言模型的智能导航系统,通过高效视频上下文导航减少冗余搜索,提升长视频理解的效率与准确性。

Comments 17 pages, 9 figures, 8 tables, accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08120 2026-04-10 cs.CV cs.AI cs.CL cs.LG 88%

Small Vision-Language Models are Smart Compressors for Long Video Understanding

小视觉-语言模型是长视频理解的智能压缩器

Junjie Fei, Jun Chen, Zechun Liu, Yunyang Xiong, Chong Zhou, Wei Wen, Junlin Han, Mingchen Zhuge, Saksham Suri, Qi Qian, Shuming Liu, Lemeng Wu, Raghuraman Krishnamoorthi, Vikas Chandra, Mohamed Elhoseiny, Chenchen Zhu

机构 * Meta AI King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

AI总结 本文提出Tempo框架,利用小视觉-语言模型压缩长视频,通过自适应令牌分配实现高效压缩,实验显示其在极端长视频任务中表现优异。

Comments Project page and demo are available at https://FeiElysia.github.io/tempo-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07071 2026-03-11 cs.CV 88%

VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding

VirtueBench: 在长视频理解中评估在不确定性下的可信度

Xueqing Yu, Bohan Li, Yan Li, Zhenheng Yang

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

AI总结 VirtueBench通过评估模型在不确定性下的可信度,揭示了长视频理解中模型拒绝行为的差异及可靠性问题。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00141 2026-03-06 cs.CV cs.AI 88%

FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding

FLoC:基于设施定位的高效视觉令牌压缩用于长视频理解

Janghoon Cho, Jungsoo Lee, Munawar Hayat, Kyuwoong Hwang, Fatih Porikli, Sungha Choi

机构 * Qualcomm AI Research(高通人工智能研究)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

AI总结 FLoC提出了一种基于设施定位函数的高效视觉令牌压缩方法,通过懒惰贪心算法实现紧凑且多样化的令牌选择,有效提升长视频理解的处理效率和性能。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13564 2026-02-10 cs.CV 88%

State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal Models

具有门控注意力和可学习采样的状态空间分层压缩用于大多模态模型中的小时级视频理解

Geewook Kim, Minjoon Seo

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于状态空间模型和门控注意力的高效压缩方法,用于减少大模型中小时级视频的token消耗,同时保持性能。

Comments AAAI 2026 (Oral). Project page: https://github.com/naver-ai/mambamia

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10200 2025-12-23 cs.CV 88%

LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents

LVAgent: 通过多轮动态协作的MLLM代理实现长视频理解

Boyu Chen, Zhengrong Yue, Siran Chen, Zikang Wang, Yang Liu, Peng Li, Yali Wang

机构 * Shenzhen Key Lab of Computer Vision and Pattern Recognition, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳计算机视觉与模式识别重点实验室,深圳先进技术研究院,中国科学院) Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China(人工智能产业研究院(AIR),清华大学,北京,中国) Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University, Beijing, China(计算机科学与技术系,人工智能研究院,清华大学,北京,中国) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

AI总结 LVAgent通过多轮动态协作的MLLM代理提升长视频理解性能,实现80%的准确率并在LongVideoBench上提升13.3%

Comments accepted in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16595 2025-11-27 cs.CV cs.AI cs.CL 88%

TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding

TimeViper: 一种用于高效长视频理解的混合Mamba-Transformer视觉语言模型

Boshen Xu, Zihan Xiao, Jiaze Li, Jianzhong Ju, Zhenbo Luo, Jian Luan, Qin Jin

机构 * AIM3 Lab, Renmin University of China(中国人民大学人工智能实验室) MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

AI总结 TimeViper是一种混合Mamba-Transformer模型,通过TransV模块实现高效长视频理解,提升多模态处理能力。

Comments Project page: https://xuboshen.github.io/TimeViper; Code: https://github.com/xiaomi-research/timeviper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27280 2025-11-25 cs.CV cs.AI cs.LG 88%

FOCUS: Efficient Keyframe Selection for Long Video Understanding

FOCUS: 长视频理解中的高效关键帧选择

Zirui Zhu, Hailun Xu, Yang Luo, Yong Liu, Kanchan Sarkar, Zhenheng Yang, Yang You

机构 * National University of Singapore(新加坡国立大学) TikTok

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

AI总结 FOCUS通过两阶段探索-利用策略,在严格标记预算下高效选择关键帧,提升长视频理解的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13589 2025-11-25 cs.CV 88%

AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding

AdaVideoRAG:多情境自适应检索增强高效长视频理解

Zhucun Xue, Jiangning Zhang, Xurong Xie, Yuxuan Cai, Yong Liu, Xiangtai Li, Dacheng Tao

机构 * Zhejiang University(浙江大学) Youtu Lab, Tencent(腾讯优图实验室) Huazhong University of Science and Technolog(华中科技大学) Nanyang Technological University(南洋理工大学)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

AI总结 AdaVideoRAG通过自适应检索增强框架提升长视频理解效率与准确性,支持多层级知识检索与深度语义分析。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20622 2025-10-24 cs.CV 88%

SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding

Yuan Sheng, Yanbin Hao, Chenxu Li, Shuo Wang, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学) Hefei University of Technology(合肥工业大学)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21496 2025-09-03 cs.CV cs.AI 88%

ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding

Hao Lu, Jiahao Wang, Yaolun Zhang, Ruohui Wang, Xuanyu Zheng, Yepeng Tang, Dahua Lin, Lewei Lu

机构 * Sensetime(秒氏科技)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01546 2025-08-05 cs.CV 88%

E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation

Zeyu Xu, Junkang Zhang, Qiang Wang, Yi Liu

机构 * Zeyu Xu(作者) Junkang Zhang(作者) Qiang Wang(作者) Yi Liu(作者)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18478 2025-04-29 cs.CV 88%

Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Xiangrui Liu, Yan Shu, Zheng Liu, Ao Li, Yang Tian, Bo Zhao

机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) University of Trento(特伦多大学)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16082 2025-04-23 cs.CV 88%

MR. Video: "MapReduce" is the Principle for Long Video Understanding

Ziqi Pang, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20472 2025-03-27 cs.CV cs.AI 88%

From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment

Yucheng Suo, Fan Ma, Linchao Zhu, Tianyi Wang, Fengyun Rao, Yi Yang

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08576 2025-03-12 cs.CV 88%

RAG-Adapter: A Plug-and-Play RAG-enhanced Framework for Long Video Understanding

Xichen Tan, Yunfan Ye, Yuanjing Luo, Qian Wan, Fang Liu, Zhiping Cai

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

Comments 37 pages, 36 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10471 2025-03-11 cs.CV cs.AI 88%

VCA: Video Curious Agent for Long Video Understanding

Zeyuan Yang, Delin Chen, Xueyang Yu, Maohao Shen, Chuang Gan

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21271 2025-03-03 cs.CV cs.AI cs.LG 88%

Adaptive Keyframe Sampling for Long Video Understanding

Xi Tang, Jihao Qiu, Lingxi Xie, Yunjie Tian, Jianbin Jiao, Qixiang Ye

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

Comments CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04264 2025-01-03 cs.CV cs.AI cs.CL 88%

MLVU: Benchmarking Multi-task Long Video Understanding

Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Zhengyang Liang, Shitao Xiao, Minghao Qin, Xi Yang, Yongping Xiong, Bo Zhang, Tiejun Huang, Zheng Liu

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06182 2024-12-12 cs.CV 88%

Towards Long Video Understanding via Fine-detailed Video Story Generation

Zeng You, Zhiquan Wen, Yaofo Chen, Xin Li, Runhao Zeng, Yaowei Wang, Mingkui Tan

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18938 2024-12-04 cs.CV cs.AI 88%

From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Heqing Zou, Tianze Luo, Guiyang Xie, Victor, Zhang, Fengmao Lv, Guangcong Wang, Junyang Chen, Zhuochen Wang, Hansheng Zhang, Huaijian Zhang

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12846 2024-11-26 cs.CV 88%

DrVideo: Document Retrieval Based Long Video Understanding

Ziyu Ma, Chenhui Gou, Hengcan Shi, Bin Sun, Shutao Li, Hamid Rezatofighi, Jianfei Cai

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03384 2024-07-23 cs.CV 88%

LongVLM: Efficient Long Video Understanding via Large Language Models

Yuetian Weng, Mingfei Han, Haoyu He, Xiaojun Chang, Bohan Zhuang

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

Comments Accepted by ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12724 2023-10-20 cs.CV 88%

Query-aware Long Video Localization and Relation Discrimination for Deep Video Understanding

Yuanxing Xu, Yuting Wei, Bin Wu

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

Comments ACM MM 2023 Grand Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07872 2026-05-11 cs.CV cs.AI 87%

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models

视频理解奖励建模:一个稳健的基准和高效的奖励模型

Yuancheng Wei, Linli Yao, Lei Li, Haojie Zhang, Hao Zhou, Fandong Meng, Xu Sun

机构 * South China University of Technology(华南理工大学) Peking University(北京大学) The University of Hong Kong(香港大学) Tencent(腾讯)

专题命中 视频理解 :video understanding(title,summary_cn);分类 cs.CV

AI总结 本文提出Video Understanding Reward Bench基准和VideoDRM/VideoGRM模型,通过大规模高质量数据提升视频理解奖励建模性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04325 2024-06-07 cs.CV 87%

ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, Li Yuan, Yu Qiao, Dahua Lin, Feng Zhao, Jiaqi Wang

专题命中 视频理解 :video understanding(title,abstract);video generation(abstract);text-to-video(abstract);video-language(abstract)

Comments Project Page: https://sharegpt4video.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03918 2026-08-05 cs.CV cs.AI 新提交 86%

When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

何时与何地查看:用于高效长视频理解的自适应视觉证据调度

Ke Li, Jiayu Chen, Maoliang Li, Zihao Zheng, Hailong Zou, Hengyi Zhang, Xuanzhe Liu, Xiang Chen

专题命中 视频理解 :video understanding(title,abstract);long video(title);分类 cs.CV

AI总结 提出无需训练的EcoFrame框架,通过熵门控预算调度和注意力引导候选提议实现高效视觉证据调度,在多个长视频理解基准上实现精度与效率的更优权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09064 2026-06-09 cs.CV cs.AI 新提交 86%

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding

看得更多,思考更深:面向长视频理解的查询扩展视觉证据与答案线索引导反思

Shuning Wang, Zhiheng Wu, YiNuo Lu, Naiming Liu, Chen Jia, Bowen Liu, Shuo Nie, Weijie Zhu, Yumeng Zhang

机构 * Baidu Inc.(百度公司) Harbin Institute of Technology(哈尔滨工业大学) Hong Kong University of Science and Technology(香港科技大学)

专题命中 视频理解 :video understanding(title,abstract);long video(title);分类 cs.CV

AI总结 提出CoVER框架,通过动态收集查询扩展视觉证据和答案特定视觉反馈验证草稿答案,实现从答案中心生成到证据中心和视觉可验证推理的转变,在长视频理解任务上超越同规模模型及部分闭源模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31598 2026-06-01 cs.CV 86%

Linear Scaling Video VLMs for Long Video Understanding

面向长视频理解的线性缩放视频视觉语言模型

Cristobal Eyzaguirre, Jiajun Wu, Juan Carlos Niebles

机构 * Stanford University(斯坦福大学)

专题命中 视频理解 :video understanding(title,abstract);long video(title);分类 cs.CV

AI总结 提出StateKV方法,通过固定容量的重要性驱动循环状态实现线性时间视频预填充,在保持接近全自注意力性能的同时显著降低计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏