arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 6737 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1369 篇

2512.18671 2025-12-23 cs.CV 79%

SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse

SmartSight: 通过时间注意力崩溃缓解视频大语言模型中的幻觉而不影响视频理解

Yiming Sun, Mi Zhang, Feifei Li, Geng Hong, Min Yang

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 SmartSight通过时间注意力崩溃技术,在不牺牲视频理解能力的前提下,有效降低视频大语言模型的幻觉问题。

Comments AAAI26 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07576 2025-12-23 cs.CV 79%

Super Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding

超级编码网络:多模态编码器的递归关联用于视频理解

Boyu Chen, Siran Chen, Kunchang Li, Qinglin Xu, Yu Qiao, Yali Wang

机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) the School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 本文提出超级编码网络,通过递归关联多模态编码器提升视频理解性能,显著提升跟踪、识别、聊天和编辑等任务效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17108 2025-12-22 cs.LG cs.MM 79%

Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse

Atom:通过模块化重用实现高效的设备端视频-语言流水线

Kunjal Panchal, Saayan Mitra, Somdeb Sarkhel, Haoliang Wang, Ishita Dasgupta, Gang Wu, Hui Guan

专题命中 视频理解 :video-language(title,abstract);分类 cs.MM

AI总结 Atom通过模块化重用提升设备端视频-语言流水线效率,实现27-33%的执行速度提升,同时保持性能稳定。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02262 2025-12-19 cs.CV 79%

From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding

从帧到片段:训练免费的自适应关键片段选择用于长形式视频理解

Guangyu Sun, Archit Singhal, Burak Uzkent, Mubarak Shah, Chen Chen, Garin Kessler

机构 * Amazon(亚马逊公司) University of Central Florida(中央佛罗里达大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 F2C通过自适应关键片段选择提升长视频理解,比均匀采样在多个基准上表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09143 2025-12-17 cs.CV 79%

Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding

Exo2Ego:基于外部知识引导的多模态大语言模型用于第一人称视频理解

Haoyu Zhang, Qiaohui Chu, Meng Liu, Haoxiang Shi, Yaowei Wang, Liqiang Nie

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 Exo2Ego通过迁移学习提升内向视频理解能力,利用外向知识增强模型性能。

Comments This paper is accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12560 2025-12-16 cs.CV cs.AI 79%

StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding

StreamingAssistant: 为加速在线视频理解的高效视觉标记修剪

Xinqi Jin, Hanxun Yu, Bohan Yu, Kebin Liu, Jian Liu, Keda Tao, Yixuan Pei, Huan Wang, Fan Dang, Jiangchuan Liu, Weiqiang Wang

机构 * Tsinghua University(清华大学) Zhejiang University(浙江大学) Ant Group(蚂蚁集团) Westlake University(西湖大学) Beijing Jiaotong University(北京交通大学) Simon Fraser University(西蒙弗雷泽大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 StreamingAssistant通过高效视觉标记修剪技术,提升在线视频理解的准确性和效率,减少GPU内存使用和计算延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06335 2025-12-10 cs.CV 79%

Harnessing Object Grounding for Time-Sensitive Video Understanding

利用物体接地提升时间敏感视频理解

Tz-Ying Wu, Sharath Nittur Sridhar, Subarna Tripathi

机构 * Intel(英特尔公司)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 本文提出 GO-Tokenizer 以提升视频大型语言模型的时间敏感视频理解能力,通过实时编码紧凑的物体信息,提高模型性能并减少噪声影响。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19552 2025-12-08 cs.CV 79%

iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning

iFinder: 结构化零样本视觉基于LLM的地面定位用于行车记录仪视频推理

Manyi Yao, Bingbing Zhuang, Sparsh Garg, Amit Roy-Chowdhury, Christian Shelton, Manmohan Chandraker, Abhishek Aich

机构 * NEC Laboratories, America(NEC美国实验室) University of California, Riverside(加州大学河滨分校) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 视频理解 :video reasoning(title);video understanding(abstract);分类 cs.CV

AI总结 iFinder通过结构化语义接地框架,利用行车记录仪视频中的关键线索提升LLM在驾驶视频推理中的性能。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15470 2025-12-05 cs.CV cs.AI 79%

EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining

EgoDTM: 向3D感知的自体视频-语言预训练迈进

Boshen Xu, Yuting Mei, Xinbi Liu, Sipeng Zheng, Qin Jin

机构 * AIM3 Lab, Renmin University of China(中国人民大学人工智能3实验室)

专题命中 视频理解 :video-language(title,abstract);分类 cs.CV

AI总结 EgoDTM通过结合大规模3D感知视频预训练和视频-文本对比学习,提升视频-语言模型的3D感知能力,实现更丰富的空间理解。

Comments Code: https://github.com/xuboshen/EgoDTM

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04025 2025-12-04 cs.CV cs.AI cs.LG 79%

PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation

PSA: 基于金字塔稀疏注意力的高效视频理解和生成

Xiaolong Li, Youping Gu, Xi Lin, Weijie Wang, Bohan Zhuang

机构 * ZIP Lab, Zhejiang University(浙江大学浙大信息实验室)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 PSA通过多级池化键值表示实现高效视频理解和生成,相比现有稀疏注意力方法在效率和质量上表现更优。

Comments Tech report

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18773 2025-12-01 cs.CV 79%

Spacewalk-18: A Benchmark for Multimodal and Long-form Procedural Video Understanding in Novel Domains

Spacewalk-18:多模态和长形式程序视频理解的基准测试,用于新领域

Zitian Tang, Rohan Myer Krishnan, Zhiqiu Yu, Chen Sun

机构 * Brown University(布朗大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 Spacewalk-18是一个用于多模态和长形式程序视频理解的新领域基准测试,通过步骤识别和视频问答任务评估模型在新领域和长时序上下文中的泛化能力。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12336 2025-11-27 cs.CV 79%

Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding

对多模态大语言模型在视频理解中的可信度进行基准测试

Youze Wang, Zijun Chen, Ruoyu Chen, Shishen Gu, Wenbo Hu, Jiayang Liu, Yinpeng Dong, Hang Su, Jun Zhu, Meng Wang, Richang Hong

机构 * Hefei University of Technology(合肥工业大学) Tsinghua University(清华大学) Institute of Science Tokyo(东京科学研究所)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 本研究提出Trust-videoLLMs基准,评估23种视频LLMs在真实性、鲁棒性、安全性和隐私等方面的表现,揭示其在动态场景理解及现实风险缓解中的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17432 2025-11-26 cs.CV cs.CL 79%

Video Understanding with Large Language Models: A Survey

利用大语言模型进行视频理解:综述

Yolo Y. Tang, Jing Bi, Siting Xu, Luchuan Song, Susan Liang, Teng Wang, Daoan Zhang, Jie An, Jingyang Lin, Rongyi Zhu, Ali Vosoughi, Chao Huang, Zeliang Zhang, Pinxin Liu, Mingqian Feng, Feng Zheng, Jianguo Zhang, Ping Luo, Jiebo Luo, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) The University of Hong Kong(香港大学) Southern University of Science and Technology(南方科技大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 本文综述了利用大语言模型进行视频理解的最新进展,探讨了Vid-LLMs的分类、功能及应用场景,并指出了未来研究方向。

Comments Accepted to IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18831 2025-11-25 cs.CV 79%

VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction

VideoCompressa: 通过联合时间压缩和空间重建实现数据高效的视频理解

Shaobo Wang, Tianle Niu, Runkang Yang, Deshan Liu, Xu He, Zichen Wen, Conghui He, Xuming Hu, Linfeng Zhang

机构 * EPIC Lab, SJTU(上海交通大学EPIC实验室) KTH(瑞典皇家理工学院) Shanghai AI Laboratory(上海人工智能实验室) HKUST(香港科技大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 VideoCompressa通过联合时间压缩和空间重建,实现视频数据高效合成,显著提升数据效率和模型性能。

Comments 15 pages, 6 tables, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17943 2025-11-25 cs.CV 79%

SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System

SciEducator: 基于Deming循环多智能体系统的科学视频理解与教育

Zhiyu Xu, Weilong Yan, Yufei Shi, Xin Meng, Tao He, Huiping Zhuang, Ming Li, Hehe Fan

机构 * Jinan University(济南大学) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) Peking University(北京大学) University of Electronic Science and Technology of China(电子科技大学) South China University of Technology(华南理工大学) Guangming Laboratory(光明实验室) Zhejiang University(浙江大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 SciEducator通过Deming循环多智能体系统实现科学视频的自演化理解与教育,生成多模态教学内容并超越现有大语言模型和视频智能体。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13644 2025-11-18 cs.CV 79%

CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding

Shrenik Patel, Daivik Patel

机构 * Rutgers University(罗切斯特大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12530 2025-11-18 cs.CV 79%

ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding

Yuan Zhou, Litao Hua, Shilong Jin, Wentao Huang, Haoran Duan

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments Accepted to AAAI 2026. Code is available at: https://github.com/robin-hlt/AAAI26-ReaSon

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04668 2025-11-17 cs.CV 79%

SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding

Ellis Brown, Arijit Ray, Ranjay Krishna, Ross Girshick, Rob Fergus, Saining Xie

机构 * New York University(纽约大学) Boston University(波士顿大学) AllenAI Vercept

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments Project page: https://ellisbrown.github.io/sims-v

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16781 2025-11-14 cs.CV cs.AI cs.CL 79%

Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features

Shihao Ji, Zihui Song

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments This paper is being withdrawn because we have identified a significant error in the implementation of our self-supervised clustering approach. Specifically, our feature aggregation step inadvertently leaked temporal information across frames, which violates the core assumption of our training-free method. We sincerely apologize to the research community

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17047 2025-11-11 cs.CV cs.AI 79%

Controllable Hybrid Captioner for Improved Long-form Video Understanding

Kuleen Sasse, Efsun Sarioglu Kayi, Arun Reddy

机构 * Johns Hopkins University Applied Physics Laboratory(约翰霍普金斯大学应用物理实验室)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01481 2025-10-28 cs.CV cs.LG 79%

VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding

Zongxia Li, Xiyang Wu, Guangyao Shi, Yubin Qin, Hongyang Du, Fuxiao Liu, Tianyi Zhou, Dinesh Manocha, Jordan Lee Boyd-Graber

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12160 2025-10-15 cs.CV 79%

State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding

Jiahuan Zhou, Kai Zhu, Zhenyu Cui, Zichen Liu, Xu Zou, Gang Hua

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) the Huazhong University of Science and Technology(华中科技大学) Amazon.com, Inc(亚马逊公司)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24008 2025-10-07 cs.CV cs.AI 79%

FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning

Haonan Ge, Yiwei Wang, Kai-Wei Chang, Hang Wu, Yujun Cai

机构 * University of California, Merced(加州大学默塞德分校) University of California, Los Angeles(加州大学洛杉矶分校) The University of Queensland(昆士兰大学)

专题命中 视频理解 :video reasoning(title);video understanding(abstract);分类 cs.CV

Comments Underreview

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02778 2025-10-06 cs.CV 79%

AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding

Xian Zhang, Zexi Wu, Zinuo Li, Hongming Xu, Luqi Gong, Farid Boussaid, Naoufel Werghi, Mohammed Bennamoun

机构 * The University of Western Australia(西澳大学) Dalian University of Technology(大连理工大学) Khalifa University(卡利夫大学) Zhejiang Lab(浙江实验室)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23672 2025-09-30 cs.CV 79%

Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding

Xixi Jiang, Chen Yang, Dong Zhang, Pingcheng Dong, Xin Yang, Kwang-Ting Cheng

机构 * Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(电子与计算机工程系,香港科学与技术大学) School of Electronic Information and Communications, Huazhong University of Science and Technology(电子信息与通信学院,华中科技大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21451 2025-09-29 cs.CV cs.CL 79%

VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding

Abdul Waheed, Zhen Wu, Dareen Alharthi, Seungone Kim, Bhiksha Raj

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15800 2025-09-22 cs.CV cs.AI 79%

ChronoForge-RL: Chronological Forging through Reinforcement Learning for Enhanced Video Understanding

Kehua Chen

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14199 2025-09-19 cs.CV cs.AI cs.CL cs.LG 79%

Dense Video Understanding with Gated Residual Tokenization

Haichao Zhang, Wenhao Chai, Shwai He, Ang Li, Yun Fu

机构 * Northeastern University(东北大学) Princeton University(普林斯顿大学) University of Maryland(马里兰大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12145 2025-09-16 cs.CV 79%

Open-ended Hierarchical Streaming Video Understanding with Vision Language Models

Hyolim Kang, Yunsu Park, Youngbeom Yoo, Yeeun Choi, Seon Joo Kim

机构 * Yonsei University(延世大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13919 2025-09-03 cs.CV cs.AI cs.CL cs.LG cs.RO 79%

Temporal Preference Optimization for Long-Form Video Understanding

Rui Li, Xiaohan Wang, Yuhui Zhang, Orr Zohar, Zeyu Wang, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏