arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2026-08-18 至 2026-08-18 共收录 8 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 8 篇

2608.16320 2026-08-18 cs.CV 新提交 79%

StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding

StreamOPD:用于流视频理解的时空线索门控后训练方案

Keming Wu, Baoyi Wang, Kaichen Zhang, Xiang An, Zuhao Yang, Sudong Wang, Haowei Zhu, Tingxuan Huang, Hongcheng Gao, Bin Wang

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 本研究提出 StreamOPD 方案,结合可验证流视频数据、思考模式 OPD 与指令模式部署,提升了 StreamingBench、OVO-Bench 等流视频理解基准性能,为相关开源研究提供参考。

Comments Project page: this https URL (https://unix-ai-lab.github.io/StreamOPD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14718 2026-08-18 cs.CV cs.CL 新提交 79%

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

VideoGAIA:面向通用人工智能助手的智能体视频理解基准

Fan Zhang, Guangming Yao, Jinyang Wu, Hao Wu, Zheng Lian, Xinyu Geng, Jingdong Chen, Yi Yuan, Pheng-Ann Heng

机构 * The Chinese University of Hong Kong(香港中文大学) Ant Group(蚂蚁集团) Tsinghua University(清华大学) Tongji University(同济大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 VideoGAIA是面向通用AI助手的智能体视频理解基准,构建多轮工具增强交互任务,现有MLLMs在其上准确率不足60%,可评估下一代模型并推动相关转型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16285 2026-08-18 cs.CV cs.AI 新提交 57%

Audio-Visual Segmentation via Depth-Guided Collaborative Modeling

基于深度引导协同建模的视听分割

Zhaojin Fu, Yuyang Hong, Qi Yang, Zili Wang, Kun Ding, Shiming Xiang, Bin Fan

机构 * School of Intelligent Science and Technology, University of Science and Technology Beijing(北京科技大学智能科学与技术学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

AI总结 该研究针对现有视听分割方法未建模几何线索的问题,提出三模态框架DGCM-AVS,通过深度感知动态调制器与深度引导渐进融合模块优化,在AVSS数据集上实现M_J、M_F指标的显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15614 2026-08-18 cs.CV cs.AI 新提交 57%

EgoGazeLite: On-Device Egocentric Gaze Prediction for Token-Efficient Multimodal LLM Video Input

EgoGazeLite:面向 token 高效多模态大语言模型视频输入的设备内自我中心注视预测

Matteo Stoiber, Niels Buus Lassen

机构 * Copenhagen Business School(哥本哈根商学院)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

AI总结 EgoGazeLite 是轻量双进程注视预测器,可在消费级硬件上实时运行,无需眼动追踪硬件即可实现 token 高效的自我中心视频理解,性能与真实注视裁剪无显著差异。

Comments 16 pages. Accepted at the WearableAI Workshop, ECCV 2026 (Archival Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14854 2026-08-18 cs.CV 新提交 57%

Zero-MELO: Test-Time Evidence Calibration with Multimodal LLMs for Zero-Shot Micro-Gesture Recognition

Zero-MELO:基于多模态大语言模型的测试时证据校准用于零样本微手势识别

Chengyan Wang, Hanliang Xie, Yueyi Yang, Haoyu Chen

机构 * University of Oulu(奥卢大学) Peking University(北京大学)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

AI总结 该研究针对多模态大语言模型在微手势识别中局部证据不足、分数偏差的瓶颈,提出Zero-MELO框架,结合树搜索、测试时校准与多线索融合,在iMiGUE和MA-52数据集上显著优于Qwen2.5-VL基线。

Comments Accepted by ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08974 2026-08-18 cs.CV cs.AI 版本更新 57%

Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models

追踪真相:面向视频大语言模型的物体感知时空监控

Tri Cao, Khoi Le, Thong Nguyen, Cong-Duy Nguyen, Quynh Vo, Anh Tuan Luu, Chunyan Miao, See-Kiong Ng, Shuicheng Yan, Bryan Hooi

机构 * National University of Singapore(新加坡国立大学) VinUniversity(文大学) Nanyang Technological University(南洋理工大学)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

AI总结 本文提出STEMO-Bench基准测试,通过分解查询验证时空监控能力,改进视频大语言模型的时空推理一致性。

Comments The authors are withdrawing this manuscript due to errors identified in the experimental evaluation and result aggregation, which affect several reported quantitative results and some conclusions. These issues require substantial re-evaluation of the experiments and analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12478 2026-08-18 cs.CV cs.LG 版本更新 57%

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning

数据更少,收敛更快:面向多模态指令微调的目标驱动数据优化

Rujie Wu, Haozhe Zhao, Hai Ci, Yizhou Wang

机构 * Peking University(北京大学) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) National University of Singapore(新加坡国立大学)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

AI总结 本文提出目标驱动数据优化框架GDO,通过优化训练样本实现更快收敛和更高精度,适用于多模态指令微调任务。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23649 2026-08-18 cs.RO cs.CV 版本更新 57%

RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion

RoboMirror: 在模仿之前理解以实现视频到仿人运动

Zhe Li, Boan Zhu, Yangyang Wei, Shuanghao Bai, Yuheng Ji, Yibo Peng, Tao Huang, Pengwei Wang, Zhongyuan Wang, S.-H. Gary Chan, Chang Xu, Cheng Chi, Jianfei Yang, Shanghang Zhang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

AI总结 RoboMirror通过视觉语言模型将视频提炼为视觉运动意图,直接生成物理合理的仿人运动,无需姿态重建,实现远程存在并提升任务成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏