arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2026-08-19 至 2026-08-19 共收录 3 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1 篇

2604.01460 2026-08-19 cs.CV 版本更新 57%

Reinforcing Consistency in Video MLLMs with Structured Rewards

通过结构化奖励强化视频MLLMs的一致性

Yihao Quan, Zeru Shi, Jinman Zhao, Ruixiang Tang

机构 * Rutgers University(罗格斯大学) University of Toronto(多伦多大学)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

AI总结 研究通过结构化奖励提升视频MLLMs的一致性,发现传统监督不足,提出结合事实和时间单元的奖励机制,提升视频理解准确性。

Comments Accepted by COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 1 篇

2608.14790 2026-08-19 cs.CV 版本更新 57%

Qwen-Video-Edit: Instruction-Based Video Editing by Repurposing an Image Editing Model

Qwen-Video-Edit:通过复用图像编辑模型实现基于指令的视频编辑

Yunpeng Bai, Yossi Gandelsman, Michaël Gharbi, Qixing Huang

机构 * UT Austin(德克萨斯大学奥斯汀分校) Reve(Reve公司)

专题命中 视频生成 :video diffusion(abstract);分类 cs.CV

AI总结 该研究提出Qwen-Video-Edit,通过复用Qwen-Image-Edit图像编辑模型,经少量适配实现基于指令的视频编辑,证明图像编辑先验可迁移至视频编辑任务。

Comments Project Page: this https URL (https://yunpeng1998.github.io/Qwen-Video-Edit-Page) Code: this https URL (https://github.com/yunpeng1998/Qwen-Video-Edit) Model: this https URL (https://huggingface.co/yunpeng1998/Qwen-Video-Edit)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频数据与评测 1 篇

2601.21282 2026-08-19 cs.CV 版本更新 57%

WorldBench: Benchmarking Physical Understanding of World Models by Isolating Physics Concepts

WorldBench: 为世界模型诊断评估进行物理辨析

Rishi Upadhyay, Howard Zhang, Jim Solomon, Ayush Agrawal, Yunhao Ba, Alex Wong, Celso M de Melo, Achuta Kadambi

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Sony AI(索尼人工智能) Yale University(耶鲁大学) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 视频数据与评测 :video generation(abstract);分类 cs.CV

AI总结 WorldBench通过概念特定的解耦评估,提升世界模型的物理推理能力评估的准确性和可扩展性。

Comments Webpage: this https URL (https://world-bench.github.io/)

详情

展开后加载摘要…

URL PDF HTML 收藏