arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2026-07-21 至 2026-07-21 共收录 2 信号源:cs.CV, eess.IV, cs.MM

1. 长视频与时序推理 2 篇

2606.19341 2026-07-21 cs.CV cs.CL cs.SD 版本更新 70%

Native Active Perception as Reasoning for Omni-Modal Understanding

原生主动感知作为全模态理解的推理

Zhenghao Xing, Ruiyang Xu, Yuxuan Wang, Jinzheng He, Ziyang Ma, Qize Yang, Yunfei Chu, Jin Xu, Junyang Lin, Chi-Wing Fu, Pheng-Ann Heng

机构 * The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) Qwen Team, Alibaba Group(阿里巴巴集团Qwen团队)

专题命中 长视频与时序推理 :video understanding(abstract);long video(abstract);分类 cs.CV

AI总结 提出OmniAgent,一种基于POMDP迭代观察-思考-行动循环的原生全模态智能体,通过主动感知将推理复杂度与视频时长解耦,在多个基准上达到开源模型最优性能。

Comments Accepted at ICML 2026. Code and models: https://github.com/harryhsing/omniagent

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01483 2026-07-21 cs.RO cs.AI 版本更新 50%

Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering

VL-KnG:从单目视频构建持久时空知识图谱以实现具身场景理解

Mohamad Al Mdfaa, Svetlana Lukina, Timur Akhtyamov, Arthur Nigmatzyanov, Dmitrii Nalberskii, Sergey Zagoruyko, Gonzalo Ferrer

机构 * Applied AI Institute, Moscow, Russia(莫斯科应用人工智能研究所)

专题命中 长视频与时序推理 :long video(abstract)

AI总结 VL-KnG通过构建时空知识图谱实现高效具身场景理解,采用LLM驱动的时空对象关联和图增强检索,无需3D重建,支持常数时间推理,优于现有前沿VLMs。

详情

展开后加载摘要…

URL PDF HTML 收藏