arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 4300 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 4300 篇

2605.23993 2026-05-29 cs.CV cs.AI cs.LG 97%

Nano World Models: A Minimalist Implementation of Future Video Prediction

纳米世界模型:未来视频预测的极简实现

Siqiao Huang, Partha Kaushik, Michael Chen, Hengkai Pan, Kaiwen Geng, Omar Chehab, Fernando Moreno-Pino, Max Simchowitz

机构 * DeepMind

专题命中 通用世界模型 :world model(title,summary_cn);world models(title,summary_cn);world model(title,summary_cn);world models(title,summary_cn)

AI总结 提出Nano World Models,一个基于扩散强迫的极简代码库,用于未来视频预测,支持可控研究世界模型的设计选择,并通过实验分析预测参数化、架构规模等因素对视频预测质量的影响。

Comments Project page: https://simchowitzlabpublic.github.io/nano-world-model/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18746 2026-03-30 cs.CV cs.AI 96%

Any4D: Open-Prompt 4D Generation from Natural Language and Images

Any4D: 从自然语言和图像生成开放提示的4D生成

Hao Li, Qiao Sun

专题命中 通用世界模型 :world model(summary_cn,abstract);world models(summary_cn,abstract);embodied world model(summary_cn,abstract);world model(summary_cn,abstract)

AI总结 本文提出Primitive Embodied World Models,通过限制视频生成时间范围,实现语言与视觉表示的细粒度对齐,降低学习复杂度,提升数据效率,并减少推理延迟,支持复杂任务的组合泛化。

Comments The authors identified issues in the 4D generation pipeline and evaluation that affect result validity. To ensure scientific accuracy, we will revise the methodology and experiments thoroughly before resubmitting. This version should not be cited or relied upon

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20926 2026-06-02 cs.SE 96%

Learning Reasoning World Models for Parallel Code

学习并行代码的推理世界模型

Gautam Singh, Arjun Guha, Bhavya Kailkhura, Harshitha Menon

专题命中 通用世界模型 :world model(title,summary_cn);world models(title,summary_cn);world model(title,summary_cn);world models(title,summary_cn)

AI总结 提出Parallel-Code World Models (PCWMs),通过推理LLM直接预测并行代码的工具输出,以解决并行代码训练数据稀缺问题,在数据竞争预测和性能分析任务上取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02912 2025-11-25 cs.MA cs.AI cs.LG cs.SY eess.SY 96%

Communicating Plans, Not Percepts: Scalable Multi-Agent Coordination with Embodied World Models

传达计划,而非感知:基于具身世界模型的可扩展多智能体协调

Brennen A. Hill, Mant Koh En Wei, Thangavel Jishnuanandh

机构 * Department of Computer Science University of Wisconsin-Madison(计算机科学系 明尼苏达大学) Department of Computer Science National University of Singapore(计算机科学系 新加坡国立大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);embodied world model(title,abstract);world model(title,abstract)

AI总结 本文提出基于具身世界模型的意图通信方法,通过端到端学习与工程化设计对比,展示在复杂环境下更优的协调能力。

Comments Published in the Proceedings of the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Scaling Environments for Agents (SEA). Additionally accepted for presentation in the NeurIPS 2025 Workshop: Embodied World Models for Decision Making (EWM) and the NeurIPS 2025 Workshop: Optimization for Machine Learning (OPT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05495 2025-05-12 cs.CV cs.RO 96%

Learning 3D Persistent Embodied World Models

Siyuan Zhou, Yilun Du, Yuncong Yang, Lei Han, Peihao Chen, Dit-Yan Yeung, Chuang Gan

机构 * Institution1(机构1) Institution2(机构2)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);embodied world model(title,abstract);world model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12639 2026-04-14 cs.CV 96%

RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization

RoboStereo: 双塔4D具身世界模型用于统一策略优化

Ruicheng Zhang, Guangyu Chen, Zunnan Xu, Zihao Liu, Zhizhou Zhong, Mingyang Zhang, Jun Zhou, Xiu Li

机构 * Tsinghua University(清华大学) X Square Robot HKUST(香港科技大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);embodied world model(title,abstract);world model(title,abstract)

AI总结 本文提出RoboStereo,通过双塔4D具身世界模型实现统一策略优化,解决几何幻觉和缺乏统一优化框架的问题,实验显示在细粒度操作任务中平均相对提升超过97%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09452 2026-01-15 cs.CV 96%

MAD: Motion Appearance Decoupling for efficient Driving World Models

MAD:用于高效驾驶世界模型的运动外观解耦

Ahmad Rahimi, Valentin Gerard, Eloi Zablocki, Matthieu Cord, Alexandre Alahi

机构 * Sorbonne Université(索邦大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);driving world model(title,abstract);world model(title,abstract)

AI总结 MAD通过解耦运动学习与外观合成,高效地将通用视频扩散模型转化为可控的驾驶世界模型,实现低计算成本和高性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26782 2026-04-22 cs.LG cs.AI cs.CV 96%

Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models

克隆确定性世界:潜在几何在长周期世界模型中的关键作用

Zaishuo Xia, Yukuan Lu, Xinyi Li, Yifan Xu, Yubei Chen

机构 * UC Davis(加州大学戴维斯分校) Open Path AI Foundation(Open Path AI基金会)

专题命中 通用世界模型 :world model(title,summary_cn);world models(title,summary_cn);world model(title,summary_cn);world models(title,summary_cn)

AI总结 本文提出Geometrically-Regularized World Models (GRWM),通过时间对比学习原理对潜在空间进行几何正则化,提升世界模型的克隆精度和稳定性,解决长周期高保真建模中的几何结构瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10781 2026-07-14 cs.CV cs.RO 新提交 96%

Is Energy Guidance All You Need? Training-Free Norm Injection for Driving World Models

你所需要的只是能量引导吗?用于驱动世界模型的无训练规范注入

Xiyan Su, Frank Diermeyer, Markus Lienkamp

机构 * Institute of Automotive Technology, Technical University of Munich(慕尼黑工业大学汽车技术研究所)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);driving world model(title,abstract);world model(title,abstract)

AI总结 研究基于大型视频扩散主干的驾驶世界模型难以控制的问题,提出通过可微能量函数在采样时操控规划轨迹,无需重新训练主干,以Open-Sora 2.0 MM-DiT主干模型为例验证,指出跨流耦合是端到端可控关键。

Comments Accepted to Robotics: Science and Systems 2026 Robot World Models Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16234 2026-08-18 cs.CV 新提交 96%

GaussianDWM++: Language-Grounded 3D Gaussian Driving World Model for Unified Scene Understanding, Editing, and Multi-Modal Generation

GaussianDWM++:用于统一场景理解、编辑与多模态生成的语言接地3D高斯驾驶世界模型

Tianchen Deng, Xuefeng Chen, Shuang Wu, Qu Chen, Jiajun Zhu, Bo Dai, Jianfei Yang, Hesheng Wang

机构 * School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院) Tsinghua University(清华大学) Chongqing Afari Intelligent Drive Co., Ltd.(重庆阿法瑞智能驾驶有限公司) Nanyang Technological University(南洋理工大学)

专题命中 通用世界模型 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 该研究提出GaussianDWM++,通过基础特征高斯分词器等模块构建统一框架,在多类驾驶任务中实现场景理解、编辑与多模态生成,性能达当前最优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24322 2026-05-26 cs.CV 96%

Causal Physics Steering in Video World Models via Concept Activation Vectors

通过概念激活向量在视频世界模型中进行因果物理引导

Nahid Alam

机构 * Oreon Labs(Oreon实验室) Cohere Labs Community(Cohere实验室社区)

专题命中 通用世界模型 :world model(title,abstract);video world model(title,abstract);world models(title,abstract);world model(title,abstract)

AI总结 提出一种无需训练的方法,利用物理涌现区(PEZ)的概念激活向量(CAV)在推理时引导视频模型的物理期望,无需修改模型权重。

Comments In proceedings of CVPR 2026 workshop on Video World Model

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12705 2025-06-19 cs.RO cs.AI cs.LG 96%

DreamGen: Unlocking Generalization in Robot Learning through Video World Models

Joel Jang, Seonghyeon Ye, Zongyu Lin, Jiannan Xiang, Johan Bjorck, Yu Fang, Fengyuan Hu, Spencer Huang, Kaushil Kundalia, Yen-Chen Lin, Loic Magne, Ajay Mandlekar, Avnish Narayan, You Liang Tan, Guanzhi Wang, Jing Wang, Qi Wang, Yinzhen Xu, Xiaohui Zeng, Kaiyuan Zheng, Ruijie Zheng, Ming-Yu Liu, Luke Zettlemoyer, Dieter Fox, Jan Kautz, Scott Reed, Yuke Zhu, Linxi Fan

机构 * NVIDIA University of Washington(华盛顿大学) KAIST(韩国科学技术院) UCLA(加州大学洛杉矶分校) UCSD(加州大学圣地亚哥分校) CalTech(加州理工学院) NTU(国立台湾大学) University of Maryland(马里兰大学) UT Austin(得克萨斯大学奥斯汀分校)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

Comments See website for videos: https://research.nvidia.com/labs/gear/dreamgen

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22286 2026-03-24 cs.CV cs.AI cs.CL cs.LG 96%

WorldCache: Content-Aware Caching for Accelerated Video World Models

WorldCache: 用于加速视频世界模型的内容感知缓存

Umair Nawaz, Ahmed Heakl, Ufaq Khan, Abdelrahman Shaker, Salman Khan, Fahad Shahbaz Khan

机构 * Mohamed bin Zayed University of Artificial Intelligence, UAE(穆罕默德·本·扎耶德人工智能大学,阿联酋)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 WorldCache通过引入运动自适应阈值、显著性加权漂移估计和最优近似方法,实现动态场景下的高效特征重用,提升推理速度并保持高质量。

Comments 33 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01528 2026-03-10 cs.CV cs.AI cs.RO 96%

DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving

DrivingGen:自主驾驶中生成视频世界模型的综合性基准

Yang Zhou, Hao Shao, Letian Wang, Zhuofan Zong, Hongsheng Li, Steven L. Waslander

机构 * University of Toronto(多伦多大学) CUHK MMLab(香港中文大学MMLab)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title);world model(title,abstract)

AI总结 DrivingGen提出首个综合性基准,用于评估生成驾驶世界模型的视觉真实性、轨迹合理性、时间一致性和可控性,揭示通用与专用模型间的权衡。

Comments ICLR 2026 Poster; Project Website: https://drivinggen-bench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11419 2025-04-29 cs.AI cs.NE 96%

Embodied World Models Emerge from Navigational Task in Open-Ended Environments

Li Jin, Liu Jia

机构 * Tsinghua Laboratory of Brain and Intelligence(清华大学脑科学与智能实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);embodied world model(title,abstract);world model(title,abstract)

Comments Research on explainable meta-reinforcement learning AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14022 2026-08-17 cs.CV cs.AI 新提交 96%

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

ForgeWM:面向少步骤动作条件视频世界模型的渐进式因果训练

Xinye Li, Lingshuai Lin, Lei Wang, Liuzhou Zhang, Jialin Cui, Qingshan Li, Guanchu Wang, Qingbin Liu, Xi Chen, Jiang Bian, Wai Lam

机构 * CUHK(香港中文大学) Tencent PCG(腾讯平台与内容事业群) FDU(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) HKUST(香港科技大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 ForgeWM 是一种渐进式框架,通过多阶段技术将双向动作条件视频生成器转化为少步骤世界模型,在 Minecraft 和 FPS 游戏任务中实现了更优的可控少步骤视频生成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10135 2026-07-27 cs.CV cs.AI 版本更新 96%

BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression

BiWM:利用双向自回归推进开源交互式视频世界模型

Shaohao Rui, Xiaofeng Mao, Zhanyu Zhang, Peijia Lin, Yansong Zhu, Yibo Zhang, Haibin Wan, Zhangrui Zhao, Weijie Ma

机构 * LynnReal AI Shanghai Innovation Institute(上海创新研究院) Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 提出BiWM框架,通过双向自回归范式将预训练视频骨干转化为交互式世界模型,仅需两阶段训练(微调+分布匹配蒸馏),支持多尺度模型和长程生成,优于现有因果流水线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10090 2026-05-26 cs.AI cs.CL cs.LG 96%

Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning

Agent World Model: 用于智能体强化学习的无限合成环境

Zhaoyang Wang, Canwen Xu, Boyi Liu, Yite Wang, Siwei Han, Zhewei Yao, Huaxiu Yao, Yuxiong He

机构 * University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 通用世界模型 :world model(title,title_cn);world model(title,title_cn);world-model(abstract,abstract_cn);world-model(abstract,abstract_cn)

AI总结 提出Agent World Model (AWM)全合成环境生成管道,通过代码驱动和数据库支持的环境进行大规模强化学习,使智能体在多样日常场景中泛化。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04002 2026-02-25 cs.CV cs.RO eess.IV 96%

NRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World Models

NRSeg: 通过驾驶世界模型实现噪声鲁棒的BEV语义分割学习

Siyu Li, Fei Teng, Yihong Cao, Kailun Yang, Zhiyong Li, Yaonan Wang

机构 * School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University(人工智能与机器人学院和机器人视觉感知与控制技术国家工程研究中心,湖南大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);driving world model(title,abstract);world model(title,abstract)

AI总结 NRSeg通过驾驶世界模型生成的合成数据增强BEV语义分割学习,提出PGCM、BiDPP和HLSE模块以提升模型鲁棒性和分割性能。

Comments Accepted to IEEE Transactions on Image Processing (TIP). The source code will be made publicly available at https://github.com/lynn-yu/NRSeg

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08971 2026-02-12 cs.CV cs.RO 96%

WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models

WorldArena: 一个用于评估具身世界模型感知与功能效用的统一基准

Yu Shang, Zhuohang Li, Yiding Ma, Weikang Su, Xin Jin, Ziyou Wang, Lei Jin, Xin Zhang, Yinzhou Tang, Haisheng Su, Chen Gao, Wei Wu, Xihui Liu, Dhruv Shah, Zhaoxiang Zhang, Zhibo Chen, Jun Zhu, Yonghong Tian, Tat-Seng Chua, Wenwu Zhu, Yong Li

机构 * Tsinghua University, Beijing, China(清华大学) Shanghai Jiao Tong University, Shanghai, China(上海交通大学) The University of Hong Kong, Hong Kong SAR, China(香港大学) Princeton University, Princeton, NJ, USA(普林斯顿大学) Chinese Academy of Sciences, Beijing, China(中国科学院) University of Science(科学技术大学) Peking University, Beijing, China(北京大学) National University of Singapore, Singapore(新加坡国立大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);embodied world model(title,abstract);world model(title,abstract)

AI总结 WorldArena是一个统一的基准,用于评估具身世界模型的感知和功能效用,揭示高视觉质量与强具身任务能力之间的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07050 2026-02-10 cs.CV cs.AI 96%

Interpreting Physics in Video World Models

视频世界模型中的物理解释

Sonia Joseph, Quentin Garrido, Randall Balestriero, Matthew Kowal, Thomas Fel, Shahab Bakhtiari, Blake Richards, Mike Rabbat

机构 * FAIR, Meta Superintelligence Labs Mila \& McGill University Kempner Institute, Harvard University

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 该研究通过分析视频模型内部的物理表示,发现其采用分布式表示而非分解表示,揭示了物理信息在模型中的涌现机制和组织方式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20840 2025-11-25 cs.RO cs.AI cs.MM 96%

Learning Primitive Embodied World Models: Towards Scalable Robotic Learning

学习原始具身世界模型:迈向可扩展的机器人学习

Qiao Sun, Liujia Yang, Wei Tang, Wei Huang, Kaixin Xu, Yongchao Chen, Mingyu Liu, Jiange Yang, Haoyi Zhu, Yating Wang, Tong He, Yilun Chen, Xili Dai, Nanyang Ye, Qinying Gu

机构 * Shanghai AI Lab(上海人工智能实验室) Fudan(复旦大学) SJTU(上海交通大学) NJUST(南京理工大学) THU(清华大学) Harvard(哈佛大学) ZJU(浙江大学) NJU(南京大学) USTC(中国科学技术大学) Tongji(同济大学) HKUST (GZ)(香港科技大学(广州))

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);embodied world model(title,abstract);world model(title,abstract)

AI总结 提出原始具身世界模型(PEWM)以解决具身数据稀疏和高维问题,通过短视界视频生成实现细粒度对齐、降低学习复杂度并提高数据效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15511 2025-11-24 cs.RO cs.AI 96%

AeroVerse: UAV-Agent Benchmark Suite for Simulating, Pre-training, Finetuning, and Evaluating Aerospace Embodied World Models

AeroVerse:用于模拟、预训练、微调和评估航空航天具身世界模型的UAV-Agent基准套件

Fanglong Yao, Yuanchang Yue, Youzhi Liu, Xian Sun, Kun Fu

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所) University of Chinese Academy of Sciences(中国科学院大学) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院) Key Laboratory of Target Cognition and Application Technology(TCAT), Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所目标认知与应用技术重点实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);embodied world model(title,abstract);world model(title,abstract)

AI总结 AeroVerse是一个用于模拟、预训练、微调和评估航空航天具身世界模型的基准套件,包含多个数据集和评估指标,旨在推动航空航天具身智能的发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11601 2026-08-13 cs.CV 新提交 96%

How Can Driving World Models Do Counterfactual Prediction?

驾驶世界模型如何能进行反事实预测?

Jiaru Zhang, Can Cui, Yi Xu, Xin Ye, Ruqi Zhang, Ziran Wang

机构 * Purdue University(普渡大学) Bosch Center for Artificial Intelligence(博世人工智能中心)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);driving world model(title,abstract);world model(title,abstract)

AI总结 本文指出驾驶世界模型的反事实预测目标与直接动作条件预测存在根本差距,构建基准验证该问题,并提出简单无训练流程提升反事实预测效果,呼吁开发更好的相关方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19990 2026-07-08 cs.AI 新提交 96%

Reward as An Agent for Embodied World Models

奖励作为具身世界模型的智能体

Pu Li, Zhigang Lin, Qiang Wu, Yongxuan Lv, Fei Wang, Shan You

机构 * ACE Robotics(ACE机器人)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);embodied world model(title,abstract);world model(title,abstract)

AI总结 提出奖励智能体框架和动态感知 rollout 多样化方法,通过鲁棒验证支持更广泛探索,缓解奖励黑客问题,提升世界模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23105 2026-06-23 cs.CV 新提交 96%

Compression and Retrieval: Implicit Memory Retrieval for Video World Models

压缩与检索:视频世界模型的隐式记忆检索

Zhan Peng, Jie Ma, Huiqiang Sun, Chong Gao, Zhijie Xue, Zhiyu Pan, Zhiguo Cao, Jun Liang, Jing Li

机构 * Huazhong University of Science and Technology(华中科技大学) HUJING Digital Media & Entertainment Group(虎鲸数字媒体与娱乐集团) Sun Yat-sen University(中山大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 提出注意力驱动的隐式记忆检索机制CaR,通过位置编码注入视角信息实现灵活检索,并引入轻量级上下文压缩网络,在构建的SceneFly数据集上取得SOTA结果并展现强泛化性。

Comments Project page: 3DV-Team/CaR" target="_blank" rel="noopener">https://github.com/Orange-3DV-Team/CaR

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09828 2026-06-09 cs.CV 新提交 96%

Latent Spatial Memory for Video World Models

视频世界模型的潜在空间记忆

Weijie Wang, Haoyu Zhao, Yifan Yang, Feng Chen, Zeyu Zhang, Yefei He, Zicheng Duan, Donny Y. Chen, Yuqing Yang, Bohan Zhuang

机构 * Zhejiang University(浙江大学) Microsoft Research(微软研究院) Adelaide University(阿德莱德大学) Monash University(莫纳什大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 提出潜在空间记忆框架Mirage,通过在扩散潜在空间中直接构建和查询3D缓存,避免像素空间重建,实现高效视频生成,速度提升10.57倍,内存减少55倍。

Comments Project Page: https://aka.ms/latent-spatial-memory, Code: https://github.com/microsoft/LatentSpatialMemory

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02436 2026-06-02 cs.CV 96%

Geometry-Aware Implicit Memory for Video World Models

几何感知隐式记忆用于视频世界模型

Zhengxuan Wei, Xu Guo, Xinghui Li, Xunzhi Xiang, Min Wei, Yiran Zhu, Qiulin Wang, Xintao Wang, Pengfei Wan, Xiangwang Hou, Qi Fan

机构 * School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院) Kling Team, Kuaishou Technology(快手技术 Kling 团队) Tsinghua University(清华大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 提出GIM-World框架,通过轻量级Transformer编码器将可变长度历史压缩为固定大小的记忆令牌,并利用相机可查询的几何头在训练期间从冻结的基础模型中蒸馏3D场景结构,从而在长时程视频生成中保持几何和视觉一致性。

Comments Project page: https://gim-world.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30263 2026-05-29 cs.CV 96%

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

minWM: 用于实时交互式视频世界模型的全栈开源框架

Min Zhao, Hongzhou Zhu, Bokai Yan, Zihan Zhou, Yimin Chen, Wenqiang Sun, Kaiwen Zheng, Guande He, Xiao Yang, Chongxuan Li, Fan Bao, Jun Zhu

机构 * ShengShu(盛书) THU(清华大学) RUC(中国人民大学) HKUST(香港科技大学) UT-Austin(德克萨斯大学奥斯汀分校)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 提出minWM全栈开源框架,通过因果强制/因果强制++流水线将双向视频扩散模型转化为可控制、低延迟的自回归世界模型,支持相机控制与多种骨干架构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18564 2026-04-22 cs.CV 96%

MultiWorld: Scalable Multi-Agent Multi-View Video World Models

多世界:可扩展的多智能体多视角视频世界模型

Haoyu Wu, Jiwen Yu, Yingtian Zou, Xihui Liu

机构 * The University of Hong Kong(香港大学) Sreal AI

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 多世界框架通过多智能体控制模块和全局状态编码器,实现多智能体多视角视频世界建模,提升视频保真度和多视角一致性。

Comments 15 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏