arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 194 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 194 篇

2606.10135 2026-07-27 cs.CV cs.AI 版本更新 96%

BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression

BiWM:利用双向自回归推进开源交互式视频世界模型

Shaohao Rui, Xiaofeng Mao, Zhanyu Zhang, Peijia Lin, Yansong Zhu, Yibo Zhang, Haibin Wan, Zhangrui Zhao, Weijie Ma

机构 * LynnReal AI Shanghai Innovation Institute(上海创新研究院) Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 提出BiWM框架,通过双向自回归范式将预训练视频骨干转化为交互式世界模型,仅需两阶段训练(微调+分布匹配蒸馏),支持多尺度模型和长程生成,优于现有因果流水线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26182 2026-07-22 cs.CV cs.AI cs.LG 版本更新 95%

Lifting Embodied World Models for Planning and Control

提升具身世界模型用于规划与控制

Alex N. Wang, Trevor Darrell, Pavel Izmailov, Yutong Bai, Amir Bar

机构 * Computer Science, New York University(纽约大学计算机科学系) BAIR, UC Berkeley(伯克利大学BAIR)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);embodied world model(title);world model(title,abstract)

AI总结 本文提出一种轻量级策略,将高层动作映射到低层关节动作序列,结合冻结的世界模型,实现预测未来观察的提升世界模型,有效降低规划复杂度。

Comments Accepted to ECCV2026. Edited policy masking

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17796 2026-06-25 cs.CV cs.AI 版本更新 95%

CustomX: Unified Character, Action, and Scene Customization in Video World Models

CustomX: 视频世界模型中的统一角色、动作与场景定制

Yitong Wang, Fangyun Wei, Hongyang Zhang, Bo Dai, Yan Lu

机构 * Fudan University(复旦大学) Microsoft Research(微软研究院) University of Waterloo(滑铁卢大学) The University of Hong Kong(香港大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title);world model(title,abstract)

AI总结 提出CustomX,结合静态世界生成与可控实体模型,支持用户指定角色在3D场景中执行开放动作,通过条件自回归视频生成保持视觉保真度。

Comments Accepted to ECCV 2026. Project page: https://snowflakewang.github.io/CustomX_Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02953 2026-08-11 cs.CV 版本更新 95%

RealWeather: Realistic and Scene-Faithful Weather Translation with Driving World Models

RealWeather:基于驾驶世界模型的逼真且场景忠实的天气转换

Yuwei Ning, Liangzhi Wang, Yi Xiao, Zhenhua Wu, Yun Pang, Mingkun Chang, Jichang Li, Guanbin Li

专题命中 通用世界模型 :world model(title,abstract);driving world model(title,abstract);world models(title);world model(title,abstract)

AI总结 RealWeather是一种驾驶世界模型,通过渐进式逼真度引导和场景忠实度强化学习优化实现逼真且场景忠实的天气转换,在视觉逼真度、结构保留等方面优于现有方法,支持长尾天气场景生成与零样本泛化。

Comments Under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16713 2026-06-12 cs.CV cs.AI 版本更新 95%

GeoWorld-VLM: Geometry from World Models for Vision-Language Models

GeoWorld-VLM:从世界模型中获取几何结构用于视觉-语言模型

Renjie Gu, Kaichen Zhou, Yan Luo, Mengyu Wang

机构 * Harvard AI and Robotics Lab(哈佛人工智能与机器人实验室) Kempner Institute for the Study of Natural and Artificial Intelligence(凯普纳自然与人工智能研究 institute) Harvard University(哈佛大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 GeoWorld-VLM通过将冻结的摄像机条件视频世界模型的几何结构转移到视觉-语言模型中,提升空间关系推理能力,实验显示在两个不同架构上均提升了约4%的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18135 2026-08-18 cs.CV 版本更新 95%

World-in-World: World Models in a Closed-Loop World

世界中的世界:闭环世界中的世界模型

Jiahan Zhang, Muqing Jiang, Nanru Dai, Taiming Lu, Arda Uzunoglu, Shunchi Zhang, Yana Wei, Jiahao Wang, Vishal M. Patel, Paul Pu Liang, Daniel Khashabi, Cheng Peng, Rama Chellappa, Tianmin Shu, Alan Yuille, Yilun Du, Jieneng Chen

机构 * JHU(约翰·霍普金斯大学) PKU(北京大学) Princeton(普林斯顿大学) MIT(麻省理工学院) Harvard(哈佛大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究人员推出首个具身场景闭环世界模型基准平台World-in-World,发现视觉质量不保障任务成功,后训练缩放比升级预训练视频生成器更有效,增加推理计算可提升闭环性能。

Comments ICLR 2026 Oral. Add acknowledgement in arxiv v2. Code is at https://github.com/World-In-World/world-in-world

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26712 2026-08-18 cs.RO 版本更新 95%

ActSWM: Action-Sensitive World Models for Long-Horizon Planning in Open-World Games

ActSWM:面向开放世界游戏长视界规划的动作敏感型世界模型

Zhenfeng Gan, ZiTong Zeng, Jiajun Cheng, Yeke Song, Yongyi Tang, Xueqian Wang

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 针对开放世界游戏长视界规划的上下文崩塌问题,提出动作敏感型世界模型ActSWM,通过约束隐式回滚保留动作依赖差异,提升了任务成功率与动作恢复能力。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15142 2026-08-11 cs.AI cs.LG 版本更新 95%

Concept-Guided Spatial Regularization for World Models in Atari Pong

《雅达利乒乓球游戏中用于世界模型的概念引导空间正则化》

Yukuan Lu, Zaishuo Xia, Weyl Lu, Yubei Chen

机构 * UC Davis(加州大学戴维斯分校)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究雅达利乒乓球游戏中五个视觉世界模型智能体,发现其存在诸多问题。提出概念引导空间正则化(CGSReg),实验表明该方法在部分模型中改善了闭环展开和像素空间零样本MBRL,但不能解决所有世界模型瓶颈。

Comments Revised manuscript with updated presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12920 2026-07-13 cs.MA cs.AI cs.CL 版本更新 94%

Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue

通过对话对齐世界模型实现具身多智能体协调

Vardhan Dongre, Dilek Hakkani-Tür

机构 * Siebel School of Computing & Data Science(计算机与数据科学学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究通过对话机制探索具身智能体的世界模型对齐,发现对话能减少冲突但降低任务成功率,提出评估世界模型对齐的框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01801 2026-06-15 cs.CV cs.AI 版本更新 94%

Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention

快速自回归视频扩散与世界模型:基于时间缓存压缩与稀疏注意力

Dvir Samuel, Issar Tzachor, Matan Levy, Michael Green, Gal Chechik, Rami Ben-Ari

机构 * Hebrew University of Jerusalem(特拉维夫大学) Google Research(谷歌研究)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出FAST-AR框架,通过TempCache压缩KV缓存、AnnCA加速交叉注意力、AnnSA稀疏化自注意力,实现自回归视频扩散模型5-10倍加速,同时保持视觉质量并稳定GPU内存使用。

Comments Accepted to ICML 2026. Project Page: https://dvirsamuel.github.io/fast-auto-regressive-video/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21282 2026-08-19 cs.CV 版本更新 94%

WorldBench: Benchmarking Physical Understanding of World Models by Isolating Physics Concepts

WorldBench: 为世界模型诊断评估进行物理辨析

Rishi Upadhyay, Howard Zhang, Jim Solomon, Ayush Agrawal, Yunhao Ba, Alex Wong, Celso M de Melo, Achuta Kadambi

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Sony AI(索尼人工智能) Yale University(耶鲁大学) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 WorldBench通过概念特定的解耦评估,提升世界模型的物理推理能力评估的准确性和可扩展性。

Comments Webpage: https://world-bench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13552 2026-08-17 cs.CV 版本更新 94%

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

PlayWorld:基于智能体玩家的长程目标世界模型基准测试

Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao

机构 * The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Zhejiang University(浙江大学) Kuaishou Technology(快手科技)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 该研究针对现有世界模型跨模型公平比较的难题,推出含171个场景的PlayWorld基准,通过多模态智能体玩家从多维度评估9种先进世界模型,发现其长程交互式目标表现仍不可靠。

Comments project page: https://kxding.github.io/project/PlayWorld/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06257 2026-08-11 cs.CV cs.HC 版本更新 94%

MASS: Multiplayer World Models with Authoritative Shared State

MASS:具有权威共享状态的多人世界模型

Ziqi Cai, Siqi Yang, Yimu Wang, Zixian Gao, Yunheng Liu, Shuchen Weng, Erwin Wu, Kaipeng Zhang, Boxin Shi

机构 * Alaya Lab(Alaya实验室) Peking University(北京大学) Institute of Science Tokyo(东京科学大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 该研究提出MASS模型,通过将世界动态与视图渲染解耦,解决现有视频世界模型在多人环境中的缺陷,在多人Snake基准测试中实现更优状态精度,为多智能体世界模拟提供实用基础。

Comments Project Page: https://alaya-lab.github.io/MASS/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30880 2026-07-29 cs.CL cs.AI 版本更新 94%

PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments

PatchWorld:可执行世界模型的免梯度优化

Jiaxin Bai, Yue Guo, Yifei Dong, Jiaxuan Xiong, Tianshi Zheng, Yixia Li, Tianqing Fang, Yufei Li, Yisen Gao, Haoyu Huang, Zhongwei Xie, Hong Ting Tsang, Zihao Wang, Lihui Liu, Jeff Z. Pan, Yangqiu Song

机构 * Hong Kong Baptist University(香港 Baptist 大学) Independent Researcher(独立研究员) HKUST(香港科技大学) Beijing Institute of Technology(北京理工大学) Southern University of Science and Technology(南方科技大学) Wayne State University(韦恩州立大学) University of Edinburgh(爱丁堡大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出 PatchWorld 框架,通过反例引导的代码修复将离线轨迹转化为可执行的 Python 世界模型,实现无需梯度优化的符号信念状态程序,在 AgentGym 环境中达到 76.4% 的宏观成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23164 2026-07-02 cs.LG 版本更新 94%

MetaOthello: A Controlled Study of Multiple World Models in Transformers

MetaOthello: Transformer中多个世界模型的控制研究

Aviral Chawla, Galen Hall, Juniper Lovato

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 通过引入共享语法但规则或分词不同的Othello变体套件,训练小型GPT处理混合数据,发现Transformer不划分容量为孤立子模型,而是收敛于跨变体因果转移的共享棋盘状态表示。

Comments Camera-ready version. Accepted to the 43rd International Conference on Machine Learning (ICML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24354 2026-06-23 cs.CV 版本更新 94%

SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation

SparseWorld: 通过具有稀疏场景表示的世界模型增强端到端自动驾驶

Ruoyu Wang, Jingke Wang, Yukai Ma, Yuehao Huang, Shuangming Lei, Guanglin Xu, Aixue Ye, Yong Liu

机构 * Institute of Cyber-Systems and Control, Zhejiang University(浙江大学控制系统研究所) Labs, Huawei(华为2012实验室) State Key Laboratory of Industrial Control Technology(国家工业控制技术重点实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出SparseWorld,一种基于稀疏场景表示的轻量级世界模型,通过自回归预测未来地图元素和周围智能体,并利用预测结果优化下游运动预测和轨迹规划,在nuScenes数据集上实现0.05%的碰撞率,达到开放循环规划指标的最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18313 2026-06-23 cs.CV 版本更新 94%

OmniNWM: Omniscient Driving Navigation World Models

OmniNWM:全知驾驶导航世界模型

Bohan Li, Zhuang Ma, Dalong Du, Baorui Peng, Zhujin Liang, Zhenqiang Liu, Xianda Guo, Zheng Zhu, Chao Ma, Yueming Jin, Xin Jin, Hao Zhao, Wenjun Zeng

机构 * Shanghai Jiao Tong University(上海交通大学) Eastern Institute of Technology(东部技术研究所) PhiGent National University of Singapore(新加坡国立大学) Tsinghua University(清华大学) Wuhan University(武汉大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出OmniNWM全知全景导航世界模型,在统一概率框架下处理状态、动作和奖励三个维度,实现多模态全景视频生成、零样本轨迹控制及内在奖励机制,在生成保真度和控制精度上达到SOTA。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05138 2026-06-09 cs.AI 版本更新 94%

Executable World Models for ARC-AGI-3 in the Era of Coding Agents

可执行世界模型在编码智能体时代的ARC-AGI-3应用

Sergey Rodionov

机构 * SingularityNET

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出一种编码智能体系统,通过维护可执行Python世界模型、验证观察、重构简化抽象和模型内规划,在ARC-AGI-3游戏中取得初步成果,GPT-5.5高推理下完全解决15个游戏。

Comments 13 pages. Accepted for publication at AGI-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28455 2026-08-17 cs.RO cs.AI cs.LG 版本更新 94%

Event-Conditioned Diagnostics of Kinematic, Contact, and Object-Permanence Fields in Passive Object-State World Models

被动对象状态世界模型中运动学、接触和物体恒存场的事件条件诊断

Yang Liu, Yuming Chen

机构 * College of Intelligent Robitcs and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出一种诊断协议,测试被动对象状态世界模型中的隐式物理场是否按事件类型组织,并验证场对齐表示对预测的功能影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04546 2026-08-20 cs.RO cs.AI cs.CV cs.LG 版本更新 94%

Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models

Mask2Real-WM:作为可控灵巧世界模型的模拟到现实桥梁的分割掩码

Riccardo O. Feingold, Davide Liconti, Chenyu Yang, Robert K. Katzschmann

机构 * Soft Robotic Lab, Department of Mechanical and Process Engineering ETH Zurich, Switzerland(苏黎世联邦理工学院机械与过程工程系软机器人实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出Mask2Real-WM,一种用于灵巧操纵的两阶段动作条件世界模型,将像素预测解耦为动力学和渲染模型,利用分割空间较小的模拟到现实差距,通过合成数据预训练和真实演示微调实现各自由度动作可控。

Comments 23 pages, 24 figures, 4 tables. Preprint. Project page: https://srl-ethz.github.io/Mask2Real-WM/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07303 2026-08-13 cs.LG 版本更新 94%

Bootstrap Theory of Representational Emergence: Explanatory Insufficiency as a Driver of Representation Learning and World Models

表征涌现的自举理论:解释不充分性作为表征学习与世界模型的驱动力

Jacques Raynal, Pierre Slangen, Elsa Raynal, Jacques Margerit

机构 * Laboratory of Bioengineering and Nanosciences (LBN), University of Montpellier(生物工程与纳米科学实验室(LBN),蒙彼利埃大学) EuroMov Digital Health in Motion, University of Montpellier, IMT Mines Alès(EuroMov数字健康运动,蒙彼利埃大学,IMT矿山阿尔勒) Certified Sophrologist, Sensorimotor Practice, Montpellier, France(认证Sophrologist,运动觉实践,蒙彼利埃,法国) Emeritus Professor, University of Montpellier(荣誉教授,蒙彼利埃大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出表征涌现自举理论(TBER),将解释不充分性视为新表征涌现的积极信号,通过五阶段递归过程驱动表征创新,应用于表征学习、世界模型和科学发现。

Comments 27 pages, 25 references, no figures or tables. Conceptual framework on representational emergence and explanatory insufficiency, with implications for representation learning, world models, autonomous AI, and scientific discovery

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04464 2026-07-16 cs.LG cs.AI 版本更新 94%

Operator-on-F complements value-equivalence: a planning-time diagnostic for latent world models

算子对F补充值等价性:潜在世界模型的规划时诊断

Donna Vakalis

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 研究针对基于模型的强化学习中世界模型评估问题,引入算子对F诊断,通过模型自身预测器比较模型与环境的k步潜在前推,揭示其与规划回报的强关联及在跨架构比较中的作用,补充而非替代值等价性诊断。

Comments Accepted at RLC 2026 WM Workshop. V2 places the diagnostic in Koopman representation theory; references expanded. Results unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00412 2026-07-03 cs.AI cs.RO 版本更新 94%

Physically Native World Models: A Hamiltonian Perspective on Generative World Modeling

物理原生世界模型:生成式世界建模的哈密顿视角

Sen Cui, Jingheng Ma

机构 * Tsinghua University(清华大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出哈密顿世界模型,通过结构化潜相空间和哈密顿动力学演化实现物理可靠、动作可控且长期稳定的未来预测,用于具身决策。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22281 2026-06-17 cs.CV cs.AI cs.CL cs.LG cs.RO 版本更新 94%

ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model

ThinkJEPA:赋予潜在世界模型大型视觉-语言推理能力

Haichao Zhang, Yijiang Li, Shwai He, Tushar Nagarajan, Mingfei Chen, Jianglin Lu, Ang Li, Yun Fu

机构 * Northeastern University(东北大学) University of California San Diego(加州大学圣地亚哥分校) University of Maryland(马里兰大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Washington(华盛顿大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出ThinkJEPA框架,结合密集JEPA分支与稀疏VLM思考者分支,通过分层金字塔表示提取模块,实现细粒度运动建模与长程语义引导,在手部操作轨迹预测任务上超越基线。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11395 2026-06-12 cs.LG cs.AI 版本更新 94%

ARROW: Augmented Replay for RObust World models

ARROW:增强重放用于鲁棒世界模型

Abdulaziz Alyahya, Abdallah Al Siyabi, Markus R. Ernst, Luke Yang, Levin Kuhlmann, Gideon Kowadlo

机构 * Imam Mohammad Ibn Saud Islamic University (IMSIU)(伊玛姆·穆罕默德·本·沙特伊斯兰大学) Monash University(莫纳什大学) University of New South Wales, Sydney(新南威尔士大学,悉尼) Cerenaut

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出ARROW算法,一种基于模型的持续强化学习方法,通过高效的重放缓冲区减少灾难性遗忘,提升在无共享结构任务和有共享结构任务中的表现。

Comments 36 pages and 11 figures (includes Appendix)

Journal ref Transactions on Machine Learning Research, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02800 2026-06-24 cs.CV cs.AI cs.LG cs.MM cs.RO 版本更新 94%

Cosmos 3: Omnimodal World Models for Physical AI

Cosmos 3:面向物理AI的全模态世界模型

NVIDIA, :, Aditi, Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, Aarti Basant, Mukesh Beladiya, Mohammad Qazim Bhat, Zaid Pervaiz Bhat, Dan Blick, Vanni Brighella, Han Cai, Tiffany Cai, Eric Cameracci, Jiaxin Cao, Yulong Cao, Mark Carlson, Carlos Casanova, Ting-Yun Chang, Yan Chang, Yu-Wei Chao, Prithvijit Chattopadhyay, Roshan Chaudhari, Chieh-Yun Chen, Junyu Chen, Ke Chen, Qizhi Chen, Wenkai Chen, Xiaotong Chen, Yu Chen, An-Chieh Cheng, Click Cheng, Xiu Chia, Jeana Choi, Chaeyeon Chung, Wenyan Cong, Yin Cui, Magdalena Dadela, Nalin Dadhich, Wenliang Dai, Joyjit Daw, Alperen Degirmenci, Rodrigo Vieira Del Monte, Robert Denomme, Sameer Dharur, Marco Di Lucca, Ke Ding, Wenhao Ding, Yifan Ding, Yuzhu Dong, Nicole Drumheller, Yilun Du, Aigul Dzhumamuratova, Aleksandr Efitorov, Hamid Eghbalzadeh, Naomi Eigbe, Imad El Hanafi, Hassan Eslami, Benedikt Falk, Jiaojiao Fan, Jim Fan, Amol Fasale, Sergiy Fefilatyev, Liang Feng, Francesco Ferroni, Sanja Fidler, Xiao Fu, Vikram Fugro, Prashant Gaikwad, TJ Galda, Katelyn Gao, Yihuai Gao, Wenhang Ge, Sreyan Ghosh, Arushi Goel, Vivek Goel, Akash Gokul, Rama Govindaraju, Jinwei Gu, Miguel Guerrero, Elfie Guo, Aryaman Gupta, Siddharth Gururani, Hugo Hadfield, Song Han, Ankur Handa, Zekun Hao, Mohammad Harrim, Ali Hassani, Nathan Hayes-Roth, Yufan He, Chris Helvig, Cyrus Hogg, Madison Huang, Michael Huang, Sophia Huang, Yufan Huang, Jacob Huffman, DeLesley Hutchins, Suneel Indupuru, Boris Ivanovic, Arihant Jain, Joel Jang, Ryan Ji, Yanan Jian, Dongfu Jiang, Jingyi Jin, Atharva Joshi, Nikhilesh Joshi, Pranjali Joshi, Andy Ju, Jaehun Jung, Weiwei Kang, Scott Kassekert, Jan Kautz, Ashna Khetan, Julia Kiczka, Slawek Kierat, Gwanghyun Kim, Kuno Kim, Sunny Kim, Kezhi Kong, Xin Kong, Zhifeng Kong, Tomasz Kornuta, Egor Krivov, Hui Kuang, Saurav Kumar, Chia-Wen Kuo, George Kurian, Wojciech Kutak, JF Lafleche, Himangshu Lahkar, Omar Laymoun, Jayjun Lee, Sanggil Lee, Gabriele Leone, Boyi Li, Freya Li, Jiajun Li, Jinfeng Li, Ling Li, Pengcheng Li, Shangru Li, Tingle Li, Xiaolong Li, Xuan Li, Zhaoshuo Li, Zhiqi Li, Hao Liang, Maosheng Liao, Chen-Hsuan Lin, Tsung-Yi Lin, Ming-Yu Liu, Sifei Liu, Zihan Liu, Hai Loc Lu, Xiangyu Lu, Alice Luo, Ruipu Luo, Wenjie Luo, Jiangran Lyu, Martin Ding Ma, Nic Ma, Qianli Ma, Dawid Majchrowski, Louis Marcoux, Miguel Martin, Qing Miao, Ashkan Mirzaei, Shreyas Misra, Kaichun Mo, Durra Mohsin, Hyejin Moon, Pawel Morkisz, Saeid Motiian, Kirill Motkov, Seungjun Nah, Yashraj Narang, Deepak Narayanan, Thabang Ngazimbi, Julian Ouyang, Shubham Pachori, David Page, Yatian Pang, Sehwi Park, Mahesh Patekar, Mostofa Patwary, Marco Pavone, Trung Pham, Wei Ping, Soha Pouya, Shrimai Prabhumoye, Varun Praveen, Delin Qu, Hesam Rabeti, Morteza Ramezanali, Marilyn Reeb, Xuanchi Ren, Kristen Rumley, Wojciech Rymer, Jun Saito, Yeongho Seol, John Shao, Piyush Shekdar, Tianwei Shen, Humphrey Shi, Min Shi, Stella Shi, Kevin Shih, Mohammad Shoeybi, Mateusz Sieniawski, Shuran Song, Alexander Sotelo, Amir Sotoodeh, Sunil Srinivasa, Vignesh Srinivasakumar, Bartosz Stefaniak, Rahul Heinrich Steiger, Shangkun Sun, Jiaxiang Tang, Shitao Tang, Yangyang Tang, Yue Tang, Tolou Tavakkoli, Kayley Ting, Krzysztof Tomala, Wei-Cheng Tseng, Jibin Varghese, Sergei Vasilev, Thomas Volk, Raju Wagwani, Roger Waleffe, Andrew Z. Wang, Boxiang Wang, Haoxiang Wang, Qiao Wang, Shihao Wang, Shijie Wang, Ting-Chun Wang, Yan Wang, Yu Wang, Rohit Watve, David Wehr, Fangyin Wei, Xinshuo Weng, Jay Zhangjie Wu, Kedi Wu, Hongchi Xia, Summer Xiao, Tianjun Xiao, Kevin Xie, Daguang Xu, Jiashu Xu, Mengyao Xu, Ruqing Xu, Xingqian Xu, Yao Xu, Dinghao Yang, Dong Yang, Hans Yang, Xiaodong Yang, Xuning Yang, Yichu Yang, Yurong You, Zhiding Yu, Hao Yuan, Simon Yuen, Xiaohui Zeng, Pengcuo Zeren, Cindy Zha, Haotian Zhang, Jenny Zhang, Jing Zhang, Liangkai Zhang, Paris Zhang, Shun Zhang, Xuanmeng Zhang, Zhizheng Zhang, Ann Zhao, Yilin Zhao, Yuliya Zhautouskaya, Charles Zhou, Fengzhe Zhou, Shilin Zhu, Yuke Zhu, Dima Zhylko, Artur Zolkowski

机构 * NVIDIA

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出基于统一混合Transformer架构的全模态世界模型Cosmos 3,联合处理语言、图像、视频、音频和动作序列,在理解和生成任务上达到新最优,为具身智能体提供可扩展的通用骨干。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16663 2026-07-07 cs.RO cs.CV cs.LG cs.SY eess.SY 版本更新 94%

Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models

使用潜在空间生成世界模型减轻自动驾驶模仿学习中的协变量转移

Alexander Popov, Alperen Degirmenci, David Wehr, Shashank Hegde, Ryan Oldja, Alexey Kamenev, Bertrand Douillard, David Nistér, Urs Muller, Ruchi Bhargava, Stan Birchfield, Nikolai Smolyanskiy

机构 * NVIDIA

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出用潜在空间生成世界模型解决自动驾驶协变量转移问题,训练时利用世界模型减轻该问题,无需大量训练数据,还引入新感知编码器,实验显示相比之前有显著改进。

Comments 8 pages, 6 figures, original September 2024, accepted at ICRA 2025 Workshop "Robots in the Wild", for associated video file, see https://youtu.be/7m3bXzlVQvU

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16592 2026-06-16 cs.RO cs.AI cs.CV cs.ET 版本更新 94%

Human Cognition in Machines: A Unified Perspective of World Models

机器中的人类认知:世界模型的统一视角

Timothy Rupprecht, Pu Zhao, Amir Taherin, Arash Akbari, Arman Akbari, Yumei He, Tooba Imtiaz, Sean Duffy, Juyi Lin, Yixiao Chen, Rahul Chowdhury, Enfu Nan, Yixin Shen, Yifan Cao, Haochen Zeng, Weiwei Chen, Geng Yuan, Jennifer Dy, Sarah Ostadabbas, Xuan Zhang, David Kaeli, Edmund Yeh, Yanzhi Wang

机构 * Northeastern University(东北大学) EmbodyX Inc.(EmbodyX公司) Tulane University(路易斯安那州立大学) Cornell University(康奈尔大学) University of Georgia(佐治亚大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出统一框架整合记忆、感知等认知功能,指出动机和元认知研究不足,并引入认知世界模型新类别。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08732 2026-06-08 cs.RO cs.LG 版本更新 94%

Latent Geometry Beyond Search: Amortizing Planning in World Models

超越搜索的潜在几何:在世界模型中摊销规划

Hoang Nguyen, Xiaohao Xu, Xiaonan Huang

机构 * Department of Robotics, University of Michigan, Ann Arbor(密歇根大学机器人系,安阿伯)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出在正则化潜在几何下,将规划摊销为潜在逆动力学映射,以轻量级GC-IDM替代在线搜索,在七个环境协议中匹配或超越CEM,决策成本降低100-130倍。

Comments 31 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15156 2026-08-20 cs.RO cs.AI 版本更新 94%

Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models

用于学习世界模型中反事实滚动的低秩动力学有效潜在载体

Yang Liu, Yuming Chen

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文针对学习世界模型的反事实滚动问题,提出基于低秩动力学有效潜在载体的方法,在两物体碰撞环境中验证秩4修补块可实现稳定的12步自主反事实滚动,且效果可复现。

Comments Revised version: removed an inconclusive development-only event-relative phase analysis; the main rank-4 carrier, fresh-checkpoint replication, B1/B2 temporal reuse, position-edit, and joint-edit conclusions are unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏