arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-07-08 至 2026-07-08 共收录 13 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 9 篇

2606.19990 2026-07-08 cs.AI 新提交 96%

Reward as An Agent for Embodied World Models

奖励作为具身世界模型的智能体

Pu Li, Zhigang Lin, Qiang Wu, Yongxuan Lv, Fei Wang, Shan You

机构 * ACE Robotics(ACE机器人)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);embodied world model(title,abstract);world model(title,abstract)

AI总结 提出奖励智能体框架和动态感知 rollout 多样化方法,通过鲁棒验证支持更广泛探索,缓解奖励黑客问题,提升世界模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06401 2026-07-08 cs.AI 新提交 94%

A Definition and Roadmap for World Models

世界模型的定义与路线图

Xinyuan Chen, Haoyu Guo, Shi Guo, Bingqi Jiang, Chunhua Shen, Xing Shen, Tianfan Xue, Yufei Xue, Mulin Yu, Weinan Zhang, Bin Zhao, Bowen Zhou, Ming Zhou

机构 * Shanghai AI Laboratory(上海人工智能实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文针对人工智能领域对世界模型定义、预测内容及构建方式缺乏共识的问题,给出科学定义,讨论关键技术,提供分阶段路线图,以助力有效世界模型的开发。

Comments Technical report, 58 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05352 2026-07-08 cs.CV cs.AI cs.LG 新提交 94%

Multiplayer Interactive World Models with Representation Autoencoders

基于表示自编码器的多智能体交互世界模型

Anthony Hu, Václav Volhejn, Adrien Ramanana Rahary, Chris Mulder, Aditya Makkar, Alyx Liao, Amélie Royer, Manu Orsini, Adam Jelley, Eloi Alonso, Florian Laurent, Fredrik Norén, James Swingos, Jan Hünermann, Kent Rollins, Lucas Hosseini, Matthieu Le Cauchois, Maxim Peter, Pim de Witte, Tim Brown, Vincent Micheli, Moritz Böhle, Gabriel de Marmiesse, Viktoriia Sharmanska, Lucia Specia, Michael Black, Patrick Pérez

机构 * Epic Games École nationale des ponts et chaussées(法国国家桥梁与道路学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 针对复杂物理交互的高动态环境,该研究提出首个多智能体世界模型,以多智能体动作流为条件训练隐扩散模型,可实时生成稳定长时四智能体对局,开源相关资源。

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28757 2026-07-08 cs.CV cs.RO 版本更新 94%

A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models

基于物理的多智能体动力学世界模型基准

Nuo Chen, Lulin Liu, Zihao Li, Ziyao Zeng, Zihao Zhu, Wenyan Cong, Junyuan Hong, Yunhao Yang, Zhengzhong Tu, Yan Wang, Boris Ivanovic, Marco Pavone, Zhangyang Wang, Yang Zhou, Zhiwen Fan

机构 * Texas A&M University(德克萨斯大学) University of Minnesota(明尼苏达大学) Marquette University(马奎特大学) Yale University(耶鲁大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) Massachusetts General Hospital(麻省总医院) Harvard Medical School(哈佛医学院) NVIDIA(英伟达) Stanford University(斯坦福大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出CrashTwin框架,通过多智能体碰撞场景数据集和校准无关重建流程,从时空一致性、动量与动能守恒、世界动力学完整性三个维度评估世界模型的物理可信度。

Comments 34 pages, 9 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05966 2026-07-08 cs.RO 新提交 91%

Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model Failure

想象的展开是运动学的,而非动力学的:对长期世界模型失败的诊断

Finn Rasmus Schäfer, Korbinian Moller, Yuan Gao, Christian Oefinger, Sebastian Schmidt, Johannes Betz

机构 * Autonomous Vehicle Systems Lab(自动驾驶车辆系统实验室) Technical University of Munich(慕尼黑技术大学) Data Analytics and Machine Learning Group(数据分析与机器学习小组)

专题命中 通用世界模型 :world-model(title);world-model(title);world model(abstract,comments);world models(abstract)

AI总结 研究针对世界模型长期失败问题,提出用运动学与动力学重新构建的方法,通过想象的运动学一致性误差(iKCE)及扰动协议进行诊断,在DreamerV3检查点实例化,区分了运动学与动力学想象,揭示模型长期失败特征。

Comments 9 Pages Workshop Paper accepted at RSS Robot World Model Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05577 2026-07-08 cs.AI cs.CL cs.IR 新提交 88%

Narrative World Model: Narratology-Grounded Writer Memory for Long-Form Fiction

叙事世界模型:基于叙事学的长篇小说作者记忆

Mohammad Saifullah, Thomas Kornmaier, Taaha Kazi, Vasu Sharma, Aditya Sanjiv Kanade, Aanand Kumar Yadav

机构 * PocketFM(口袋FM)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI

AI总结 研究针对长篇小说作者记忆问题,提出叙事世界模型(NWM),结合叙事学时间状态图与查询条件混合检索,在多跳叙事学问答上显著优于基线,优势源于其独特表征,非提取方式。

Comments 23 pages, 4 figures; 9-page main text plus appendix. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06291 2026-07-08 cs.CV cs.HC 新提交 86%

AlayaWorld: Long-Horizon and Playable Video World Generation

AlayaWorld:长期可玩视频世界生成

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao

机构 * Alaya Lab(阿莱亚实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 研究旨在解决游戏世界构建难题,核心方法是提出AlayaWorld全栈开源框架,可实现开放式实时交互,统一开发流程,主要贡献是为生成世界模型研究和应用奠定实践基础。

Comments Authors are listed alphabetically by the first name and their role. See the contribution section for details

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.02965 2026-07-08 cs.LG cs.SY eess.SP eess.SY stat.ML 版本更新 81%

Joint Energy Management and Coordinated AIGC Workload Scheduling for Distributed Data Centers: A Diffusion-Aided Reward Shaping Approach

分布式数据中心的联合能源管理与协同AIGC工作负载调度:一种扩散辅助奖励塑造方法

Yang Fu, Peng Qin, Liming Chen, Zihao Zhang, Hao Yu, Yifei Wang

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 针对AIGC在数据中心带来的能源管理与工作负载调度挑战,提出联合调度框架,考虑多种能源资源,制定系统效用最大化问题。因奖励稀疏限制DRL算法,开发扩散模型辅助奖励塑造方法,实验表明该方案能有效适应波动和异质性,实现高效学习收敛和系统效用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23212 2026-07-08 cond-mat.dis-nn cond-mat.stat-mech nlin.AO 版本更新 65%

Replication and Information Extraction in a Minimal Agent-Environment Model

最小化智能体 - 环境模型中的复制与信息提取

Sebastiano Ariosto, Jerome Garnier-Brun, Luca Saglietti, Davide Straziota

专题命中 通用世界模型 :environment model(title)

AI总结 研究在无明确指导或奖励下从数据提取信息的问题,通过对分类规则施加简单性偏差,利用统计力学工具刻画自发学习相变,扩展到智能体群体后发现交互可重塑学习相边界,开辟仅通过标签交换的分散学习途径。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身与机器人 4 篇

2607.06559 2026-07-08 cs.RO 新提交 95%

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

RynnWorld-4D:用于机器人操作的4D具身世界模型

Haoyu Zhao, Xingyue Zhao, Siteng Huang, Xin Li, Deli Zhao, Zhongyu Li

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hong Kong Embodied AI Lab(香港具身人工智能实验室) CUHK(香港中文大学) Hupan Lab(湖畔实验室)

专题命中 具身与机器人 :world model(title,abstract);world models(title);embodied world model(title);world model(title,abstract)

AI总结 研究针对开放世界机器人操作,提出RynnWorld-4D生成模型,通过RGB-DF捕捉4D动态,其具三分支架构,还整理数据集并提出RynnWorld-4D-Policy,实验证明该模型在时空预测及实际操作任务中表现出色。

Comments Project Page: https://alibaba-damo-academy.github.io/RynnWorld-4D.github.io, Github: https://github.com/alibaba-damo-academy/RynnWorld-4D

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06501 2026-07-08 cs.RO 新提交 81%

Hypothesis-driven Model Expansion under Uncertainty for Open-World Robot Planning

不确定性下用于开放世界机器人规划的假设驱动模型扩展

Anxing Xiao, Hanbo Zhang, Tianrun Hu, David Hsu

机构 * School of Computing, National University of Singapore(新加坡国立大学计算机学院) Smart Systems Institute, National University of Singapore(新加坡国立大学智能系统研究所)

专题命中 具身与机器人 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 针对开放世界机器人规划中传统方法失效的问题,提出假设驱动模型扩展框架,利用基础模型生成假设,结合自动规划与验证反馈进行迭代优化,实现自主知识扩展,推动家用服务机器人实际部署。

Comments Accepted to Robotics: Science and Systems (RSS) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17649 2026-07-08 cs.CV cs.AI cs.RO 版本更新 73%

SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios

SWITCH:在长时程具身场景中评估和处理具象接口的基准测试

Juntao Cheng, Wanyue Zhang, Zhiwei Yu, Shuo Ren, Zheqi He, Shaoxuan Xie, Guocai Yao, Jieru Lin, Börje F. Karlsson, Jiajun Zhang

专题命中 具身与机器人 :world model(abstract);world model(abstract);分类 cs.AI、cs.CV、cs.RO

AI总结 SWITCH基准测试通过1170个时间交互视频,评估具象接口在真实第一人称环境中的闭环交互推理,揭示多模态模型在细粒度视觉-时间感知、结果验证和错误恢复方面的不足。

Comments The dataset is available at https://huggingface.co/datasets/BAAI-Agents/SWITCH

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20182 2026-07-08 cs.RO cs.MA 版本更新 71%

IndoorR2X: Indoor Robot-to-Everything Coordination with LLM-Driven Planning

IndoorR2X: 基于大语言模型的室内机器人与万物协调

Fan Yang, Soumya Teotia, Shaunak A. Mehta, Prajit KrisshnaKumar, Quanting Xie, Jun Liu, Yueqi Song, Wenkai Li, Atsunori Moteki, Kanji Uchino, Yonatan Bisk

机构 * Fujitsu Research(富士通研究所) Carnegie Mellon University(卡内基梅隆大学)

专题命中 具身与机器人 :world model(abstract);world model(abstract);分类 cs.RO、cs.MA

AI总结 IndoorR2X通过整合移动机器人和静态物联网设备的观测数据,构建全局语义状态,实现可扩展的场景理解、减少冗余探索并利用大语言模型进行高层协调,提升多机器人系统的效率和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏