arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 4308 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 4308 篇

2605.09131 2026-05-12 cs.AI cs.MA 90%

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments

MCP-Cosmos:用于MCP环境复杂任务执行的世界模型增强型智能体

Giridhar Ganapavarapu, Dhaval Patel

机构 * IBM

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 MCP-Cosmos融合世界模型与MCP框架,通过引入生成世界模型提升智能体在复杂任务中的执行能力,实验显示其在工具成功率和参数准确性等方面有显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11751 2026-04-14 cs.RO cs.AI 90%

Grounded World Model for Semantically Generalizable Planning

基于语义泛化的世界模型用于规划

Quanyi Li, Lan Feng, Haonan Zhang, Wuyang Li, Letian Wang, Alexandre Alahi, Harold Soh

机构 * Independent(独立) EPFL(瑞士联邦理工学院洛桑) Beihang University(北京航空航天大学) University of Toronto(多伦多大学) NUS(新加坡国立大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出基于视觉-语言对齐潜在空间的世界模型,用于改进视觉-运动模型控制,实现更广泛的语义泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03353 2026-04-13 cs.MA cs.AI 90%

Decentralized Collective World Model for Emergent Communication and Coordination

去中心化的集体世界模型用于涌现通信与协调

Kentaro Nomura, Tatsuya Aoki, Tadahiro Taniguchi, Takato Horii

机构 * The University of Osaka(大阪大学) Kyoto University(京都大学) Ritsumeikan University(立命馆大学) The University of Tokyo(东京大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出一种去中心化的多智能体世界模型,通过时间扩展的集体预测编码实现通信与协调的同步发展,通过对比学习促进信息对齐,展示了在双智能体轨迹绘制任务中优于非通信模型的协调能力。

Comments Accepted at IEEE ICDL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25981 2026-03-30 cs.RO cs.AI cs.CL 90%

Policy-Guided World Model Planning for Language-Conditioned Visual Navigation

基于策略引导的世界模型规划用于语言条件的视觉导航

Amirhosein Chahe, Lifeng Zhou

机构 * Drexel University(德雷塞尔大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出PiJEPA框架,结合学习导航策略与潜在世界模型规划,通过微调Octo策略和冻结预训练视觉编码器,提升语言指导下的视觉导航准确性与指令遵循性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19675 2026-03-23 cs.CV cs.RO 90%

DynFlowDrive: Flow-Based Dynamic World Modeling for Autonomous Driving

DynFlowDrive: 基于流的动态世界建模用于自动驾驶

Xiaolu Liu, Yicong Li, Song Wang, Junbo Chen, Angela Yao, Jianke Zhu

机构 * Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出DynFlowDrive,通过流基于的动力学建模提升自动驾驶中的场景预测可靠性,引入稳定性感知的多模式轨迹选择策略,实验证明在nuScenes和NavSim基准上效果显著。

Comments 18 pages, 6 figs

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17652 2026-03-19 cs.RO cs.CV 90%

VectorWorld: Efficient Streaming World Model via Diffusion Flow on Vector Graphs

VectorWorld: 通过向量图上的扩散流实现高效的流式世界模型

Chaokang Jiang, Desen Zhou, Jiuming Liu, Kevin Li Sun

机构 * University of Cambridge, Cambridge, United Kingdom(剑桥大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 VectorWorld通过向量图上的扩散流实现高效的流式世界模型,解决了自动驾驶政策闭环评估中的初始化不匹配、采样延迟和运动可行性问题,提升了地图结构精度和闭环运行稳定性。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14497 2026-03-18 cs.CV cs.RO 90%

WorldVLM: Combining World Model Forecasting and Vision-Language Reasoning

WorldVLM:结合世界模型预测与视觉语言推理

Stefan Englmeier, Katharina Winter, Fabian B. Flohr

机构 * Munich University of Applied Sciences(慕尼黑应用科学大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 WorldVLM结合视觉语言模型与世界模型,通过统一架构提升自动驾驶中的环境预测与决策能力,解决空间理解受限问题。

Comments 8 pages, 6 figures, 5 tables; submitted to IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13615 2026-03-17 cs.CV cs.RO 90%

Egocentric World Model for Photorealistic Hand-Object Interaction Synthesis

第一人称世界模型用于逼真手-物体交互合成

Dayou Li, Lulin Liu, Bangya Liu, Shijie Zhou, Jiu Feng, Ziqi Lu, Minghui Zheng, Chenyu You, Zhiwen Fan

机构 * Texas A&M University(德克萨斯A&M大学) University of Minnesota(明尼苏达大学) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) University of California, Los Angeles(加州大学洛杉矶分校) University of Texas at Austin(得克萨斯大学奥斯汀分校) Amazon(亚马逊) State University of New York at Stony Brook(纽约州立大学石溪分校)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出EgoHOI模型,通过动作信号模拟物理一致的交互,克服了高速头部运动、严重遮挡和高自由度手部运动带来的挑战,实验验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10390 2026-02-12 cs.LG cs.AI 90%

Affordances Enable Partial World Modeling with LLMs

affordances 使 LLMs 能实现部分世界建模

Khimya Khetarpal, Gheorghe Comanici, Jonathan Richens, Jeremy Shar, Fei Xia, Laurent Orseau, Aleksandra Faust, Doina Precup

机构 * Google Deepmind(谷歌DeepMind)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出利用 affordances 构建部分世界模型,通过在多任务中引入分布稳健的 affordances,提高搜索效率并提升奖励表现。

Comments 18 pages, 5 figures, 2 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06521 2026-02-09 cs.CV cs.RO 90%

DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving

DriveWorld-VLA: 基于视觉-语言-动作的统一潜在空间世界建模用于自动驾驶

Feiyang jia, Lin Liu, Ziying Song, Caiyan Jia, Hangjun Ye, Xiaoshuai Hao, Long Chen

机构 * School of Computer Science(计算机科学学院) Beijing Key Laboratory of Traffic Data Mining(交通数据挖掘重点实验室) Beijing Jiaotong University(北京交通大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 DriveWorld-VLA通过统一视觉-语言-动作与潜在空间世界建模,提升自动驾驶中的决策与前瞻性想象能力。

Comments 20 pages, 7 tables, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03569 2026-02-04 cs.AI cs.LG 90%

EHRWorld: A Patient-Centric Medical World Model for Long-Horizon Clinical Trajectories

EHRWorld: 一个以患者为中心的医疗世界模型用于长周期临床轨迹

Linjie Mu, Zhongzhen Huang, Yannian Gu, Shengqian Qin, Shaoting Zhang, Xiaofan Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 EHRWorld通过因果序列范式训练,有效解决长周期临床模拟中的误差累积问题,提升医疗世界模型的稳定性与可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02110 2026-02-03 cs.LG cs.CV 90%

An Empirical Study of World Model Quantization

世界模型量化经验研究

Zhongqian Fu, Tianyi Zhao, Kai Han, Hang Zhou, Xinghao Chen, Yunhe Wang

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本研究通过DINO-WM案例系统探讨世界模型量化的影响,发现量化方法对模型稳定性、精度和任务成功率有显著影响,揭示了不同量化策略下的失败模式及部署指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17507 2026-01-27 cs.RO 90%

MetaWorld: Skill Transfer and Composition in a Hierarchical World Model for Grounding High-Level Instructions

MetaWorld: 一个用于地面指令基础的分层世界模型中的技能迁移与组合

Yutong Shen, Hangxu Liu, Kailin Pei, Ruizhe Xia, Tongtong Feng

机构 * Beijing University of Technology(北京理工大学) Fudan University(复旦大学) Tsinghua University(清华大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);latent dynamics(abstract);model-based RL(abstract)

AI总结 MetaWorld通过整合语义规划与物理控制,利用专家策略迁移提升人形机器人在定位-操作任务中的性能。

Comments 8 pages, 4 figures, Submitted to ICLR 2026 World Model Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03905 2026-01-09 cs.AI cs.CL cs.LG 90%

Current Agents Fail to Leverage World Model as Tool for Foresight

当前智能体无法利用世界模型作为前瞻性工具

Cheng Qian, Emre Can Acikgoz, Bingxuan Li, Xiusi Chen, Yuji Zhang, Bingxiang He, Qinyu Luo, Dilek Hakkani-Tür, Gokhan Tur, Yunzhu Li, Heng Ji

机构 * UIUC(伊利诺伊大学香槟分校) THU(清华大学) JHU(约翰霍普金斯大学) Columbia(哥伦比亚大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文研究了当前智能体能否利用世界模型作为工具提升前瞻性认知,发现其在模拟调用和结果整合方面存在显著不足。

Comments 36 Pages, 13 Figures, 17 Tables (Meta data updated)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21887 2026-01-06 cs.RO cs.AI 90%

Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space

空域世界模型用于3D空间中的长视距视觉生成与导航

Weichen Zhang, Peizhi Tang, Xin Zeng, Fanhang Man, Shiquan Yu, Zichao Dai, Baining Zhao, Hongjin Chen, Yu Shang, Wei Wu, Chen Gao, Xinlei Chen, Xin Wang, Yong Li, Wenwu Zhu

机构 * Tsinghua University(清华大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 ANWM是一种空域导航世界模型,通过预测未来视觉观察来提升无人机在大规模3D空间中的长视距导航性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23541 2025-12-30 cs.RO cs.AI 90%

Act2Goal: From World Model To General Goal-conditioned Policy

Act2Goal:从世界模型到通用目标条件政策

Pengfei Zhou, Liliang Chen, Shengcong Chen, Di Chen, Wenzhi Zhao, Rongjun Jin, Guanghui Ren, Jianlan Luo

机构 * Agibot Research(阿吉博特研究所)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 Act2Goal通过结合视觉世界模型与多尺度时间控制,实现目标条件政策的鲁棒长时间尺度操作,提升机器人在新环境中的自主适应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19133 2025-12-23 cs.RO cs.CV 90%

WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving

WorldRFT: 通过强化微调的潜在世界模型进行自动驾驶的规划

Pengxuan Yang, Ben Lu, Zhongpu Xia, Chao Han, Yinfeng Gao, Teng Zhang, Kun Zhan, XianPeng Lang, Yupeng Zheng, Qichao Zhang

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 WorldRFT通过强化学习微调提升自动驾驶规划性能,实现安全性和效率的双重优化。

Comments AAAI 2026, first version

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08478 2025-12-10 cs.CV cs.AI cs.GR 90%

Visionary: The World Model Carrier Built on WebGPU-Powered Gaussian Splatting Platform

Visionary: 基于WebGPU驱动的高斯点云平台的世界模型载体

Yuning Gong, Yifei Liu, Yifan Zhan, Muyao Niu, Xueying Li, Yuanjun Liao, Jiaming Chen, Yuanyuan Gao, Jiaqi Chen, Minming Chen, Li Zhou, Yuning Zhang, Wei Wang, Xiaoqing Hou, Huaxi Huang, Shixiang Tang, Le Ma, Dingwen Zhang, Xue Yang, Junchi Yan, Yanchi Zhang, Yinqiang Zheng, Xiao Sun, Zhihang Zhong

机构 * Shanghai AI Laboratory(上海人工智能实验室) Sichuan University(四川大学) The University of Tokyo(东京大学) Shanghai Jiao Tong University(上海交通大学) Northwestern Polytechnical University(西北工业大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 Visionary是一款基于WebGPU的实时高斯点云渲染平台,通过高效推断和浏览器集成,提升3DGS方法的部署效率与多样性。

Comments Project page: https://visionary-laboratory.github.io/visionary

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19430 2025-12-05 cs.RO cs.CV 90%

GigaBrain-0: A World Model-Powered Vision-Language-Action Model

GigaBrain-0:一种由世界模型驱动的视觉-语言-动作模型

GigaBrain Team, Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang, Guosheng Zhao, Haoyun Li, Jie Li, Jiagang Zhu, Lv Feng, Peng Li, Qiuping Deng, Runqi Ouyang, Wenkang Qin, Xinze Chen, Xiaofeng Wang, Yang Wang, Yifan Li, Yilong Li, Yiran Ding, Yuan Xu, Yun Ye, Yukun Zhou, Zhehao Dong, Zhenan Wang, Zhichao Liu, Zheng Zhu

机构 * GigaBrain Team(GigaBrain团队)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 GigaBrain-0通过世界模型生成数据,减少对真实机器人数据的依赖,提升视觉-语言-动作模型的泛化能力和现实世界性能。

Comments https://gigabrain0.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03692 2025-12-03 cs.CV cs.RO 90%

LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences

LiDARCrafter:从LiDAR序列动态构建4D世界模型

Ao Liang, Youquan Liu, Yu Yang, Dongyue Lu, Linfeng Li, Lingdong Kong, Huaici Zhao, Wei Tsang Ooi

机构 * WorldBench Team(WorldBench团队)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 LiDARCrafter通过自然语言输入生成和编辑4D LiDAR数据,实现高保真度、可控性和时间一致性,为自动驾驶数据增强和模拟提供新方法。

Comments AAAI 2026 Oral Presentation; 38 pages, 18 figures, 12 tables; Project Page at https://lidarcrafter.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01119 2025-12-02 cs.LG cs.AI 90%

World Model Robustness via Surprise Recognition

通过惊喜识别提升世界模型鲁棒性

Geigh Zollicoffer, Tanush Chopra, Mingkuan Yan, Xiaoxu Ma, Kenneth Eaton, Mark Riedl

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 通过惊喜识别提升世界模型在噪声环境下的鲁棒性,增强自动驾驶模拟中智能体的稳定性与性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20156 2025-11-26 cs.CV cs.RO 90%

Map-World: Masked Action planning and Path-Integral World Model for Autonomous Driving

Map-World: 遮蔽动作规划与路径积分世界模型用于自动驾驶

Bin Hu, Zijian Lu, Haicheng Liao, Chengran Yuan, Bin Rao, Yongkang Li, Guofa Li, Zhiyong Cui, Cheng-zhong Xu, Zhenning Li

机构 * University of Macau(澳门大学) National University of Singapore(新加坡国立大学) Purdue University(普渡大学) Chongqing University(重庆大学) Beihang University(北航大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 MAP-World通过结合遮蔽动作规划和路径加权世界模型,实现了高效的多模式自动驾驶规划,无需强化学习即可达到最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10819 2025-11-21 cs.AI cs.LG 90%

PoE-World: Compositional World Modeling with Products of Programmatic Experts

PoE-World: 通过程序专家的乘积进行组合世界建模

Wasu Top Piriyakulkij, Yichao Liang, Hao Tang, Adrian Weller, Marta Kryven, Kevin Ellis

机构 * Cornell University(康奈尔大学) University of Cambridge(剑桥大学) The Alan Turing Institute(艾伦·图灵研究所) Dalhousie University(达尔豪斯大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 PoE-World通过程序专家的乘积有效建模复杂非网格世界领域,利用LLMs学习世界模型并实现高效泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17482 2025-11-18 cs.CV cs.AI 90%

SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries

Chenxu Dang, Haiyan Liu, Jason Bao, Pei An, Xinyue Tang, PanAn, Jie Ma, Bingchuan Sun, Yan Wang

机构 * AIR

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

Comments Accepted by AAAI2026 Code: https://github.com/MSunDYY/SparseWorld

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27002 2025-11-03 cs.LG cs.AI 90%

Jasmine: A Simple, Performant and Scalable JAX-based World Modeling Codebase

Mihir Mahajan, Alfred Nguyen, Franz Srambical, Stefan Bauer

机构 * p(doom)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

Comments Blog post: https://pdoom.org/jasmine.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07591 2025-10-24 cs.LG cs.AI 90%

DMWM: Dual-Mind World Model with Long-Term Imagination

Lingyi Wang, Rashed Shelim, Walid Saad, Naren Ramakrishnan

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02722 2025-09-09 cs.AI 90%

Planning with Reasoning using Vision Language World Model

Delong Chen, Theo Moutakanni, Willy Chung, Yejin Bang, Ziwei Ji, Allen Bolourchi, Pascale Fung

机构 * Meta FAIR

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23126 2025-08-27 cs.RO 90%

ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation

Suning Huang, Qianzhong Chen, Xiaohan Zhang, Jiankai Sun, Mac Schwager

机构 * Stanford University(斯坦福大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11762 2025-08-21 cs.LG cs.RO 90%

MUVO: A Multimodal Generative World Model for Autonomous Driving with Geometric Representations

Daniel Bogdoll, Yitian Yang, Tim Joseph, Melih Yazgan, J. Marius Zöllner

机构 * FZI Research Center for Information Technology, Germany(德国弗赖堡信息科技研究中心) Karlsruhe Institute of Technology, Germany(德国卡尔斯鲁厄理工学院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

Comments Daniel Bogdoll and Yitian Yang contributed equally. Accepted for publication at IV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06096 2025-08-11 cs.RO cs.AI 90%

Bounding Distributional Shifts in World Modeling through Novelty Detection

Eric Jing, Abdeslam Boularias

机构 * Department of Computer Science, Rutgers University(计算机科学系,罗格斯大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

Comments 7 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏