arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 6450 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 4308 篇

2602.07030 2026-02-10 cs.LG 91%

Neural Sabermetrics with World Model: Play-by-play Predictive Modeling with Large Language Model

基于世界模型的神经棒球统计学:利用大语言模型进行逐局预测建模

Young Jin Ahn, Yiyang Du, Zheyuan Zhang, Haisen Kang

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出基于世界模型的神经棒球统计学,利用大语言模型预测棒球比赛发展,实验证明其在预测投球和挥棒决策上的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18234 2026-08-20 cs.RO cs.AI cs.LG 新提交 91%

GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

GigaBrain-WBC-0.5:一种用于与环境交互的鲁棒全身控制的行为世界模型

Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI、cs.LG、cs.RO

AI总结 该研究提出首个用于人形机器人全身控制的行为世界模型GigaBrain-WBC-0.5,通过训练因果Transformer联合预测动作、状态与指令分布,在多场景下实现鲁棒控制,成功率优于现有基线。

Comments 20 pages, 8 figures, 4 tables. Technical report. Project page: https://shepherd1226.github.io/gigabrain-wbc-0.5/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08600 2026-08-18 cs.CV cs.AI cs.LG 版本更新 91%

Population-Scalable Multi-Agent World Modeling

支持种群规模扩展的多智能体世界建模

Renjie Zhao, Yuxiang Wu, Mingyu Zhang, Jiaxin Li, Sisi Li, Yimin Sheng, Tianxi Tan, Zhenkai Zhang, Jianyi Zhu, Yong-Lu Li

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 针对多智能体世界模型的种群可扩展性问题,提出无需重训练即可扩展至任意智能体数量的Khora模型,通过解耦世界状态演化与渲染实现跨视图一致性,验证了泛化性并构建了实时交互系统。

Comments Technical report. Project page: https://rhos.ai/research/khora. Online demo: https://ophilus.ai/khora

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15065 2026-08-05 cs.RO cs.CV cs.LG 版本更新 91%

DriftWorld: Fast World Modeling through Drifting

DriftWorld:通过漂移实现快速世界建模

Susie Lu, Haonan Chen, Weirui Ye, Yilun Du

机构 * Massachusetts Institute of Technology(麻省理工学院) Harvard University(哈佛大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 研究针对预测性世界模型生成展开慢的问题,提出基于漂移生成模型的DriftWorld,训练时学习动作条件漂移以快速生成未来帧,在机器人操作基准测试中实现快速准确决策,还能作离线模拟器,性能优于基于扩散的基线。

Comments Website at https://susie-lu.github.io/driftworld/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05966 2026-07-08 cs.RO 新提交 91%

Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model Failure

想象的展开是运动学的,而非动力学的:对长期世界模型失败的诊断

Finn Rasmus Schäfer, Korbinian Moller, Yuan Gao, Christian Oefinger, Sebastian Schmidt, Johannes Betz

机构 * Autonomous Vehicle Systems Lab(自动驾驶车辆系统实验室) Technical University of Munich(慕尼黑技术大学) Data Analytics and Machine Learning Group(数据分析与机器学习小组)

专题命中 通用世界模型 :world-model(title);world-model(title);world model(abstract,comments);world models(abstract)

AI总结 研究针对世界模型长期失败问题,提出用运动学与动力学重新构建的方法,通过想象的运动学一致性误差(iKCE)及扰动协议进行诊断,在DreamerV3检查点实例化,区分了运动学与动力学想象,揭示模型长期失败特征。

Comments 9 Pages Workshop Paper accepted at RSS Robot World Model Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13769 2026-06-16 cs.RO cs.CV cs.LG 新提交 91%

$μ_0$: A Scalable 3D Interaction-Trace World Model

$\mu_0$: 一种可扩展的3D交互轨迹世界模型

Seungjae Lee, Yoonkyo Jung, Jusuk Lee, Jonghun Shin, Amir Hossein Shahidzadeh, Yao-Chih Lee, H. Jin Kim, Jia-Bin Huang, Furong Huang

机构 * University of Maryland, College Park(马里兰大学帕克分校) Seoul National University(首尔大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 提出基于3D轨迹的可扩展世界模型$\mu_0$,通过预测交互点轨迹实现跨本体机器人学习,无需动作标签,性能媲美有监督模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06192 2026-05-08 cs.CV cs.AI cs.RO 91%

EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields

EA-WM:事件感知生成世界模型与结构运动-视觉动作场

Zhaoyang Yang, Yurun Jin, Lizhe Qi, Cong Huang, Kai Chen

机构 * Fudan University(复旦大学) Zhongguancun Academy(中关村学院) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院) University of Science and Technology of China(中国科学技术大学) DeepCybo(深瞳)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 EA-WM通过结构化运动-视觉动作场闭环连接运动控制与视觉感知,提升机器人空间几何和细粒度交互动态的生成精度。

Comments Preprint. 22 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29090 2026-04-01 cs.LG cs.CV cs.RO 91%

HCLSM: Hierarchical Causal Latent State Machines for Object-Centric World Modeling

HCLSM:面向对象的世界建模的分层因果潜在状态机

Jaber Jaber, Osama Jaber

机构 * RightNow AI

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 HCLSM通过分层时序动态和因果结构学习,改进了基于视频预测未来状态的世界模型,实现了更精确的对象中心化表示和事件边界学习。

Comments 10 pages, 3 tables, 4 figures, 1 algorithm. Code: https://github.com/rightnow-ai/hclsm

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26741 2026-03-31 cs.CV cs.AI cs.RO 91%

Language-Conditioned World Modeling for Visual Navigation

基于语言的视觉导航世界建模

Yifei Dong, Fengyi Wu, Yilong Dai, Lingdong Kong, Guangyu Chen, Xu Zhu, Qiyu Hu, Tianyu Wang, Johnalbert Garnica, Feng Liu, Siyu Huang, Qi Dai, Zhi-Qi Cheng

机构 * University of Washington(华盛顿大学) National University of Singapore(新加坡国立大学) Clemson University(克莱姆森大学) Drexel University(德雷塞尔大学) Microsoft Research(微软研究院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文研究了基于语言的视觉导航问题,提出LCVN数据集并开发了两种模型家族,通过语言 grounding、未来状态预测和动作生成实现统一任务学习。

Comments 19 pages, 6 figures, Code: https://github.com/F1y1113/LCVN

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13024 2026-03-16 cs.CV cs.AI cs.LG eess.IV 91%

SAW: Toward a Surgical Action World Model via Controllable and Scalable Video Generation

SAW:通过可控且可扩展的视频生成实现手术动作世界模型

Sampath Rapuri, Lalithkumar Seenivasan, Dominik Schneider, Roger Soberanis-Mukul, Yufan He, Hao Ding, Jiru Xu, Chenhao Yu, Chenyan Jing, Pengfei Guo, Daguang Xu, Mathias Unberath

机构 * Johns Hopkins University(约翰霍普金斯大学) NVIDIA

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 SAW通过四个轻量信号条件的视频扩散模型生成真实手术动作视频,解决手术AI和模拟中的数据稀缺、稀有事件合成及仿真到现实的差距问题,实现高时间一致性和视觉质量。

Comments The manuscript is under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05438 2026-03-06 cs.CV cs.AI cs.RO 91%

Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model

8个标记的规划:一种用于潜在世界模型的紧凑离散标记器

Dongwon Kim, Gawon Seo, Jinsung Lee, Minsu Cho, Suha Kwak

机构 * KAIST(韩国科学技术院) POSTECH(POSTECH大学) RLWRLD

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出CompACT,一种将观测压缩为8个标记的离散标记器,显著降低计算成本并提升规划效率,为潜在世界模型的实际应用提供可行方案。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11021 2026-02-12 cs.RO cs.AI cs.CV 91%

ContactGaussian-WM: Learning Physics-Grounded World Model from Videos

ContactGaussian-WM: 从视频中学习物理基础的世界模型

Meizhong Wang, Wanxin Jin, Kun Cao, Lihua Xie, Yiguang Hong

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 ContactGaussian-WM通过从稀疏视频中学习物理定律,提升机器人规划与模拟中的复杂动态场景建模能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04075 2026-02-04 cs.LG cs.AI cs.CV 91%

Accurate and Efficient World Modeling with Masked Latent Transformers

基于掩码潜在变换器的精准高效世界建模

Maxime Burchi, Radu Timofte

机构 * Computer Vision Lab, CAIDAS \& IFI, University of Würzburg, Germany

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 EMERALD通过高效掩码潜在变换器实现精准高效的世界建模,首次在1000万环境步骤内超越人类专家性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08931 2026-01-28 cs.CV cs.AI cs.LG 91%

Astra: General Interactive World Model with Autoregressive Denoising

Astra:通用交互世界模型与自回归去噪

Yixuan Zhu, Jiaqi Feng, Wenzhao Zheng, Yuan Gao, Xin Tao, Pengfei Wan, Jie Zhou, Jiwen Lu

机构 * Tsinghua University(清华大学) Kuaishou Technology(快手科技)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 Astra提出了一种通用交互世界模型,通过自回归去噪和动作感知适配器实现长周期视频预测与多样化交互。

Comments Accepted in ICLR 2026. Code is available at: https://github.com/EternalEvan/Astra

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11225 2025-12-15 cs.CV cs.AI cs.LG 91%

VFMF: World Modeling by Forecasting Vision Foundation Model Features

VFMF:通过预测视觉基础模型特征进行世界建模

Gabrijel Boduljak, Yushi Lan, Christian Rupprecht, Andrea Vedaldi

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出了一种基于VFM特征的生成预测器,通过自回归流匹配在潜在空间中实现更准确的预测,提升世界建模的效率和效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21690 2025-11-27 cs.RO cs.CV cs.LG 91%

TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos

TraceGen:在3D轨迹空间中进行世界建模,实现跨躯体视频的学习

Seungjae Lee, Yoonkyo Jung, Inkook Chun, Yao-Chih Lee, Zikui Cai, Hongjia Huang, Aayush Talreja, Tan Dat Dao, Yongyuan Liang, Jia-Bin Huang, Furong Huang

机构 * University of Maryland, College Park(马里兰大学 College Park分校)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 TraceGen通过3D轨迹空间实现跨躯体视频学习,利用紧凑的轨迹表示提升机器人任务学习效率与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20922 2025-10-27 cs.MA cs.AI cs.LG 91%

Revisiting Multi-Agent World Modeling from a Diffusion-Inspired Perspective

Yang Zhang, Xinran Li, Jianing Ye, Shuang Qiu, Delin Qu, Xiu Li, Chongjie Zhang, Chenjia Bai

机构 * Tsinghua University(清华大学) The Hong Kong University of Science and Technology(香港科技大学) Washington University in St. Louis(圣路易斯华盛顿大学) City University of Hong Kong(香港城市大学) Fudan University(复旦大学) Institute of Artificial Intelligence (TeleAI)(人工智能研究所) China Telecom(中国电信) Shenzhen Research Institute of Northwestern Polytechnical University(西北工业大学深圳研究院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

Comments Accepted at NIPS'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14977 2025-10-17 cs.CV cs.AI cs.LG 91%

Terra: Explorable Native 3D World Model with Point Latents

Yuanhui Huang, Weiliang Chen, Wenzhao Zheng, Xin Tao, Pengfei Wan, Jie Zhou, Jiwen Lu

机构 * Tsinghua University(清华大学) Kuaishou Technology(快手科技)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

Comments Project Page: https://huang-yh.github.io/terra/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09737 2025-09-15 cs.CV cs.AI cs.LG 91%

World Modeling with Probabilistic Structure Integration

Klemen Kotar, Wanhee Lee, Rahul Venkatesh, Honglin Chen, Daniel Bear, Jared Watrous, Simon Kim, Khai Loong Aw, Lilian Naing Chen, Stefan Stojanov, Kevin Feigelis, Imran Thobani, Alex Durango, Khaled Jedoui, Atlas Kazemian, Dan Yamins

机构 * Stanford NeuroAI Lab(斯坦福神经人工智能实验室)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18897 2025-08-21 cs.RO cs.AI 91%

MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis

Xiaowei Chi, Kuangzhi Ge, Jiaming Liu, Siyuan Zhou, Peidong Jia, Zichen He, Yuzhen Liu, Tingguang Li, Lei Han, Sirui Han, Shanghang Zhang, Yike Guo

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20523 2025-03-27 cs.CV cs.AI cs.RO 91%

GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving

Lloyd Russell, Anthony Hu, Lorenzo Bertoni, George Fedoseev, Jamie Shotton, Elahe Arani, Gianluca Corrado

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06170 2025-03-13 cs.AI cs.CV cs.RO 91%

Object-Centric World Model for Language-Guided Manipulation

Youngjoon Jeong, Junha Chun, Soonwoo Cha, Taesup Kim

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08261 2025-02-18 cs.RO cs.AI cs.LG 91%

FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model

Chongkai Gao, Haozhuo Zhang, Zhixuan Xu, Zhehao Cai, Lin Shao

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10373 2024-12-16 cs.CV cs.AI cs.LG 91%

GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction

Sicheng Zuo, Wenzhao Zheng, Yuanhui Huang, Jie Zhou, Jiwen Lu

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

Comments Code is available at: https://github.com/zuosc19/GaussianWorld

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12399 2024-10-31 cs.LG cs.AI cs.CV 91%

Diffusion for World Modeling: Visual Details Matter in Atari

Eloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto, Amos Storkey, Tim Pearce, François Fleuret

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

Comments NeurIPS 2024 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09631 2024-03-15 cs.CV cs.AI cs.CL cs.RO 91%

3D-VLA: A 3D Vision-Language-Action Generative World Model

Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, Chuang Gan

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

Comments Project page: https://vis-www.cs.umass.edu/3dvla/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.03622 2023-11-08 cs.RO cs.AI cs.CV cs.LG 91%

TWIST: Teacher-Student World Model Distillation for Efficient Sim-to-Real Transfer

Jun Yamada, Marc Rigter, Jack Collins, Ingmar Posner

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);latent dynamics(abstract);model-based RL(abstract)

Comments 7 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16955 2026-08-19 cs.MA cs.LG 新提交 90%

WONDER: A Radio World Model-based Negotiation Framework for Multi-Agent UAV Coverage Optimization

WONDER:一种基于无线电世界模型的多智能体无人机覆盖优化协商框架

Jiahao Huang, Rongpeng Li, Zhifeng Zhao, Guoru Ding, Honggang Zhang

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 针对灾后无线覆盖中断问题,提出基于JEPA无线电世界模型与PPO演员交替更新的WONDER框架,在含62个城市场景的RadioDynamics仿真中,其平衡得分0.870,覆盖优于STACCA且保持无人机全连通。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08553 2026-08-11 cs.CV cs.LG cs.MM 新提交 90%

MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

MotionCraft:用于视频超分辨率的基于稀疏注意力的潜在世界建模

Rong Fu, Chunlei Meng, Yangchen Zeng, Xiaowen Ma, Yongtai Liu, Wangyu Wu, Shuo Yin, Zijian Zhang, Sicheng Li, Yingrui Ji, Chenhao Wang, Simon Fong

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 MotionCraft是一种可控视频超分辨率框架,通过结合鲁棒运动融合、潜在世界Transformer与自适应稀疏注意力,实现了时间一致的高质量视频重建,且能灵活权衡时间平滑性与重建保真度。

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07548 2026-08-11 cs.RO cs.CV 新提交 90%

SC$^{2}$-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments

SC²-WM:用于连续环境下视觉语言导航的带闭环反馈的自校正世界模型

Xuan Yao, Yuze Zhu, Junyu Gao, Zongmeng Wang, Changsheng Xu

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 针对连续环境视觉语言导航中开环执行的状态漂移问题,提出带闭环反馈的自校正世界模型SC²-WM,通过状态级计划优化与测试时模型自适应提升导航鲁棒性和泛化性。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏