arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 6450 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 4308 篇

2606.27364 2026-06-26 cs.CV 新提交 86%

PhysiFormer: Learning to Simulate Mechanics in World Space

PhysiFormer: 在世界空间中学习模拟力学

Yiming Chen, Yushi Lan, Andrea Vedaldi

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 提出PhysiFormer,一种直接在世界坐标下通过去噪扩散过程预测3D物体顶点轨迹的扩散Transformer,无需归纳偏置即可生成物理合理的刚性和弹性运动,并泛化到混合材料、未见几何和更多物体。

Comments Project page: https://yimingc9.github.io/physiformer

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09535 2026-04-13 cs.CV 86%

EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks

EgoTL:用于长时任务的目视思考链

Lulin Liu, Dayou Li, Yiqing Liang, Sicong Jiang, Hitesh Vijay, Hezhen Hu, Xuhai Xu, Zirui Liu, Srinivas Shakkottai, Manling Li, Zhiwen Fan

机构 * UMN(明尼苏达大学) TAMU(德克萨斯农工大学) Brown University(布朗大学) McGill University(麦吉尔大学) UT Austin(德克萨斯大学奥斯汀分校) Columbia University(哥伦比亚大学) Northwestern University(西北大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 本文提出EgoTL,通过构建目视思考链管道,提升长时任务中视觉语言模型和世界模型的推理与规划能力,发现基础模型在作为目视助手或开放世界模拟器方面仍有不足。

Comments https://ego-tl.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18422 2026-02-23 cs.CV 86%

Generated Reality: Human-centric World Simulation using Interactive Video Generation with Hand and Camera Control

生成现实:基于交互式视频生成的人本世界模拟

Linxi Xie, Lisong C. Sun, Ashley Neall, Tong Wu, Shengqu Cai, Gordon Wetzstein

机构 * Stanford University(斯坦福大学) NYU Shanghai(纽约大学上海分校) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 本文提出了一种基于交互式视频生成的人本世界模拟系统,通过结合头部和手部姿态控制,提升虚拟环境的交互性和用户控制感。

Comments Project page here: https://codeysun.github.io/generated-reality

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21840 2025-10-28 cs.CV cs.GR 86%

Improving the Physics of Video Generation with VJEPA-2 Reward Signal

Jianhao Yuan, Xiaofeng Zhang, Felix Friedrich, Nicolas Beltran-Velez, Melissa Hall, Reyhane Askari-Hemmat, Xiaochuang Han, Nicolas Ballas, Michal Drozdzal, Adriana Romero-Soriano

机构 * FAIR, Meta Superintelligence Labs(FAIR、Meta超智能实验室) University of Oxford(牛津大学) Mila - Québec AI Institute(魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) Columbia University(哥伦比亚大学) McGill University(麦吉尔大学) Canada CIFAR AI Chair(加拿大CIFAR人工智能 chair)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

Comments 2 pages

Journal ref Winning entry of the ICCV 2025 Physics IQ Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01793 2025-10-06 cs.LG cs.AI 86%

STORI: A Benchmark and Taxonomy for Stochastic Environments

Aryan Amit Barsainyan, Jing Yu Lim, Dianbo Liu

专题命中 通用世界模型 :world model(abstract,comments);world models(abstract,comments);world model(abstract,comments);world models(abstract,comments)

Comments v2. New mathematical formulation and renamed notation; added additional experiments and a detailed analytical case study on error behaviors in world models under different stochasticity types; link to code repository for reproducibility: https://github.com/ARY2260/stori

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10191 2025-04-15 cs.CL cs.AI 86%

Localized Cultural Knowledge is Conserved and Controllable in Large Language Models

Veniamin Veselovsky, Berke Argin, Benedikt Stroebl, Chris Wendler, Robert West, James Evans, Thomas L. Griffiths, Arvind Narayanan

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13518 2024-07-19 cs.LG 86%

Model-based Policy Optimization using Symbolic World Model

Andrey Gorodetskiy, Konstantin Mironov, Aleksandr Panov

专题命中 通用世界模型 :world model(title);world model(title);model-based reinforcement learning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19878 2024-05-31 cs.LG cs.GT 86%

Learning from Random Demonstrations: Offline Reinforcement Learning with Importance-Sampled Diffusion Models

Zeyu Fang, Tian Lan

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26040 2026-07-29 cs.LG stat.ML 新提交 86%

Reinformed Dreamer: An Asymmetric World Model Efficiently Trained through Latent Guidance

强化信息梦行者:通过潜在引导有效训练的非对称世界模型

Gaspard Lambrechts, Adrien Bolland, Daniel Ebi, Damien Ernst

专题命中 通用世界模型 :world model(title);world model(title);model-based reinforcement learning(abstract);分类 cs.LG

AI总结 研究基于模型的强化学习中,非对称学习对观测及特权信息表示的影响。针对‘信息梦行者’局限性,提出用潜在引导的新目标,形成‘强化信息梦行者’算法,实验显示其比之前非对称方法有更持续改进。

Comments 8 pages, 18 pages total, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04541 2026-07-07 cs.CV cs.AI cs.LG cs.RO 新提交 86%

CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining

CRISP:一种通过基于预测的世界模型预训练实现驾驶的时空相机-雷达主干网络

Jingyu Song, Yi Liu, Katherine A. Skinner

机构 * Department of Robotics, University of Michigan(密歇根大学机器人系)

专题命中 通用世界模型 :world-model(title);world-model(title);分类 cs.AI、cs.LG、cs.CV

AI总结 研究通过基于预测的表示学习预训练时空相机-雷达主干网络CRISP,利用历史多视图图像和雷达扫描,通过预测未来激光雷达点云学习统一鸟瞰图表示,介绍相关组件,实验表明其能提升点云预测并有效迁移至下游任务。

Comments 17 pages, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12299 2025-04-17 cs.AI cs.CV cs.LG 86%

Adapting a World Model for Trajectory Following in a 3D Game

Marko Tot, Shu Ishida, Abdelhak Lemkhenter, David Bignell, Pallavi Choudhury, Chris Lovett, Luis França, Matheus Ribeiro Furtado de Mendonça, Tarun Gupta, Darren Gehring, Sam Devlin, Sergio Valcarcel Macua, Raluca Georgescu

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.AI、cs.LG、cs.CV;dynamics model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09859 2024-03-18 cs.LG 86%

MAMBA: an Effective World Model Approach for Meta-Reinforcement Learning

Zohar Rimon, Tom Jurgenson, Orr Krupnik, Gilad Adler, Aviv Tamar

专题命中 通用世界模型 :world model(title);world model(title);model-based RL(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.09968 2018-12-27 cs.LG cs.AI cs.NE 86%

VMAV-C: A Deep Attention-based Reinforcement Learning Algorithm for Model-based Control

Xingxing Liang, Qi Wang, Yanghe Feng, Zhong Liu, Jincai Huang

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22363 2026-06-23 cs.AI cs.LG cs.RO 新提交 85%

Reference-Free Assessment of Physical Consistency in World Model-based Video Generation

基于世界模型的视频生成中物理一致性的无参考评估

Yun Oh, Sukmin Yun

机构 * Hanyang University ERICA(汉阳大学ERICA校区)

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.AI、cs.LG、cs.RO

AI总结 提出结合相对与绝对方法的无参考度量,用于评估生成视频的物理一致性,无需人工投票或真实参考,通过DROID-SLAM和SEA-RAFT量化不一致性,过滤后视频任务成功率提升超8%,并实现时空定位。

Comments Accepted to the 2nd 3D-LLM/VLA Workshop, CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12197 2025-02-10 cs.RO cs.AI 85%

Towards Interpretable Visuo-Tactile Predictive Models for Soft Robot Interactions

Enrico Donato, Thomas George Thuruthel, Egidio Falotico

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments IEEE RAS EMBS 10th International Conference on Biomedical Robotics and Biomechatronics (BioRob 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05781 2024-09-04 cs.LG cs.AI cs.CV 85%

CURLing the Dream: Contrastive Representations for World Modeling in Reinforcement Learning

Victor Augusto Kich, Jair Augusto Bottega, Raul Steinmetz, Ricardo Bedin Grando, Ayano Yorozu, Akihisa Ohya

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.AI、cs.LG、cs.CV

Comments Paper accepted for 24th International Conference on Control, Automation and Systems (ICCAS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.01163 2021-12-03 cs.LG cs.AI cs.RO 85%

Robust Robotic Control from Pixels using Contrastive Recurrent State-Space Models

Nitish Srivastava, Walter Talbott, Martin Bertran Lopez, Shuangfei Zhai, Josh Susskind

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments NeurIPS Deep Reinforcement Learning Workshop 2021. Code can be found at https://github.com/apple/ml-core

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14650 2026-08-18 cs.LG cs.RO 新提交 85%

Paired Exact-Reset Evaluation of a Prediction-Derived Medium-to-Full World-Model Cascade

预测衍生的中到全世界模型级联的配对精确重置评估

Malo de Pastor

机构 * Télécom SudParis(南巴黎电信学院) Institut Polytechnique de Paris(巴黎理工学院)

专题命中 通用世界模型 :world-model(title);world-model(title);分类 cs.LG、cs.RO

AI总结 该研究提出配对精确重置评估协议,在PushT任务上验证预测接口路由器可降低决策成本,其优势限于低计算价格,不支持因果充分性等特性。

Comments 20 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09381 2026-08-11 cs.RO 新提交 85%

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

JEPA-WAM:基于联合嵌入世界建模的视觉-语言-动作策略学习

Yihan Lin, Jiawei He, Shifeng Bao, Chen Zhao, Yang Li, Xiaobo Wang, Yan Wang, Cheng Chi, Jing Zhang

机构 * XYZ Embodied AI(XYZ具身智能)

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.RO;predictive model(abstract)

AI总结 该研究提出JEPA-WAM,一种构建于预训练V-JEPA空间的隐式WAM,通过共享预测器耦合隐式转移预测与动作生成,在LIBERO-Plus等数据集上取得优异的机器人操作策略性能,且泛化能力强。

Comments 22 pages, 7 figures. Project page: https://spritewithoutice.github.io/JEPA_WAM/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01397 2026-08-04 cs.RO cs.CV 新提交 85%

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

SG-WAM:几何感知策略空间中的自引导世界建模

Ruiteng Zhao, Zhengshen Zhang, Yue Su, Wenshuo Wang, Jiahui Li, Zhiyuan Yang, Francis E. H. Tay, Marcelo H. Ang, Haiyue Zhu

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.CV、cs.RO

AI总结 SG-WAM是自引导的几何感知世界动作模型框架,通过策略衍生空间联合优化动态预测等,在LIBERO等数据集上实现高成功率,性能优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24121 2026-06-23 cs.RO 版本更新 85%

Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation

在线世界建模实现真实世界逆强化学习从观察中学习

Tyler Han, Bat Nemekhbold, Siyang Shen, Rohan Baijal, Richard Ebock, Harine Ravichandiran, Sanghun Jung, Kevin Huang, Byron Boots

机构 * University of Washington(华盛顿大学)

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.RO;simulation model(abstract)

AI总结 提出MPAIL2方法,通过在线世界建模实现从观察中逆强化学习,首次在真实世界从零学习视觉操作任务,40分钟内达到82%成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17391 2026-06-17 cs.CL cs.AI cs.LG 新提交 85%

NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama

NarrativeWorldBench:面向长程共创音频剧的前沿饱和基准与潜在世界模型

Logan Mann, Abdur Rahman, Mohammad Saifullah, Taaha Kazi, Vasu Sharma

机构 * University of California, Santa Barbara(加州大学圣塔芭芭拉分校) Pocket FM

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.AI、cs.LG

AI总结 提出NarrativeWorldBench基准,在九种叙事结构指标上评估21个模型,并引入N-VSSM变分状态空间模型,通过Mamba-2骨干和事件条件后验在200集以上维持结构化潜在状态,在长弧一致性和可控性上超越Claude Opus 4.5。

Comments 10 pages. Accepted to the ICML 2026 Workshops on High-dimensional Learning Dynamics (HiLD) and Culture x AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03604 2026-04-09 cs.CV cs.AI 85%

A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures

一种轻量级的用于基于能量的联合嵌入预测架构库

Basile Terver, Randall Balestriero, Megi Dervishi, David Fan, Quentin Garrido, Tushar Nagarajan, Koustuv Sinha, Wancong Zhang, Mike Rabbat, Yann LeCun, Amir Bar

机构 * Meta FAIR INRIA(法国国家信息与自动化研究所) New York University(纽约大学)

专题命中 通用世界模型 :world model(abstract,comments);world models(abstract,comments);world model(abstract,comments);world models(abstract,comments)

AI总结 本文提出EB-JEPA库,展示如何将图像级自监督学习的表示学习技术扩展到视频和动作条件世界模型,通过实验验证其在任务中的有效性。

Comments v2: clarify confusion in definition of JEPAs vs. regularization-based JEPAs v3: Camera-ready of ICLR world models workshop, fixed formatting and ViT config / results

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01605 2026-04-03 cs.CV cs.RO 85%

F3DGS: Federated 3D Gaussian Splatting for Decentralized Multi-Agent World Modeling

F3DGS:联邦3D高斯点云用于去中心化多智能体世界建模

Morui Zhu, Mohammad Dehghani Tezerjani, Mátyás Szántó, Márton Vaitkus, Song Fu, Qing Yang

机构 * University of North Texas(北德克萨斯大学) Budapest University of Technology and Economics(布达佩斯技术与经济大学)

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.CV、cs.RO

AI总结 F3DGS通过联邦学习实现去中心化多智能体3D重建,利用共享几何框架和可见性感知聚合解决部分观测问题,实现分布式优化。

Comments Accepted to the CVPR 2026 SPAR-3D Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10098 2026-02-17 cs.RO cs.CV 85%

VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model

VLA-JEPA: 通过潜在世界模型增强视觉-语言-动作模型

Jingwen Sun, Wenyao Zhang, Zekun Qi, Shaojie Ren, Zezhi Liu, Hanxin Zhu, Guangzhong Sun, Xin Jin, Zhibo Chen

机构 * University of Science and Technology of China(中国科学技术大学) Zhongguancun Academy, Beijing, China(中关村学院) Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) Eastern Institute of Technology, Ningbo(宁波东部科技研究院) University of Chinese Academy of Sciences(中国科学院大学) Nankai University(南开大学)

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.CV、cs.RO

AI总结 VLA-JEPA通过潜在世界模型提升视觉-语言-动作模型的泛化与鲁棒性,采用无泄露状态预测和两阶段训练策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18319 2025-11-25 cs.AI cs.LG cs.SY eess.SY 85%

Weakly-supervised Latent Models for Task-specific Visual-Language Control

弱监督潜在模型用于任务特定的视觉语言控制

Xian Yeow Lee, Lasitha Vidyaratne, Gregory Sin, Ahmed Farahat, Chetan Gupta

机构 * Industrial AI Lab, Hitachi America, Ltd.(日立美国有限公司工业人工智能实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出了一种任务特定的潜在动态模型,利用目标状态监督学习动作诱导位移,以提高空间定位任务中的视觉语言控制性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07982 2025-07-11 cs.CV cs.AI 85%

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

Haoyu Wu, Diankun Wu, Tianyu He, Junliang Guo, Yang Ye, Yueqi Duan, Jiang Bian

机构 * Microsoft Research(微软研究院) Tsinghua University(清华大学)

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.AI、cs.CV

Comments 18 pages, project page: https://GeometryForcing.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17330 2023-10-27 cs.LG cs.AI 85%

CQM: Curriculum Reinforcement Learning with a Quantized World Model

Seungjae Lee, Daesol Cho, Jonghae Park, H. Jin Kim

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.AI、cs.LG

Comments Accepted to NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00760 2023-10-03 cs.RO 85%

Uncertainty-aware hybrid paradigm of nonlinear MPC and model-based RL for offroad navigation: Exploration of transformers in the predictive model

Faraz Lotfi, Khalil Virji, Farnoosh Faraji, Lucas Berry, Andrew Holliday, David Meger, Gregory Dudek

专题命中 通用世界模型 :model-based RL(title,abstract);predictive model(title,abstract);environment model(abstract);model-based reinforcement learning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11698 2022-10-24 cs.LG cs.AI 85%

Learning Robust Dynamics through Variational Sparse Gating

Arnav Kumar Jain, Shivakanth Sujit, Shruti Joshi, Vincent Michalski, Danijar Hafner, Samira Ebrahimi-Kahou

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏