arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 6450 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 4308 篇

2503.04641 2026-02-17 cs.CV cs.AI cs.LG 83%

Simulating the Real World: A Unified Survey of Multimodal Generative Models

模拟现实世界:多模态生成模型的统一综述

Yuqi Hu, Longguang Wang, Xian Liu, Ling-Hao Chen, Yuwei Guo, Yukai Shi, Ce Liu, Anyi Rao, Zeyu Wang, Hui Xiong

机构 * Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(人工智能前沿技术研究所,香港科学与技术大学(广州)) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology Hong Kong SAR(计算机科学与工程系,香港科学与技术大学香港特别行政区) MMLab, The Hong Kong University of Science and Technology(多模态实验室,香港科学与技术大学) School of Electronics and Communication Engineering, Shenzhen Campus of Sun Yat-sen University(电子与通信工程学院,中山大学深圳校区) The Chinese University of Hong Kong, Hong Kong, China(香港中文大学,香港,中国) Tsinghua University, Guangdong, China(清华大学,广东,中国) Bosch (China) Investment Co., Ltd., Shanghai, China(博世(中国)投资有限公司,上海,中国)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文首次系统性地统一研究了2D、视频、3D和4D生成,为多模态生成模型和现实世界模拟提供了统一框架的综述。

Comments Repository for the related papers at https://github.com/ALEEEHU/World-Simulator

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05696 2026-02-13 cs.LG cs.AI cs.RO 83%

A Multi-Fidelity Control Variate Approach for Policy Gradient Estimation

一种多保真度控制变量化法用于策略梯度估计

Xinjie Liu, Cyrus Neary, Kushagra Gupta, Wesley A. Suttle, Christian Ellis, Ufuk Topcu, David Fridovich-Keil

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) The University of British Columbia(不列颠哥伦比亚大学) DEVCOM Army Research Laboratory(国防部陆军研究实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出一种多保真度控制变量化方法,通过结合稀缺目标环境数据和丰富的低保真度模拟数据,提高策略梯度估计的样本效率和收敛速度,适用于有限高保真度数据但丰富低保真度数据的机器人任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22129 2025-12-30 cs.MA cs.AI cs.LG 83%

ReCollab: Retrieval-Augmented LLMs for Cooperative Ad-hoc Teammate Modeling

ReCollab:基于检索增强的LLM用于协作即兴队友建模

Conor Wallace, Umer Siddique, Yongcan Cao

机构 * Department of Electrical and Computer Engineering, University of Texas at San Antonio(电子与计算机工程系,德克萨斯大学圣安东尼奥分校)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 ReCollab通过检索增强生成技术提升即兴团队合作中队友行为建模的适应性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15493 2025-12-18 cs.LG cs.AI stat.ML 83%

Soft Geometric Inductive Bias for Object Centric Dynamics

基于几何的诱导偏置用于物体中心动力学

Hampus Linander, Conor Heins, Alexander Tschantz, Marco Perin, Christopher Buckley

机构 * VERSES AI University of Sussex(苏塞克斯大学) University of Sussex, Department of Informatics(苏塞克斯大学信息学院)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出基于几何代数神经网络的物体中心世界模型,通过柔软几何诱导偏置提升物理动力学模拟的保真度和泛化能力。

Comments 8 pages, 11 figures; 6 pages supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21418 2025-10-27 cs.LG cs.AI 83%

DreamerV3-XP: Optimizing exploration through uncertainty estimation

Lukas Bierling, Davide Pasero, Jan-Henrik Bertrand, Kiki Van Gerwen

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14758 2025-09-19 cs.RO cs.CV cs.LG cs.SY eess.SY 83%

Designing Latent Safety Filters using Pre-Trained Vision Models

Ihab Tabbara, Yuxuan Yang, Ahmad Hamzeh, Maxwell Astafyev, Hussein Sibai

机构 * Department of Computer Science and Engineering, Washington University in St. Louis(计算机科学与工程系,华盛顿大学圣路易斯分校)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05263 2025-09-09 cs.AI cs.CV cs.LG 83%

LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation

Yinglin Duan, Zhengxia Zou, Tongwei Gu, Wei Jia, Zhan Zhao, Luyi Xu, Xinzhu Liu, Yenan Lin, Hao Jiang, Kang Chen, Shuang Qiu

机构 * NetEase, Inc.(网易公司) Beihang University(北京航空航天大学) Tsinghua University(清华大学) City University of Hong Kong(香港城市大学) Independent Researcher & Technical Artists(独立研究者及技术艺术家)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10778 2025-06-13 cs.CV cs.AI cs.LG 83%

SlotPi: Physics-informed Object-centric Reasoning Models

Jian Li, Wan Han, Ning Lin, Yu-Liang Zhan, Ruizhi Chengze, Haining Wang, Yi Zhang, Hongsheng Liu, Zidong Wang, Fan Yu, Hao Sun

机构 * Renmin University of China(中国人民大学) Huawei Technologies Ltd.(华为技术有限公司)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21668 2025-04-08 cs.AI cs.CV cs.LG 83%

Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI

Danaja Rutar, Alva Markelius, Konstantinos Voudouris, José Hernández-Orallo, Lucy Cheke

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04249 2025-02-07 cs.AI cs.LG cs.MA physics.data-an stat.ML 83%

Free Energy Risk Metrics for Systemically Safe AI: Gatekeeping Multi-Agent Study

Michael Walters, Rafael Kaufmann, Justice Sefas, Thomas Kopinski

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments 9 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14853 2024-05-24 cs.LG cs.AI cs.RO 83%

Privileged Sensing Scaffolds Reinforcement Learning

Edward S. Hu, James Springer, Oleh Rybkin, Dinesh Jayaraman

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments ICLR 2024 Spotlight version

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02288 2024-05-20 cs.CV cs.AI cs.RO 83%

Prospective Role of Foundation Models in Advancing Autonomous Vehicles

Jianhua Wu, Bingzhao Gao, Jincheng Gao, Jianhao Yu, Hongqing Chu, Qiankun Yu, Xun Gong, Yi Chang, H. Eric Tseng, Hong Chen, Jie Chen

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments 45 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.06356 2024-04-10 cs.LG cs.AI cs.RO 83%

Policy-Guided Diffusion

Matthew Thomas Jackson, Michael Tryfan Matthews, Cong Lu, Benjamin Ellis, Shimon Whiteson, Jakob Foerster

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments Previously at the NeurIPS 2023 Workshop on Robot Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10812 2024-03-28 cs.LG cs.AI 83%

Learning to Act without Actions

Dominik Schmidt, Minqi Jiang

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments Accepted at ICLR 2024 (spotlight). The code can be found at http://github.com/schmidtdominik/LAPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17198 2024-01-19 cs.LG cs.AI cs.MA 83%

A Model-Based Solution to the Offline Multi-Agent Reinforcement Learning Coordination Problem

Paul Barde, Jakob Foerster, Derek Nowrouzezahrai, Amy Zhang

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11093 2023-08-23 cs.CV cs.AI cs.LG 83%

Video OWL-ViT: Temporally-consistent open-world localization in video

Georg Heigold, Matthias Minderer, Alexey Gritsenko, Alex Bewley, Daniel Keysers, Mario Lučić, Fisher Yu, Thomas Kipf

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.10295 2022-07-22 cs.LG cs.AI cs.RO 83%

Addressing Optimism Bias in Sequence Modeling for Reinforcement Learning

Adam Villaflor, Zhe Huang, Swapnil Pande, John Dolan, Jeff Schneider

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.12670 2022-03-25 cs.LG cs.AI cs.HC cs.NE cs.RO 83%

Competency Assessment for Autonomous Agents using Deep Generative Models

Aastha Acharya, Rebecca Russell, Nisar R. Ahmed

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.14641 2021-12-14 cs.LG cs.AI cs.RO 83%

Learning to Plan Optimistically: Uncertainty-Guided Deep Exploration via Latent Model Ensembles

Tim Seyde, Wilko Schwarting, Sertac Karaman, Daniela Rus

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.08577 2021-07-20 cs.LG cs.AI 83%

Structured World Belief for Reinforcement Learning in POMDP

Gautam Singh, Skand Peri, Junghyun Kim, Hyunseok Kim, Sungjin Ahn

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments Published in ICML 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.02339 2021-07-07 cs.LG cs.AI cs.RO 83%

Multi-Modal Mutual Information (MuMMI) Training for Robust Self-Supervised Deep Reinforcement Learning

Kaiqi Chen, Yong Lee, Harold Soh

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments 10 pages, Published in ICRA 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.04862 2019-04-11 cs.LG cs.AI cs.CV stat.ML 83%

SWNet: Small-World Neural Networks and Rapid Convergence

Mojan Javaheripi, Bita Darvish Rouhani, Farinaz Koushanfar

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.11388 2019-02-19 cs.LG cs.AI cs.RO 83%

Deep Intrinsically Motivated Continuous Actor-Critic for Efficient Robotic Visuomotor Skill Learning

Muhammad Burhan Hafez, Cornelius Weber, Matthias Kerzel, Stefan Wermter

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Journal ref Paladyn, Journal of Behavioral Robotics, Volume 10, Issue 1, Pages 14-29, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19492 2026-08-21 cs.LG cs.RO 新提交 82%

Beyond Multimodal Alignment: Certifying Physical Language through Response Substitution and Ordered Execution

超越多模态对齐:通过响应替换与有序执行验证物理语言

Kaizhen Tan, Xin Xu, Siru Tao, Yixiao Li, Hanzhe Hong, Yang Feng, Heqing Du

机构 * New York University(纽约大学) Carnegie Mellon University(卡内基梅隆大学) Columbia University(哥伦比亚大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出DBOSC方法,在Cluster Haptic数据集与弹塑性系统中验证了多模态物理表示的可执行含义,明确了执行器、图表等因素对动作组合的影响,区分了多项可测试的操作能力成就。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14481 2026-08-17 cs.RO cs.AI 新提交 82%

Ensuring Safe Physical AI in Urban Mobility via Hazard-Informed Synthesized Envelopes

通过危险信息合成包络确保城市移动中的安全物理人工智能

Alexei Odinokov, Rostislav Yavorskiy

机构 * Xortech DOO

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 针对异构机器人城市部署的安全挑战,提出结合危险分析与运行时执行的统一框架,通过跨层安全转换保障物理AI的城市移动安全。

Comments The 2026 International Conference on Control, Robotics Engineering and Technology (CRET 2026), https://www.cret.net/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07693 2026-08-11 cs.CV cs.AI cs.CL 新提交 82%

CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting

CosmosAlign:适配世界基础模型用于生成式交通视频预测

Quang Minh Dinh, Tuan Kiet Doan

机构 * Simon Fraser University(西蒙菲莎大学) Institut Polytechnique de Paris(巴黎综合理工学院)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 该研究提出基于Cosmos3-Nano的CosmosAlign框架,通过两阶段LoRA适配与推理优化,在AI City Challenge 2026 Track 5基准中获76.49分排名第一,实现了高质量交通视频预测。

Comments Accepted at ECCVW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29393 2026-08-11 cs.RO 版本更新 82%

AquaJEPA: An Action-Conditioned Multimodal JEPA Family for Underwater Robot Dynamics

AquaJEPA:面向水下机器人动力学的动作条件多模态预测表示

Alan-Barsag Gazzaev, Alexey Gavrilov, Sergey Muravyov

机构 * ITMO University(ITMO大学)

专题命中 通用世界模型 :world model(abstract);world-model(abstract);world model(abstract);world-model(abstract)

AI总结 该研究提出动作条件多模态预测模型AquaJEPA,在Stonefish环境中对比多种基线,经120个带计划DVL损失的配对环境实验,其闭环性能最优,配对最终误差显著优于多数基线。

Comments Submitted to IEEE ICRA 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02150 2026-08-06 cs.CV cs.AI 版本更新 82%

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs

PhyCheck:面向视频大语言模型的物理规律理解细粒度证据驱动数据集

Zhongjie Ba, Shengwang Xu, Peng Cheng, Jinyang Zou, Ting Yu, Zhibo Wang, Zhan Qin

机构 * Zhejiang University(浙江大学) ZJU-Hangzhou Global Scientific and Technological Innovation Center(浙江大学杭州国际科技创新中心) Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出PhyCheck数据集,通过粗粒度、细粒度及诊断子集提升视频大语言模型的物理规律理解,实验证实其训练效果,同时指出当前模型难以整合因果条件的问题。

Comments 15pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00115 2026-08-03 cs.CV cs.LG stat.ML 版本更新 82%

Physics from Video: Identifiability of Time-Invariant Second-Order ODEs under Minimal Trajectory Conditions

来自视频的物理:最小轨迹条件下时不变二阶ODE的可辨识性

Yuanyuan Wang, Wenjie Wang, Kun Zhang, Mingming Gong

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 研究从原始像素中辨识连续时间物理定律的结构可辨识性,证明在最小轨迹条件下,编码器-仅管道可唯一恢复二阶线性ODE参数,并引入方差底正则化器稳定无解码器目标。

Comments Accepted at ICML 2026. Updated to the camera-ready version; main results unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26579 2026-07-30 cs.RO cs.CV 新提交 82%

ContactFlow: A video action conditioning that transfers across embodiments

ContactFlow:一种可跨 embodiment 迁移的视频动作条件模型

Sami Azirar, Enrico Pallotta, Jan Nogga, Jürgen Gall, Sven Behnke, Hermann Blum

机构 * University of Bonn(波恩大学) Lamarr Institute(拉马尔研究所)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 Contact Flow 是一种与 embodiment 无关的动作表示,可跨人类演示与不同机器人 embodiment 迁移,用于构建能预测合理操纵结果的世界模型,集成到 pipeline 后在相关任务上表现良好。

详情

展开后加载摘要…

URL PDF HTML 收藏