arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 6450 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 4308 篇

2607.00673 2026-07-02 cs.RO 新提交 93%

Path Planning in Physically Viable World Models

物理可行世界模型中的路径规划

Su Ann Low, Cheng-Hsi Hsiao, Xingjian Li, Adam J. Thorpe, Ufuk Topcu, Krishna Kumar

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出物理可行世界模型,通过物理仿真生成修改后的场景,使机器人能在部署前评估未来地形变化对路径可行性的影响。

Comments 18 pages, 7 figures, submitted to CORL

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00457 2026-07-02 cs.AI 新提交 93%

Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments

面向演化环境中具身智能体的多尺度世界模型混合

Jinwoo Jang, Daniel J. Rho, Sihyung Yoon, Hyunsuk Cho, Honguk Woo

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出MuSix框架,通过尺度感知的世界模型混合与演化,解决具身智能体在动态环境中的多尺度推理和知识适应问题,在EmbodiedBench和HAZARD上超越现有方法。

Comments Accepted at ECCV 2026. 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30639 2026-06-30 cs.AI cs.CL 93%

Self-Evolving World Models for LLM Agent Planning

面向LLM智能体规划的自我进化世界模型

Xuan Zhang, Wenxuan Zhang, See-Kiong Ng, Yang Deng

机构 * National University of Singapore(国立新加坡大学) Singapore University of Technology and Design(新加坡科技设计大学) Singapore Management University(新加坡管理学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出WorldEvolver框架,通过情景记忆、语义记忆和选择性预见三个模块在测试时修正世界模型,提升预测准确性和下游规划成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22960 2026-06-30 cs.CV 93%

UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models

UCM:基于时间感知位置编码变形的相机控制与内存统一建模

Tianxing Xu, Zixuan Wang, Guangyuan Wang, Li Hu, Zhongyi Zhang, Peng Zhang, Bang Zhang, Songhai Zhang

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 UCM通过时间感知位置编码变形机制实现相机控制与长期记忆的统一建模,在真实和合成基准测试中显著提升场景一致性并实现高保真视频生成的精确控制。

Comments Project Page: https://humanaigc.github.io/ucm-webpage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27681 2026-06-29 cs.LG cs.CL 新提交 93%

Textual Belief States for World Models: Identifiable Representation Learning Under Strict Mediation

用于世界模型的文本信念状态:严格中介下的可识别表示学习

Xiang Gao, Kaiwen Dong, Yuguang Yao, Padmaja Jonnalagedda, Kamalika Das

机构 * Intuit AI Research(Intuit AI研究)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出严格中介原则解决LLM中历史旁路导致的潜在状态不可识别问题,引入离散文本潜在状态和因子化GRPO方法,在TextWorld和ScienceWorld上实现表示质量和滚动性能显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27644 2026-06-29 cs.CV 新提交 93%

CascadeOcc: Rethinking 3D Occupancy World Models with Cascaded VQ Representations

CascadeOcc: 用级联VQ表示重新思考3D占用世界模型

Kyumin Hwang, Wonhyeok Choi, Jaeyeul Kim, Jihun Park, Daehee Park, Sunghoon Im

机构 * Daegu Gyeongbuk Institute of Science and Technology (DGIST)(大邱庆北科学技术院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出CascadeOcc,一种通过级联向量量化机制在自回归框架中利用占用表示内在结构层次,实现从粗到细的3D场景预测和运动规划,在4D占用预测和运动规划基准上取得优越性能。

Comments Accepted to IEEE Signal Processing Letters (SPL), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16732 2026-06-29 cs.CV 版本更新 93%

A Comprehensive Survey on World Models for Embodied AI

具身AI世界模型综述

Xinqing Li, Xin He, Le Zhang, Min Wu, Xiaoli Li, Yun Liu

机构 * College of Computer Science and the Academy for Advanced Interdisciplinary Studies, Nankai University(南开大学计算机科学学院与前沿交叉学科研究院) School of Computer Science and Engineering, Tianjin University of Technology(天津理工大学计算机科学与工程学院) School of Information and Communication Engineering, University of Electronic Science and Technology of China(电子科技大学信息与通信工程学院) Institute for Infocomm Research (I2R), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局资讯通信研究院) Information Systems Technology and Design (ISTD) Pillar, Singapore University of Technology and Design (SUTD)(新加坡科技设计大学信息系统科技与设计系)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文系统综述了具身AI中的世界模型,提出了功能、时间建模和空间表示的三轴分类法,并总结了数据资源、评估指标及开放挑战。

Comments https://github.com/Li-Zn-H/AwesomeWorldModels

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23296 2026-06-23 cs.RO 新提交 93%

IOI: Decoupling Kinematics and Physics for Interactive World Models

IOI: 解耦运动学与物理学的交互式世界模型

Chengyu Bai, Peidong Jia, Tiecheng Guo, Yukai Wang, Rui Ma, Fangyuan Zhao, Chunkai Fan, Xiaobao Wei, Jintao Chen, Hao Wang, Ying Li, Xiaozhu Ju, Jian Tang, Shanghang Zhang

机构 * Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) Peking University(北京大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出IOI混合交互式世界模型,通过显式运动学先验与学习物理动力学解耦,实现精确控制对齐和物理合理视觉反馈,在RoboTwin基准上达到最优仿真性能与零样本泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03208 2026-06-18 cs.LG 版本更新 93%

Hierarchical Planning with Latent World Models

基于潜在世界模型的分层规划

Wancong Zhang, Basile Terver, Artem Zholus, Soham Chitnis, Harsh Sutaria, Mido Assran, Randall Balestriero, Amir Bar, Adrien Bardes, Yann LeCun, Nicolas Ballas

机构 * FAIR at Meta(Meta旗下的FAIR) New York University(纽约大学) Mila - Québec AI Institute(魁北克AI研究院) Brown University(布朗大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出HWM架构,通过多时间尺度潜在世界模型和潜在匹配实现分层模型预测控制,解决长时域任务中单层规划失败和计算爆炸问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18180 2026-06-17 cs.CV 新提交 93%

EgoCS-400K: An Egocentric Gameplay Dataset for World Models

EgoCS-400K:面向世界模型的自我中心游戏数据集

Rongjin Guo, Dong Liang, Yuhao Liu, Fang Liu, Tianyu Huang, Gerhard P. Hancke, Rynson W. H. Lau

机构 * City University of Hong Kong(香港城市大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 为支持世界模型研究,构建大规模自我中心游戏数据集EgoCS-400K,包含40万第一人称视频和1万小时游戏轨迹,支持动作条件未来预测、状态事件场景展开等交互式视觉建模任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16519 2026-06-16 cs.CV 新提交 93%

BadWorld: Adversarial Attacks on World Models

BadWorld:对世界模型的对抗攻击

Linghui Shen, Mingyue Cui, Xingyi Yang

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出BadWorld框架,通过自监督速度攻击和轨迹自适应双层优化,对自回归视觉世界模型进行无标签对抗攻击,暴露其结构脆弱性。

Comments Project Page: https://linghuiishen.github.io/BadWorld/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05963 2026-06-16 cs.LG 版本更新 93%

Next-Latent Prediction Transformers Learn Compact World Models

下一潜在预测变换器学习紧凑世界模型

Jayden Teoh, Manan Tomar, Kwangjun Ahn, Edward S. Hu, Tim Pearce, Pratyusha Sharma, Akshay Krishnamurthy, Riashat Islam, Alex Lamb, John Langford

机构 * Massachusetts Institute of Technology(麻省理工学院) Microsoft Research(微软研究院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出NextLat方法,通过潜在空间中的自监督预测训练变换器学习紧凑世界模型,提升泛化能力和推理效率。

Comments Microsoft Research Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13877 2026-06-15 cs.RO 新提交 93%

ContactWorld: What Matters in Vision-Tactile World Models for Contact-Rich Manipulation

ContactWorld: 视觉-触觉世界模型中什么对接触丰富操作至关重要

Zhiyuan Zhang, Pokuang Zhou, Kaidi Zhang, Adeesh Desai, Temitope Amosa, Davood Soleymanzadeh, Jiuzhou Lei, Minghui Zheng, Yu She

机构 * School of Industrial Engineering, Purdue University(普渡大学工业工程学院) Department of Mechanical Engineering, Texas A&M University(德克萨斯农工大学机械工程系)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 通过12项接触丰富操作任务,发现空间结构化和时间连续的表征(如点云)能显著提升规划成功率,且触觉传感的有效性依赖于跨模态表征兼容性。

Comments 32 pages, 12 figures, supplementary material included

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12979 2026-06-12 cs.LG 新提交 93%

EPM-JEPA: Operator-Side Experience Modulation in JEPA-Family World Models

EPM-JEPA:JEPA系列世界模型中的算子侧经验调制

Vedant Pandya

机构 * School of Artificial Intelligence and Data Engineering (SAIDE), Indian Institute of Technology Jodhpur(印度理工学院焦特布尔分校人工智能与数据工程学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出EPM-JEPA,通过LoRA在权重层面调制预测器,以应对测试时动态偏移;实验表明其优于无记忆基线,但效果弱于预期,并揭示了三种独立动力学过程。

Comments 16 pages, 5 figures, 9 tables, 5 code listings. Pre-registered experimental study with mechanism analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12783 2026-06-12 cs.AI 新提交 93%

A Tutorial on World Models and Physical AI

世界模型与物理AI教程

Il-Seok Oh

机构 * Department of Computer Science and Artificial Intelligence/CAIIT, Jeonju, Jeonbuk, South Korea(韩国全北全州计算机科学与人工智能系/CAIIT)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出统一框架,区分显式与隐式世界模型,并探讨其在机器人、自动驾驶等物理AI领域的应用,以及迈向通用人工智能的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12072 2026-06-11 cs.CV 新提交 93%

World Model Self-Distillation: Training World Models to Solve General Tasks

世界模型自蒸馏:训练世界模型以解决通用任务

Sebastian Stapf, Pablo Acuaviva Huertos, Aram Davtyan, Paolo Favaro

机构 * Department of Computer Science(计算机科学系)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出结合自蒸馏与强化学习的框架,从预训练视频生成器中提取任务解决能力,无需配对任务视频,在基准测试中超越原始模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15466 2026-06-09 cs.CV 版本更新 93%

Entity-Centric World Models: Interaction-Aware Masking for Causal Video Prediction

以实体为中心的世界模型:交互感知的掩码用于因果视频预测

Santosh Kumar Paidi

机构 * Genentech, Inc.(基因泰克公司)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出IA-JEPA,通过运动中心的自监督掩码策略,优先捕捉物理交互,提升因果推理任务的准确性,并在真实世界动作和物理谜题中验证了其泛化能力。

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06492 2026-06-09 cs.RO 版本更新 93%

How Well Do Latent World Models Understand Partially Observable Safety Constraints?

潜在世界模型如何理解部分可观测的安全约束?

Matthew Kim, Kensuke Nakamura, Andrea Bajcsy

机构 * UC San Diego(加州大学圣地亚哥分校) Carnegie Mellon University(卡内基梅隆大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究潜在世界模型在部分可观测安全约束下的故障模式,提出互信息度量和滚动预测度量来诊断估计间隙和预测间隙,并通过多模态监督和共形风险校准缓解问题,提高机器人操作安全性。

Comments 10 tables 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06832 2026-06-08 cs.RO 新提交 93%

STRIPS-WM: Learning Grounded Propositional STRIPS-style World Models from Images

STRIPS-WM:从图像学习基于命题的STRIPS风格世界模型

Abhiroop Ajith, Constantinos Chamzas

机构 * Worcester Polytechnic Institute(沃斯特理工学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出STRIPS-WM框架,从图像转换中学习符号化世界模型,用于机器人视觉任务规划,提升规划成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05925 2026-06-05 cs.AI 93%

Towards World Models in Biomedical Research

迈向生物医学研究的世界模型

Guangyu Wang, Jingkun Yue, Siqi Zhang, Yu Liu, Xiaoyu Wang, Mingyuan Meng, Changwei Ji, Zongbo Han, Yulin Wang, Yang Yue, Frank Fu, Ting Chen, Song Wu, Ziwei Liu, Jiangning Song, Ming Li, Gao Huang, Xiaohong Liu, Athanasios Vasilakos, Xingcai Zhang, Ping Zhang, Yong Li

机构 * State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China(网络与交换技术国家重点实验室,北京邮电大学,北京,中国) Department of Engineering Science, University of Oxford, Oxford, United Kingdom(英国牛津大学工程科学系,牛津,英国) Institute of Medical Artificial Intelligence, South China Hospital, Medical School, Shenzhen University, Shenzhen, Guangdong, China(医学人工智能研究所,南方医院,医学学院,深圳大学,深圳,广东,中国) Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence, Beijing, China(中关村学院及中关村人工智能研究院,北京,中国) Beijing National Research Center for Information Science and Technology (BNRist), Tsinghua University, 100084, Beijing, China(北京信息科学与技术国家研究中心(BNRist),清华大学,100084,北京,中国) Department of Chemical and Nano Engineering, University of California, San Diego, La Jolla, CA, USA(美国加州大学圣地亚哥分校化学与纳米工程系,La Jolla,CA,美国) Nanyang Technological University, Singapore(新加坡南洋理工大学) Monash Biomedicine Discovery Institute and Department of Biochemistry and Molecular Biology, Monash University, Melbourne, Victoria, Australia(莫纳什大学生物医学发现研究所和生物化学与分子生物学系,墨尔本,维多利亚,澳大利亚) David R. Cheriton School of Computer Science, University of Waterloo, Waterloo, Ontario, Canada(加拿大滑铁卢大学戴维·R·切里顿计算机科学学校,滑铁卢,安大略,加拿大) Department of ICT and Center for AI Research, University of Agder (UiA), Jon Lilletuns vei 9, Grimstad, Norway(挪威阿格德大学(UiA)信息与通信技术系及人工智能研究中心,Jon Lilletuns vei 9,Grimstad,挪威) Department of Electronic Engineering, Tsinghua University, Beijing, China(清华大学电子工程系,北京,中国)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出生物医学世界模型作为AI驱动发现的新范式,通过学习分子、细胞、组织和临床状态的潜在表征及干预条件动态,实现未来轨迹模拟,并探讨其在虚拟细胞、类器官、虚拟患者和手术模拟等应用中的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05015 2026-06-04 cs.RO 93%

Generalization of World Models under Environmental Variability for Vision-based Quadrotor Navigation

环境变异性下基于视觉的四旋翼导航的世界模型泛化

Luca Zanatta, Grzegorz Malczyk, Kostas Alexis

机构 * Norwegian University of Science and Technology(挪威科学与技术大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 通过基于视觉的四旋翼导航测试,研究世界模型在不同环境随机性下的鲁棒性,发现自监督预训练阶段的泛化能力是模拟到现实迁移的强预测因子,并识别出离散潜在大小和训练序列长度是关键因素。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03603 2026-06-03 cs.CV cs.CL 93%

World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning

世界模型遇见语言模型:论具体推理与抽象推理的互补性

Yucheng Zhou, Wei Tao, Yiwen Guo, Jianbing Shen

机构 * Nanyang Technological University(南洋理工大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出受控具体推理框架及PF-OPSD方法,通过结合世界模型的视觉模拟与多模态大语言模型的抽象推理,在空间前瞻和开放域物理预测任务上提升性能与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02575 2026-06-02 cs.CV 93%

From Zero to Hero: Training-Free Custom Concept Spawning in World Models

从零到英雄:世界模型中的免训练自定义概念生成

Kiymet Akdemir, Pinar Yanardag

机构 * Virginia Tech(弗吉尼亚理工学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出SPAWN方法,利用图像到视频骨干网络的结构特性,通过交换参考帧锚点与外部概念潜变量,实现无需训练即可在世界模型中生成用户指定的视觉概念。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02372 2026-06-02 cs.AI cs.CL 93%

COMAP: Co-Evolving World Models and Agent Policies for LLM Agents

COMAP:面向LLM智能体的世界模型与智能体策略协同进化

Youwei Liu, Jian Wang, Hanlin Wang, Wenjie Li

机构 * Central South University(中南大学) College of Computer Science, Sichuan University(四川大学计算机学院) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出COMAP框架,通过闭环交互协同进化文本世界模型和智能体策略,在具身任务规划、网页导航和工具使用基准上显著提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01626 2026-06-02 cs.LG 93%

IMWM: Intuition Models Complement World Models for Latent Planning

IMWM:直觉模型补充世界模型用于潜在规划

Baoqi Gao, Ruize Han, Miao Wang, Song Wang

机构 * Beihang University(北航) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 针对基于潜在世界模型的规划中搜索瓶颈问题,提出IMWM框架,通过直觉模型与三个轻量组件协作,在四个像素级任务上显著提升成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31111 2026-06-01 cs.LG 93%

Subspace-Decomposed JEPAs: Disentangling Progression and Content in Latent World Models

子空间分解的JEPA:解耦潜在世界模型中的进展与内容

Lucas Thil, Jesse Read, Rim Kaddah, Guillaume Doquet

机构 * LIX, École Polytechnique(巴黎高等学院LIX实验室) IRT SystemX(系统X研究院) Safran Tech(萨弗兰科技)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出SD-JEPA方法,通过将JEPA潜在空间分解为正交的进展子空间和内容子空间,利用余弦边际三元组损失和SIGReg正则化分别约束,在控制基准上优于LeWM基线,并证明进展坐标可作为场景感知的指南针。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30542 2026-06-01 cs.AI 93%

Physically Viable World Models: A Case for Query-Conditioned Embodied AI

物理可行的世界模型:面向查询条件具身AI的案例

Adam J. Thorpe, Stepan Tretiakov, Cheng-Hsi Hsiao, Su Ann Low, Xingjian Li, Hassan Iqbal, Neel P. Bhatt, Ufuk Topcu, Krishna Kumar

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 针对具身AI中现有世界模型预测未来观测但物理不可行的问题,提出应构建基于查询条件、识别最简物理抽象的世界模型,通过模块化分解确保可解释性和可验证性。

Comments 21 pages; Adam J. Thorpe and Stepan Tretiakov contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11389 2026-05-29 cs.AI 93%

Causal-JEPA: Learning World Models through Object-Level Latent Masking

Causal-JEPA:通过对象级潜在掩码学习世界模型

Heejeong Nam, Quentin Le Lidec, Lucas Maes, Yann LeCun, Randall Balestriero

机构 * Brown University, GalilAI(布朗大学,GalilAI) New York University(纽约大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出C-JEPA,一种通过对象级潜在掩码扩展联合嵌入预测的对象中心世界模型,在视觉问答和智能体控制任务中分别提升反事实推理20%和仅用1%潜在特征实现高效规划。

Comments Project Page: https://hazel-heejeong-nam.github.io/cjepa/ ICML 2026 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29360 2026-05-29 cs.AI 93%

MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

MiraBench: 评估机器人世界模型中的动作条件可靠性

Tianzhuo Yang, Zihan Shen, Zirui Mi, Zhaoyi Zhang, Jiayi Zhou, Jiaming Ji, Juntao Dai, Jiawei Chen, Boyuan Chen, Yaodong Yang

机构 * Institute for Artificial Intelligence, Peking University(人工智能研究院,北京大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出MiraBench基准,通过物理一致性、动作跟随保真度和乐观偏差检测三个层次评估机器人世界模型的动作条件可靠性,发现视觉保真度不能反映动作保真度、模型规模扩大不保证动作跟随改善、乐观偏差普遍存在。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23345 2026-05-29 cs.CV 93%

SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models

SCOPE: 在可玩环境中模拟跨游戏操作以构建FPS世界模型

Zizhao Tong, Yeying Jin, Hongfeng Lai, Zeqing Wang, Zhaohu Xing, Kexu Cheng, Haoran Xu, Zhao Pu, Shangwen Zhu, Ruili Feng, Jian Zhao, Yan Zhang, Hao Tang, Ling Shao

机构 * UCAS-Terminus AI Lab, University of Chinese Academy of Sciences(中国科学院大学Terminus AI实验室) Tencent(腾讯) National University of Singapore(新加坡国立大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Waterloo(多伦多大学) Shanghai Jiaotong University(上海交通大学) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机科学学院多媒体信息处理国家重点实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出SCOPE方法,通过在每个Transformer块中插入条件模块,将特征重塑为逐像素时间序列,以分离FPS游戏中局部作用域(scope)内的操作效果与全局生成,并引入跨游戏数据集CrossFPS,实现零样本迁移。

Comments Project page: https://z2tong.github.io/SCOPE/. Code is available at https://github.com/z2tong/SCOPE

详情

展开后加载摘要…

URL PDF HTML 收藏