arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 4308 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 4308 篇

0707.2199 2009-12-01 astro-ph 88%

Observational constraints on the braneworld model with brane-bulk energy exchange

M. Sadegh Movahed, Ahmad Sheykhi

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract)

Comments 17 pages and 18 figures, V2: Added comments, references, explained some topics related to the matter power spectrum as a robust constraint, accepted for publication in Mon. Not. R. Astron. Soc

Journal ref Mon. Not. R. Astron. Soc. 388, 197- 210 (2008)

详情

展开后加载摘要…

URL PDF HTML 收藏
gr-qc/0305031 2009-11-30 gr-qc 88%

Weak gravity in DGP braneworld model

Takahiro Tanaka

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract)

Comments 7 pages, 1 figure, references added

Journal ref Phys.Rev. D69 (2004) 024001

详情

展开后加载摘要…

URL PDF HTML 收藏
astro-ph/0305123 2009-11-30 astro-ph gr-qc hep-ph hep-th 88%

Second-order Perturbations of the Friedmann World Model

H. Noh, J. Hwang

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract)

Comments 61 pages; published version in Phys. Rev. D

Journal ref Phys. Rev. D 69 (2004) 104011

详情

展开后加载摘要…

URL PDF HTML 收藏
hep-th/0006184 2009-11-30 hep-th hep-ph 88%

A Brane World Model with Intersecting Branes

Matej Pavsic

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract)

Comments 14 pages

Journal ref Phys.Lett. A283 (2001) 8

详情

展开后加载摘要…

URL PDF HTML 收藏
astro-ph/9909150 2009-11-30 astro-ph gr-qc 88%

Conserved quantities in the perturbed Friedmann world model

J. Hwang

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract)

Comments 4 pages, no figure, Proceedings of The 2nd Seoul Workshop on Gravitation and Cosmology

Journal ref J.KoreanPhys.Soc.35:S633-S637,1999; J.KoreanPhys.Soc.35:S633-637,1999

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02846 2026-07-07 cs.AI 新提交 87%

Object-Centric Environment Modeling for Agentic Tasks

用于智能任务的以对象为中心的环境建模

Yiyang Li, Tianyi Ma, Zehong Wang, Yijun Ma, Yanfang Ye

机构 * University of Notre Dame(圣母大学)

专题命中 通用世界模型 :environment model(title,abstract);world model(abstract);world models(abstract);world model(abstract)

AI总结 研究大语言模型智能体经验维护难题,提出以对象为中心的环境建模(OCM),将经验组织成可执行模型,含对象知识与过程知识,在线更新并验证,实验表明其能提升智能体性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02365 2026-08-04 cs.AI cs.LG cs.RO 新提交 87%

Faster-WAM: Do World Action Models Need Deep Action Modules?

Faster-WAM:世界动作模型需要深度动作模块吗?

Liheng Ma, Rui Heng Yang, Zhanguang Zhang, Mateo Clemente, Ziwen Hu, Tongtong Cao, Yingxue Zhang

机构 * Huawei Noah’s Ark Lab(华为诺亚方舟实验室) Huawei Celia Team(华为Celia团队) Labs(2012实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 针对现有WAMs动作模块深度绑定视频骨干网络导致延迟高的问题,提出以视频为中心的DoT架构,构建Faster-WAM,实现低延迟、高性能与强泛化,较Fast-WAM提速3.2倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15407 2026-05-26 cs.AI cs.CV cs.LG 87%

IPR-1: Interactive Physical Reasoner

IPR-1:交互式物理推理器

Mingyu Zhang, Lifeng Zhuo, Tianxi Tan, Guocan Xie, Xian Nie, Yan Li, Renjie Zhao, Zizhu He, Ziyu Wang, Jiting Cai, Yong-Lu Li

机构 * CARNEGIE MELLON UNIVERSITY(卡内基梅隆大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 提出IPR模型,通过世界模型滚动评分和强化VLM策略,结合物理中心动作代码PhysCode,在1000+异构游戏基准上实现鲁棒的物理推理,性能超越GPT-5并零样本迁移至未见游戏。

Comments Accepted by CVPR 2026. 13 pages of main text and 20 pages of appendices. Project page: https://mybearyzhang.github.io/ipr-1

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05014 2026-04-08 cs.RO cs.AI cs.CV 87%

StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

StarVLA:一种积木式代码库,用于视觉-语言-动作模型开发

StarVLA Community

机构 * Von Neumann Institute, HKUST(香港科技大学冯·诺依曼研究所)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 StarVLA通过模块化架构、可重用训练策略和统一评估接口,解决VLA方法碎片化问题,提升可复现性和跨架构兼容性。

Comments Open-source VLA infra, Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20327 2026-03-24 cs.LG cs.AI cs.CV 87%

Probing the Latent World: Emergent Discrete Symbols and Physical Structure in Latent Representations

探测潜在世界:在潜在表示中涌现的离散符号与物理结构

Liu hung ming

机构 * PARRAWA AI(PARRAWA人工智能)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 本文提出AIM框架,通过被动量化探针探测V-JEPA 2潜在表示中的离散符号序列,揭示潜在空间的紧凑性及结构化符号 manifold 的可发现性。

Comments 35 pages, 6 figures, 3 tables, 26 equations; independent research report; Stage 1 of a four-stage AIM--V-JEPA 2 integration roadmap; code available at https://github.com/cyrilliu1974/JEPA

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04296 2025-02-07 cs.RO cs.CV cs.LG 87%

Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression

Lirui Wang, Kevin Zhao, Chaoqi Liu, Xinlei Chen

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

Comments Website: https://liruiw.github.io/hma/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27036 2026-07-30 cs.CV cs.LG 新提交 87%

Mitigating Compounding Error via Video Representation Regularization

通过视频表示正则化缓解复合误差

Taiye Chen, Qi Zhang, Yisen Wang

机构 * Peking University(北京大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 针对视频扩散世界模型自回归生成的复合误差问题,研究发现其与表示维度崩溃相关,提出视频表示正则化方法,在VBench指标上显著优于Diffusion Forcing,提升了长视频生成的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20867 2026-06-23 cs.CV cs.AI 新提交 87%

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation

FOCA: 面向未来的条件化用于数据高效的视觉-语言-动作适应

Duc Minh Nguyen, Nghiem Tuong Diep, Binh Gia Nguyen, Trong-Bao Ho, Doanh Le, Tan Q. Nguyen, Thien-Loc Ha, Nhiem Tran, Bao Thach, Nhat X. Tran, Tuan A. Tran, Artur Habuda, Philip Lund Møller, Tran Nguyen Le, Daniel Sonntag, Matthias Niepert, Khoa D. Doan, Vu Duong, Hung Ngo, Minh N. Vu, Duy M. H. Nguyen, An Thai Le, Ngo Anh Vien

机构 * Center for AI Research, VinUniversity, Vietnam University of Utah, USA German Research Center for Artificial Intelligence (DFKI) Technical University of Denmark, Denmark University of Oldenburg, Germany University of Stuttgart, Germany Max Planck Research School for Intelligent Systems (IMPRS-IS), Germany

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 提出FOCA框架,结合未来交互嵌入预测与目标观测隐式对齐,实现数据高效的VLA少样本适应,在LIBERO、RoboCasa和真实机器人上取得新最优结果。

Comments Accepted at ICML 2026. Project page: https://focavla.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20781 2026-06-23 cs.RO cs.CV 新提交 87%

World Action Models: A Survey

世界行动模型:综述

Qiuhong Shen, Shihua Zhang, Yue Liao, Qi Li, Zhenxiong Tan, Shizun Wang, Shuicheng Yan, Xinchao Wang

机构 * National University of Singapore(新加坡国立大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 本文综述世界行动模型(WAMs),通过两种互补视角组织现有工作,揭示其设计权衡与未来趋势。

Comments 57 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12688 2026-06-16 cs.LG cs.AI cs.DC 新提交 87%

M*: A Modular, Extensible, Serving System for Multimodal Models

M*: 一个模块化、可扩展的多模态模型服务系统

Atindra Jha, Naomi Sagan, Keisuke Kamahori, Irmak Sivgin, Rohan Sanda, Steven Gao, Mark Horowitz, Luke Zettlemoyer, Olivia Hsu, Jure Leskovec, Baris Kasikci, Stephanie Wang

机构 * Stanford University(斯坦福大学) University of Washington(华盛顿大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 提出M*系统,通过将模型表示为数据流图并引入Walk Graph抽象,支持多模态复合模型的高效服务,在多个任务上降低延迟并提升吞吐量。

Comments The codebase is available at https://github.com/mstar-project/mstar

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00090 2026-06-02 cs.RO cs.AI 87%

Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systems

物理AI中的静默故障:自主系统运行时动作授权的文献综述

Barak Or

机构 * STATE16

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 本文综述了物理AI系统中黑箱模型发出看似合理但实际错误的物理动作导致的静默故障问题,提出了运行时防护栏的分类和评估要求。

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08268 2026-05-12 cs.MA cs.AI 87%

Insider Attacks in Multi-Agent LLM Consensus Systems

多智能体大语言模型共识系统中的内部攻击

Xiaolin Sun, Zixuan Liu, Yibin Hu, Zizhan Zheng

机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,Location,Country) School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,Location,Country) Department of Computer Science, Tulane University, New Orleans, United States of America(计算机科学系, Tulane大学,新奥尔良,美国)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 研究多智能体大语言模型共识系统中的内部攻击问题,提出基于世界模型的框架,通过学习良性智能体的潜在行为状态并利用强化学习训练攻击者,有效降低共识率并延长分歧时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06679 2026-04-01 cs.AI cs.CV cs.GR 87%

MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines

MultiGen:扩散游戏引擎中可编辑多人世界的层级设计

Ryan Po, David Junhao Zhang, Amir Hertz, Gordon Wetzstein, Neal Wadhwa, Nataniel Ruiz

机构 * Stanford University(斯坦福大学) Google(谷歌)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 本文提出MultiGen,通过引入显式外部内存模块,实现可编辑多人世界中的环境控制与共享推理,提升交互性和一致性。

Comments Project page here: https://ryanpo.com/multigen/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16655 2025-11-21 cs.CV cs.LG 87%

Solving Spatial Supersensing Without Spatial Supersensing

无需空间超感知的解决方法

Vishaal Udandarao, Shyamgopal Karthik, Surabhi S. Nath, Andreas Hochlehnert, Matthias Bethge, Ameya Prabhu

机构 * Tübingen AI Center, University of Tübingen(图宾根人工智能中心,图宾根大学) Max Planck Institute for Biological Cybernetics(马克斯·普朗克生物 cybernetics 研究所)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 Cambrian-S通过定制推理策略在无需空间超感知的情况下解决视频世界模型基准。

Comments Tech Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02918 2025-09-22 cs.AI cs.LG 87%

World Modelling Improves Language Model Agents

Shangmin Guo, Omar Darwiche Domingues, Raphaël Avalos, Aaron Courville, Florian Strub

机构 * University of Edinburgh(爱丁堡大学) Cohere Vrije Universiteit Brussel(布鲁塞尔自由大学) Université de Montréal(蒙特利尔大学)

专题命中 通用世界模型 :world model(title);world model(title);environment model(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00259 2025-05-13 cs.RO cs.CV cs.GR 87%

One-Shot Real-to-Sim via End-to-End Differentiable Simulation and Rendering

Yifan Zhu, Tianyi Xiang, Aaron Dollar, Zherong Pan

机构 * Department of Mechanical Engineering and Materials Science, Yale University(机械工程与材料科学系,耶鲁大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

Comments 8 pages, 8 figures. Published at IEEE Robotics Automation Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01203 2024-07-16 cs.LG cs.AI 87%

Harnessing Discrete Representations For Continual Reinforcement Learning

Edan Meyer, Adam White, Marlos C. Machado

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

Comments 23 pages, 16 figures, accepted to RLC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14490 2026-08-17 cs.AI 新提交 86%

Twin: Playing an Unknown Game with a Test-Time Digital Twin

Twin:利用测试时数字孪生玩未知游戏

Alexy Skoutnev, Kirill Acharya, Gaston Longhitano, Madeleine Udell, Kevin Ellis, Iddo Drori

专题命中 通用世界模型 :world-model(abstract,abstract_cn);world-model(abstract,abstract_cn);world model(abstract);world model(abstract)

AI总结 该研究提出Twin系统,通过测试时数字孪生构建可执行世界模型,在ARC-AGI-3等未知游戏中通关率达97.8%,效率优于人类,核心是通过模拟交互和反例修复推断游戏规则与目标。

Comments Project website with action-by-action replays of all 25 runs: https://arc-agi-3-twin.vercel.app/ Code: AGI-3" target="_blank" rel="noopener">https://github.com/Alexyskoutnev/TWIN-ARC-AGI-3

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18171 2026-08-11 cs.LG 版本更新 86%

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

FlashRT:用于引导智能体部署实时多模态应用的智能体框架

Krish Agarwal, Zhuoming Chen, Yanyuan Qin, Zhenyu Gu, Atri Rudra, Beidi Chen

机构 * Carnegie Mellon University(卡内基梅隆大学) AMD(超威半导体公司) University at Buffalo(纽约州立大学水牛城分校)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 研究实时多模态应用部署难题,提出FlashRT智能体框架。通过新范式引导编码智能体多阶段转换,将参考实现转为高效部署,在不同GPU上显著提升性能,尤其在专家优化不成熟平台更具扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25333 2026-08-03 cs.CV 版本更新 86%

Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution

教会视频生成器记忆:为不可见状态演化引出动态记忆

Tianshuo Xu, Yichen Xie, Depu Meng, Chensheng Peng, Quentin Herau, Bo Jiang, Yihan Hu, Wei Zhan

机构 * Applied Intuition(应用直觉) University of California, Berkeley(加州大学伯克利分校)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 针对视频生成模型在观测中断时状态冻结的问题,提出ReMind框架,通过面向记忆的数据构建、事件感知训练和缓存适配,利用KV缓存机制实现动态记忆,在STEVO-Bench和恢复任务上取得最佳成绩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22098 2026-07-17 cs.CV 版本更新 86%

WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation

WorldWander:在视频生成中连接自我中心和外中心世界

Quanjian Song, Yiren Song, Kelly Peng, Yuan Gao, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 研究聚焦视频生成中自我中心与外中心世界转换,提出WorldWander框架,基于视频扩散变换器,整合上下文视角对齐与协作位置编码,精心策划数据集,实验证明该框架在视角同步、角色一致性和泛化能力上表现卓越,设定了新基准。

Comments Accepted by ECCV 2026; Code: https://github.com/showlab/WorldWander

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06291 2026-07-08 cs.CV cs.HC 新提交 86%

AlayaWorld: Long-Horizon and Playable Video World Generation

AlayaWorld:长期可玩视频世界生成

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao

机构 * Alaya Lab(阿莱亚实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 研究旨在解决游戏世界构建难题,核心方法是提出AlayaWorld全栈开源框架,可实现开放式实时交互,统一开发流程,主要贡献是为生成世界模型研究和应用奠定实践基础。

Comments Authors are listed alphabetically by the first name and their role. See the contribution section for details

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02517 2026-07-03 cs.CV 新提交 86%

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

WorldDirector: 构建具有持久动态记忆的可控世界模拟器

Hanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang, Qingyan Bai, Ka Leong Cheng, Yue Yu, Yixuan Li, Yihao Meng, Zichen Liu, Yanhong Zeng, Yujun Shen, Qifeng Chen

机构 * HKUST(香港科技大学) Ant Group(蚂蚁集团) ZJU(浙江大学) CUHK(香港中文大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 提出WorldDirector框架,通过LLM协调3D轨迹与相机运动作为视频生成控制信号,解耦语义运动编排与视觉生成,实现持久动态对象记忆和自由视角探索。

Comments Project Page: https://worlddirector.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22032 2026-07-03 cs.CV 版本更新 86%

Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving

Drive-JEPA:视频JEPA结合多模态轨迹蒸馏实现端到端驾驶

Linhan Wang, Zichong Yang, Chen Bai, Guoxiang Zhang, Xiaotong Liu, Xiaoyin Zheng, Xiao-Xiao Long, Chang-Tien Lu, Cheng Lu

机构 * Virginia Tech(弗吉尼亚理工学院) Purdue University(普渡大学) XPENG Motors(小鹏汽车) Nanjing University(南京大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 提出Drive-JEPA框架,结合视频联合嵌入预测架构(V-JEPA)与多模态轨迹蒸馏,通过自监督视频预训练和动量感知选择机制提升端到端驾驶的规划性能,在NAVSIM上达到新最优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27964 2026-06-29 cs.CV 新提交 86%

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control

Directing the World: 具有组合式人体-相机控制的快速自回归视频生成

Haoyuan Wang, Yabo Chen, Haibin Huang, Chi Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence, China Telecom (TeleAI)(中国电信人工智能研究院(TeleAI))

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 提出一种快速自回归框架,通过解耦控制学习并保持统一视频先验,实现人体运动和相机轨迹的组合控制,支持稳定长程生成与高视觉质量。

详情

展开后加载摘要…

URL PDF HTML 收藏