arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-08-19 至 2026-08-19 共收录 17 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 17 篇

2601.21282 2026-08-19 cs.CV 版本更新 94%

WorldBench: Benchmarking Physical Understanding of World Models by Isolating Physics Concepts

WorldBench: 为世界模型诊断评估进行物理辨析

Rishi Upadhyay, Howard Zhang, Jim Solomon, Ayush Agrawal, Yunhao Ba, Alex Wong, Celso M de Melo, Achuta Kadambi

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Sony AI(索尼人工智能) Yale University(耶鲁大学) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 WorldBench通过概念特定的解耦评估,提升世界模型的物理推理能力评估的准确性和可扩展性。

Comments Webpage: https://world-bench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17959 2026-08-19 cs.AI cs.LG 新提交 94%

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

面向基于神经符号世界模型的零样本任务迁移

Isidoro Tamassia, Lennert De Smet, Giuseppe Marra

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本研究提出一种神经符号世界模型,通过解耦观测重构与奖励预测,实现无需额外环境交互的零样本任务迁移,其泛化能力优于纯神经方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17163 2026-08-19 cs.LG cs.AI 新提交 94%

Q-Learning With World Models

结合世界模型的Q学习

Perry Dong, Yueru Jia, Chelsea Finn, Dorsa Sadigh

机构 * Stanford University(斯坦福大学) Peking University(北京大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本研究提出QWM框架,将世界模型与标准Q学习结合,在真实环境训练中避免复合模型偏差,在Robomimic和LIBERO基准上的样本效率与性能均优于现有SOTA方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17542 2026-08-19 cs.LG cs.AI 新提交 94%

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

无需高斯:用于JEPA世界模型的对比逆动力学

Jack Boylan, Chris Hokamp

机构 * Quantexa(昆泰克萨公司)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 该研究提出AC-MTM,用对比逆动力学替代JEPA世界模型的高斯型抗坍塌机制,在多目标视觉任务上性能优于SIGReg,且训练稳定无额外复杂组件

Comments 17 pages, 5 figures. Code: https://github.com/jackboyla/action-contrastive-jepa

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08713 2026-08-19 cs.AI cs.CV cs.RO 版本更新 93%

Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight

通过记忆增强的规划和前瞻性构建视觉导航的统一世界模型

Yifei Dong, Fengyi Wu, Guangyu Chen, Lingdong Kong, Qiyu Hu, Yuxuan Zhou, Xu Zhu, Jingdong Sun, Jun-Yan He, Qi Dai, Alexander G. Hauptmann, Zhi-Qi Cheng

机构 * University of Washington(华盛顿大学) National University of Singapore(新加坡国立大学) Apple(苹果公司) Microsoft Research(微软研究院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本文提出UniWM,一种整合视觉前瞻性与规划的统一世界模型,通过记忆机制提升导航鲁棒性和泛化能力,在多个基准测试中显著提升导航成功率。

Comments Accepted to ECCV 2026. 22 pages, 12 figures, code: https://github.com/UWMILab/UniWM

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17956 2026-08-19 cs.LG cs.AI cs.SY eess.SY 新提交 93%

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

被遗漏的模式即罕见规则:连续代码世界模型中的采样-验证危险定律

Javier Aguilar Martín

机构 * AGILabs

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 该研究揭示连续代码世界模型中采样-验证机制的危险,通过理论分析与实验发现,LLM合成的模式盲模型易被利用,接受仅能证明样本一致性,无法保证连续控制任务的性能。

Comments 92 pages, 5 figures. Code, data and result artifacts: https://github.com/JaviMaligno/code-world-models

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16955 2026-08-19 cs.MA cs.LG 新提交 90%

WONDER: A Radio World Model-based Negotiation Framework for Multi-Agent UAV Coverage Optimization

WONDER:一种基于无线电世界模型的多智能体无人机覆盖优化协商框架

Jiahao Huang, Rongpeng Li, Zhifeng Zhao, Guoru Ding, Honggang Zhang

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 针对灾后无线覆盖中断问题,提出基于JEPA无线电世界模型与PPO演员交替更新的WONDER框架,在含62个城市场景的RadioDynamics仿真中,其平衡得分0.870,覆盖优于STACCA且保持无人机全连通。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17739 2026-08-19 cs.MA 新提交 90%

Offline Multi-Agent Reinforcement Learning with a Physics-Informed World Model for Cooperative Mixed Traffic Control

面向协同混合交通控制的、融入物理信息世界模型的离线多智能体强化学习

Lu Liu, Chi Xie, Xi Xiong

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 本研究针对混合交通中部分可观测高速瓶颈的网联自动驾驶车辆协同控制问题,提出融入物理信息世界模型的离线多智能体强化学习框架,经SUMO实验验证可提升状态重构与预测准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17769 2026-08-19 eess.SP 新提交 90%

Electromagnetic World Model for 6G: A Unified Framework for Joint Environment Reconstruction and Channel Prediction

面向6G的电磁世界模型:环境重建与信道预测的统一框架

Yizhu Zhao, Li Yu, Jianhua Zhang, Yuxiang Zhang, Zhen Zhang, Guangyi Liu

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 该研究针对6G智能终端需同时实现环境感知与通信的需求,提出EMWM统一框架,融合多模态信息完成信道预测与环境重建,性能优于基线方法,具备鲁棒性与零样本泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18077 2026-08-19 cs.RO 新提交 88%

Hydra-0: Action Flow for Generalist World Modeling and Control

Hydra-0:面向通用世界建模与控制的动作流

Hongyu Li, Bowen Wen, Xinghao Zhu, Yixuan Wang, Yilun Du, Yunzhu Li, George Konidaris, Stan Birchfield, Soha Pouya, Chenran Li, Yan Chang

机构 * NVIDIA(英伟达) Brown University(布朗大学) Columbia University(哥伦比亚大学) Harvard University(哈佛大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.RO

AI总结 本研究提出Hydra-0通用世界模型,以动作流为条件实现跨实体、任务等的通用建模控制,在RoboLab基准获高相关性,还涌现出逆模式,可助力机器人控制。

Comments Project page: https://nvidia-isaac.github.io/video_to_data/hydra-0/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01372 2026-08-19 cs.LG cs.AI cs.CV 版本更新 83%

BRo-JEPA: Learning Modular Transformations in Latent Space

BRo-JEPA:在潜空间中学习模算术

Divyansh Jha, Yuanfang Xie, Brennen Yu, Varan Mehra

机构 * Georgia Institute of Technology(佐治亚理工学院) NYU Langone Health(纽约大学Langone医疗中心)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出BRo-JEPA模型,通过在潜空间中施加模10算术的循环结构,实现零样本泛化,解决了标准模型无法外推未见操作的问题。

Comments 20 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17095 2026-08-19 cs.CV 新提交 81%

Inference-Time Attention Steering for Vision-Language-Action Driving Models

面向视觉-语言-动作驾驶模型的推理时注意力引导

Darshan Nagendra Prasad, Lars Ullrich, Knut Graichen

机构 * FAU Erlangen-Nürnberg(埃尔朗根-纽伦堡大学)

专题命中 通用世界模型 :world model(abstract,abstract_cn);world model(abstract,abstract_cn);分类 cs.CV

AI总结 该研究针对VLA驾驶模型无法在推理时重定向注意力的问题,在Qwen3-VL主干上添加预softmax注意力偏差,实验验证了其对轨迹的引导效果及层相关特性。

Comments Attention Steering, Vision-Language-Action, AutonomousDriving, Inference-Time Intervention

Journal ref European Conference on Computer Vision 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13546 2026-08-19 cs.CV 版本更新 81%

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

Alaya-EVOKE:从线性缩放监督到无尽世界

Yuanyang Yin, Gongxuan Wang, Yifan Zhan, Chuanhao Li, Kaipeng Zhang, Feng Zhao

机构 * MoE Key Lab of BIPC(BIPC教育部重点实验室) USTC(中国科学技术大学) Shanghai Innovation Institute(上海创新研究院) Alaya Lab(Alaya实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 Alaya-EVOKE将持久世界状态外部化并重新设计教师模型,解决交互式世界模型的冲突需求,在WBench等基准上实现最优性能,支持开放式长时序生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11534 2026-08-19 cs.CV 版本更新 81%

Risk-Controllable Multi-View Diffusion for Driving Scenario Generation

可控风险多视角扩散用于驾驶场景生成

Hongyi Lin, Wenxiu Shi, Heye Huang, Dingyi Zhuang, Song Zhang, Yang Liu, Xiaobo Qu, Jinhua Zhao

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 RiskMV-DPO通过整合风险水平与物理风险建模,实现可控风险的多视角驾驶场景生成,提升3D检测性能并减少FID,推动安全导向的具身智能发展。

Comments 10 pages, 4 figures; accepted at the CVPR 2026 Workshop on Video Generative Models: Benchmarks and Evaluation (VGBE). Updated to the complete camera-ready version

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18005 2026-08-19 cs.CV 版本更新 81%

UrbanWorld2.0: A Multimodal Agentic Framework for Reality-Aligned 3D World Generation at City-Scale

RAISECity: 一种用于城市级现实对齐3D世界生成的多模态代理框架

Shengyuan Wang, Zhiheng Zheng, Yu Shang, Lixuan He, Yangcheng Yu, Fan Hangyu, Jie Feng, Qingmin Liao, Yong Li

机构 * College of AI, Tsinghua University(人工智能学院,清华大学) Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院,清华大学) Department of Electronic Engineering, BNRist, Tsinghua University(电子工程系,北京研究院,清华大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 RAISECity通过多模态代理框架实现城市级3D世界生成,提升现实对齐、精度和性能,适用于沉浸媒体和具身智能应用。

Comments Accepted by ACM MM 2026, the code is available at: https://github.com/tsinghua-fib-lab/UrbanWorld2.0

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20182 2026-08-19 cs.CL cs.AI 版本更新 69%

Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits

超越BFI:用于增强评估LLM人格特质可靠性与有效性的CSI

Huanhuan Ma, Haisong Gong, Xiaoyuan Yi, Xing Xie, Philip S. Yu, Dongkuan Xu

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.AI

AI总结 针对现有LLM人格特质评估工具BFI的可靠性与有效性局限,提出适配LLM的CSI,经实验验证其可靠性更高、与模型真实输出相关性超0.85,能有效评估LLM人格特质。

Comments Code available via https://github.com/dependentsign/CSI

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13113 2026-08-19 eess.SY cs.RO cs.SY 版本更新 50%

MPC for underactuated spacecraft control with a Lyapunov supervised physics-informed neural network correction layer

基于李雅普诺夫监督的物理信息神经网络校正层的欠驱动航天器MPC控制

Amirhossein Ayanmanesh Motlaghmofrad, Carlo Cena, Mauro Martini, Marcello Chiaberge

机构 * Politecnico di Torino(托斯纳理工学院) Argotec S.R.L.(Argotec公司)

专题命中 通用世界模型 :environment model(abstract);分类 cs.RO

AI总结 针对欠驱动航天器姿态控制,提出一种分层架构,结合非线性模型预测控制、物理信息神经网络和李雅普诺夫监督机制,在不确定性下降低稳态误差并保持鲁棒性。

Comments Accepted at SPAICE (AI in and for Space) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏