arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 21 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 自动驾驶 21 篇

2608.10386 2026-08-12 cs.LG cs.RO 新提交 94%

Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving

Dreamer-SAC:用于样本高效自动驾驶的潜世界模型离线策略学习

Jiazhuo Li, Linjiang Cao, Qi Liu, Xi Xiong

机构 * Tongji University(同济大学)

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出Dreamer-SAC框架,结合循环状态空间世界模型与离线策略SAC算法,在自动驾驶场景中优于DreamerV3、SAC等基线,且所需真实环境交互更少。

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14005 2026-07-16 cs.CV cs.RO 新提交 94%

M$^\text{4}$World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming

M$^\text{4}$World:用于交互式对象操纵和分钟级流的多视图多模态驾驶世界模型

Ke Cheng, Hanqiao Ye, Lei Shi, Yahui Liu, Yunhan Shen, Jingtao Dong, Zhenke Wang, Wenxuan Ao, Weixiang Xu, Kaining Huang, Shuhan Shen

专题命中 自动驾驶 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 针对现有驾驶世界生成方法局限,提出M$^\text{4}$World模型,通过灵活接口与多阶段训练实现对象操纵及长时流稳定,引入后训练与生成模型,并用新管道评估,实验证明其在驾驶模拟中有高质量、可控性与稳定性。

Comments 24 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15341 2026-06-16 cs.CV 新提交 93%

CausalDrive: Real-time Causal World Models for Autonomous Driving

CausalDrive: 用于自动驾驶的实时因果世界模型

Tianyi Yan, Huan Zheng, Dubing Chen, Meizhi Qu, Yingying Shen, Lijun Zhou, Mingfei Tu, Bing Wang, Guang Chen, Hangjun Ye, Haiyang Sun, Cheng-zhong Xu, Jianbing Shen

机构 * SKL-IOTSC, CIS, University of Macau(澳门大学协同创新研究院,科技学院) Xiaomi EV(小米汽车) CASIA(中国科学院自动化研究所)

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出CausalDrive,一种可控、实时的驾驶世界渲染器,通过因果预测和Context-Forced DMD架构实现交互式模拟,支持闭环评估、强化学习后训练和人在环仿真。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16274 2026-06-16 cs.CV 新提交 92%

GraphWorld: Long-Horizon Planning with World Models for End-to-End Autonomous Driving

GraphWorld: 基于世界模型的长时域规划实现端到端自动驾驶

Ziying Song, Caiyan Jia, Lin Liu, Lei Yang, Shengkai Zhang, Feiyang Jia, Fengda Zhao, Peiliang Wu, Shaoqing Xu, Chen Lv, Yadan Luo

机构 * Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院,交通数据挖掘与具身智能北京市重点实验室) School of Artificial Intelligence (School of Software), Yanshan University(燕山大学人工智能学院(软件学院)) School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与航空航天工程学院) University of Macau(澳门大学) The University of Queensland(昆士兰大学)

专题命中 自动驾驶 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 提出GraphWorld框架,通过潜在世界建模增强长时域规划,利用自车中心交互图建模邻车关系,并基于世界状态条件规划实现安全轨迹生成,显著降低碰撞率。

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29031 2026-08-03 cs.RO cs.AI 新提交 92%

Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving

Auto-JEPA:面向端到端自动驾驶的连续意图隐式世界模型

Jiwei Yang, Zhengxian Chen, Chaosheng Huang, Jun Li

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);world models(abstract);driving world model(abstract)

AI总结 Auto-JEPA是面向端到端自动驾驶的连续意图隐式世界模型,通过联合嵌入预测学习未来驾驶意图,无需密集未来世界建模,在NAVSIM数据集上取得优异规划性能,可聚焦规划相关视觉特征。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00298 2026-08-04 cs.AI 新提交 92%

WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation

WM-Cov:交互式世界模型式自动驾驶仿真的测试充分性

Jianxun Cui, Ping Wu, Stanisa Peric, Marko Milojkovic, Vladan Devedzic

机构 * School of Transportation Science and Engineering, Harbin Institute of Technology(哈尔滨工业大学交通科学与工程学院) Chongqing Research Institute of HIT(哈尔滨工业大学重庆研究院) Chongqing Changan Automobile Co., Ltd.(重庆长安汽车股份有限公司) University of Nis(尼什大学) University of Belgrade(贝尔格莱德大学)

专题命中 自动驾驶 :world-model(title,abstract);world-model(title,abstract);world model(abstract);world models(abstract)

AI总结 本文针对交互式世界模型式自动驾驶仿真测试的充分性问题,提出与提供方无关的 WM-Cov 评估层,经多组实验验证其可通过有效交互式证据收敛性更科学地评估测试效果。

Comments 11 pages, 6 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16354 2026-08-18 cs.AI cs.CV 新提交 91%

DriveCache: Action-Aware Caching for Driving World Model Inference

DriveCache:面向驾驶世界模型推理的动作感知缓存

Jianchun Yang, Jian Liang, Xianda Guo, Pinhan Fu, Yanlun Peng, Conglang Zhang, Wenke Huang, Mang Ye

专题命中 自动驾驶 :world model(title);driving world model(title);world model(title);driving world model(title)

AI总结 针对扩散驾驶生成器吞吐量受限的问题,提出动作感知的DriveCache控制器,利用规划运动与动态规划优化缓存,提升保真度-效率权衡,代码将公开。

Comments 9 pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20988 2026-07-24 cs.CV cs.AI 新提交 90%

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving

HyWorldVLA:一种用于自动驾驶的具有混合世界建模的视觉-语言-动作模型

Quanfu Yu, Xian Wu, Hao Xu, Liulong Ma

机构 * Automotive New Technology Research Institute, BYD Company Limited(比亚迪汽车新技术研究院)

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 研究针对自动驾驶中视觉-语言-动作模型的不足,提出HyWorldVLA框架,统一像素级监督与潜在表征学习。预训练阶段预测视频潜在并重建帧,微调阶段预测潜在特征生成轨迹,实验表明其性能优于基线,还建立了世界模型噪声鲁棒性评估新基准。

Comments 20 pages with 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10107 2026-08-12 cs.CV 新提交 88%

4D-WAM: 4D Consistent World Modeling for Autonomous Driving

4D-WAM:面向自动驾驶的4D一致世界建模

Jiacheng Fu, Yibo Yuan, Meng Tian, Yue Li, Jiangtong Zhu, Jianhua Han, Yueyi Zhang, Jianwu Fang, Jianru Xue, Hang Xu, Zhiwei Xiong

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);分类 cs.CV

AI总结 本文提出4D-WAM模型,通过几何基础模型的训练时监督与面向决策的时间步长采样策略,提升自动驾驶中世界-动作模型的4D场景一致性,在NAVSIM-v1、NAVSIM-v2基准上实现最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13410 2026-07-16 cs.RO 新提交 88%

Ego-Dynamics-Augmented World Model for Autonomous Driving with Zero-Shot Cross-Chassis Adaptation

用于零样本跨底盘自适应自动驾驶的自我动力学增强世界模型

Zhidong Wang, Jingsong Liang, Zirui Li, Zhan Chen, Han Yu, Chen Lv

机构 * School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与宇航工程学院) Collaborative Initiative, Interdisciplinary Graduate Programme, Nanyang Technological University(南洋理工大学跨学科研究生项目合作计划) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);分类 cs.RO

AI总结 研究针对自动驾驶中基于世界模型的强化学习问题,提出DynaDreamer方法,通过增强自我动力学先验改进世界模型,减少自我运动建模负担,实现零样本跨底盘自适应,实验证明该方法显著提升驾驶任务成功率。

Comments 13 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14058 2026-06-15 cs.RO 新提交 88%

ReactSim-Bench: Benchmarking Reactive Behavior World Model Simulation in Autonomous Driving

ReactSim-Bench:自动驾驶中反应性行为世界模型模拟的基准测试

Zhiyuan Zhang, Yanlun Peng, Jianing Zhang, Xianda Guo, Zehan Huang, Haoran Liu, Qifeng Li, Shaofeng Zhang, Xiaosong Jia, Junchi Yan

机构 * School of Computer Science & School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学计算机科学与技术学院、人工智能学院) Great Wall Motor(长城汽车) Institute of Trustworthy Embodied AI (TEAI), Fudan University(复旦大学可信具身人工智能研究所) School of Computer Science, Wuhan University(武汉大学计算机学院) University of Science and Technology of China(中国科学技术大学)

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);分类 cs.RO

AI总结 提出ReactSim-Bench,通过解耦自车与周围智能体控制,使用偏离日志的自车行为作为输入,评估行为世界模型模拟的反应性能力,并基于碰撞、地图和运动学指标系统评测多种模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11964 2026-07-15 cs.LG 新提交 88%

LIDAR-AD: A Decoder-Free Latent-Interaction Dreamer with Action-Residual Chains for Autonomous Driving

LIDAR-AD:一种用于自动驾驶的无解码器潜在交互梦想家与动作残差链

Yongzhi Liu, Yang Xiao, Zhong Cao, Zeng Kang, Sunan Zhang, Zhaozhi Dong, Guojun Yu, Weichao Zhuang

机构 * School of Mechanical Engineering, Southeast University(东南大学机械工程学院) Department of Civil and Environmental Engineering, University of Michigan(密歇根大学土木与环境工程系) Xheart Technology Co., Ltd.(芯驰科技有限公司)

专题命中 自动驾驶 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 研究针对自动驾驶中多源观测冗余问题,提出LIDAR-AD,用减少冗余的潜在对齐取代观测重建,将车辆控制建模为残差动作更新,经实验验证其在模拟场景和现实布局下性能优异,能提升风险感知等能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12987 2026-06-12 cs.CV cs.AI cs.LG cs.RO 新提交 83%

Diffusion Transformer World-Action Model for AV Scene Prediction

扩散Transformer世界-动作模型用于自动驾驶场景预测

Ruslan Sharifullin, Benjamin Jiang, Kai Xi Chew

机构 * Stanford University(斯坦福大学)

专题命中 自动驾驶 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出紧凑潜世界模型,结合扩散Transformer(DiT)预测未来场景,在nuScenes上实现4.8倍更好的KID,并实现动作可控性(转向ρ=0.81)。

Comments 10 pages, 9 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27504 2026-06-29 cs.CV 新提交 81%

ReWorld: Learning Better Representations for World Action Models

ReWorld:为世界动作模型学习更好的表示

Tianze Xia, Lijun Zhou, Kaixin Xiong, Jingfeng Yao, Yu Zhu, Zhenxin Zhu, Bing Wang, Guang Chen, Hangjun Ye, Wenyu Liu, Haiyang Sun, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学) Xiaomi EV(小米汽车)

专题命中 自动驾驶 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出ReWorld框架,通过优化中间表示(未来预测监督、跨模态对齐、难负样本监督)提升自动驾驶世界动作模型的规划性能,在nuScenes和NAVSIM上取得显著提升。

Comments 19 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22813 2026-06-23 cs.AI 新提交 81%

Active Inference as the Test-Time Scaling Law for Physical AI Agents

主动推理作为物理AI代理的测试时缩放定律

Omar Hashash, Christo Kurisummoottil Thomas, Walid Saad, Merouane Debbah, Karl Friston, Adeel Razi

机构 * Bradley Department of Electrical and Computer Engineering(布拉德利电气与计算机工程系) Institute for Advanced Computing(先进计算研究所) Virginia Tech(弗吉尼亚理工大学) Department of Electrical and Computer Engineering(电气与计算机工程系) Worcester Polytechnic Institute(沃思堡理工学院) Khalifa University of Science and Technology(卡利法科学与技术大学) CentraleSupelec(中央超导研究所) University Paris Saclay(巴黎萨克雷大学) Queen Square Institute of Neurology(伦敦大学学院神经科学研究所) School of Psychological Sciences(心理学系) Monash University(莫纳什大学) CIFAR Global Scholars Program(CIFAR全球学者计划) Queen Square Institute of Neurology, University College London(伦敦大学学院神经科学研究所)

专题命中 自动驾驶 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出基于主动推理第一性原理的测试时缩放定律,使物理AI代理通过世界模型推理在非平稳环境中泛化,并利用变分推理更新策略,在自动驾驶任务中提升36%以上推理效率。

Comments 53 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15820 2026-07-20 cs.SE 新提交 80%

In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing

掌控全局:关于自动驾驶系统测试现状的多公司研究

Qunying Song, Yuan Gao, Johannes Betz, Dietmar Pfahl, Mohammad Reza Mousavi, Federica Sarro

专题命中 自动驾驶 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 针对自动驾驶系统测试复杂且缺乏标准的问题,通过对九家公司专家访谈,经主题分析总结行业测试情况、挑战与趋势,提出以证据为中心的闭环测试框架,为ADS测试提供指导并指明未来方向。

Comments 34 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15620 2026-07-20 cs.RO cs.AI 新提交 71%

AEGIS: Assay-Aware Protocol Validation and Runtime Monitoring for Open-Source Liquid Handling Robots

AEGIS:用于开源液体处理机器人的检测感知协议验证和运行时监测

Priyanka V. Setty, Arvind Ramanathan, Ian Foster, Rick Stevens

机构 * Data Science and Learning Division, Argonne National Laboratory(阿贡国家实验室数据科学与学习部) Department of Computer Science, University of Chicago(芝加哥大学计算机科学系)

专题命中 自动驾驶 :world model(abstract);world model(abstract);分类 cs.AI、cs.RO

AI总结 研究开源液体处理机器人运行问题,提出AEGIS两层防护系统。一层结合检测规则库与大语言模型验证协议,二层用主成分分析模型监测运行轨迹,经实验验证有效,统一了检测感知验证与视觉监测,且开源。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08375 2026-07-10 cs.CV cs.AI 新提交 71%

WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving

WCog-VLA:用于端到端自动驾驶的双级世界认知视觉-语言-行动模型

Xuerun Yan, Zhexi Lian, Nuoheng Zhang, Shiyu Fang, Haoran Wang, Chen Lv, Jia Hu, Binyang Song

机构 * Tongji University(同济大学) Nanyang Technological University(南洋理工大学)

专题命中 自动驾驶 :world model(abstract);world model(abstract);分类 cs.AI、cs.CV

AI总结 针对现有视觉-语言-行动模型在自动驾驶中存在的局限,提出双级世界认知的WCog-VLA框架,语义层统一认知推理,生成层引入新模型加速推理,构建数据集,实验证明该模型在NAVSIM基准测试中达到最优分数。

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31209 2026-07-01 cs.AI cs.RO 新提交 71%

Long-term Traffic Simulation via Structured Autoregressive Modeling

基于结构化自回归建模的长期交通仿真

Lingyu Xiao, Zexin Feng, Xintao Yan

机构 * The University of Hong Kong(香港大学)

专题命中 自动驾驶 :world model(abstract);world model(abstract);分类 cs.AI、cs.RO

AI总结 提出RosettaSim框架,利用大规模序列模型的归纳偏置和统计先验,通过结构化自回归流建模场景拓扑、智能体状态和生成意图,实现短期准确与长期稳定仿真;并引入基于检索的交通评估(RTE)解决长期评估难题。

Comments ECCV 2026 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11120 2026-06-10 cs.AI cs.CV 新提交 71%

Monte Carlo Pass Search: Using Trajectory Generation for 3D Counterfactual Pass Evaluation in Football

蒙特卡洛传球搜索:利用轨迹生成进行足球3D反事实传球评估

Andrew Kang, Priya Narasimhan

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 自动驾驶 :world model(abstract);world model(abstract);分类 cs.AI、cs.CV

AI总结 提出蒙特卡洛传球搜索(MCPS),结合价值模型、世界模型和反事实策略,基于3D轨迹数据评估足球传球,通过两种执行盈余分数实现分布感知的传球分析。

Comments CVPR 2026, CVSports Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12990 2026-06-12 cs.LG 新提交 50%

Exposure Bias as Epistemic Underidentification in Recursive Forecasting

递归预测中的曝光偏差作为认知欠识别问题

Riku Green, Zahraa S. Abdallah, Telmo M Silva Filho

机构 * University of Bristol(布里斯托大学)

专题命中 自动驾驶 :latent dynamics(abstract);分类 cs.LG

AI总结 本文证明递归多步预测中的曝光偏差不仅是分布偏移,更是部分可观测性下的认知欠识别问题,并提出基于来源变量的误差分解与校正方法。

Comments Accepted for ICML 2026 EIML workshop

详情

展开后加载摘要…

URL PDF HTML 收藏