arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 140 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 自动驾驶 140 篇

2606.06014 2026-06-05 cs.AI cs.RO 95%

PLAN-S: Bridging Planning with Latent Style Dynamics for Autonomous Driving World Models

PLAN-S:通过潜在风格动态桥接规划以实现自动驾驶世界模型

Xiaoyun Qiu, Jingtao He, Yijie Chen, Yusong Huang, Haotian Wang, Yixuan Wang, Xinhu Zheng

机构 * Intelligent Transportation Thrust, Systems Hub, and Center of Seamless Connectivity & Connected Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(智能交通 thrust、系统中心及无缝连接与智能连接研究院,香港科学与技术大学(广州))

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);driving world model(title);world model(title,abstract)

AI总结 提出PLAN-S框架,通过从潜在表示解码风格条件语义成本图,解决自动驾驶中潜在世界模型规划的可控性问题,在nuScenes和NAVSIM上降低了碰撞率并提升了驾驶性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09059 2026-04-13 cs.CV cs.AI 94%

Learning Vision-Language-Action World Models for Autonomous Driving

学习视觉-语言-动作世界模型以实现自动驾驶

Guoqing Wang, Pin Tang, Xiangxuan Ren, Guodongfang Zhao, Bailan Feng, Chao Ma

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(上海交通大学人工智能研究院教育部人工智能重点实验室) Central Research Institute, Huawei(华为中央研究院)

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出VLA-World模型,通过结合预测想象与反思推理提升自动驾驶的前瞻性。该模型利用生成的轨迹引导图像生成,并通过反思优化轨迹预测,实验表明其在规划和未来场景生成任务中优于现有方法。

Comments Accepted by CVPR2026 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14696 2026-05-15 cs.CV 94%

EponaV2: Driving World Model with Comprehensive Future Reasoning

EponaV2:通过全面的未来推理驱动世界模型

Jiawei Xu, Zhizhou Zhong, Zhijian Shu, Mingkai Jia, Mingxiao Li, Jia-Wang Bian, Qian Zhang, Kaicheng Zhang, Jin Xie, Jian Yang, Wei Yin

机构 * PCA Lab, VCIP, College of Computer Science, Nankai University(PCA实验室、VCIP、计算机科学学院、南开大学) Horizon Robotics HKUST(香港科技大学) NJUPT(南京工程大学) NTU(国立台湾大学) Anyverse School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院、南京大学)

专题命中 自动驾驶 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 本文提出EponaV2,一种新的驾驶世界模型范式,通过全面的未来推理实现高质量规划。模型通过预测更全面的未来表示,结合3D和语义模态,提升环境理解与现实推理能力,从而改进轨迹规划。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28196 2026-05-01 cs.CV 94%

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation

HERMES++:迈向统一的驾驶世界模型用于3D场景理解和生成

Xin Zhou, Dingkang Liang, Xiwu Chen, Feiyang Tan, Dingyuan Zhang, Hengshuang Zhao, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Mach Drive University of Hong Kong(香港大学)

专题命中 自动驾驶 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 本文提出HERMES++,一种统一的驾驶世界模型,整合3D场景理解和未来几何预测。通过BEV表示、LLM增强世界查询和当前到未来链接等设计,提升驾驶场景的生成与理解能力。

Comments Extended version of ICCV 25 paper HERMES, Code: https://github.com/H-EmbodVis/HERMESV2, Project page: https://h-embodvis.github.io/HERMESV2/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14729 2025-08-14 cs.CV 94%

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation

Xin Zhou, Dingkang Liang, Sifan Tu, Xiwu Chen, Yikang Ding, Dingyuan Zhang, Feiyang Tan, Hengshuang Zhao, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) MEGVII Technology(梅格维七科技) Mach Drive(马奇驱动) The University of Hong Kong(香港大学)

专题命中 自动驾驶 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

Comments Accepted by ICCV 2025. The code is available at https://github.com/LMD0311/HERMES

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24004 2026-05-26 cs.AI cs.CV cs.LG cs.RO 94%

Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving

推理--想象--行动:基于世界模型的闭环LLM自动驾驶决策

Zhengqi Sun, Yiwen Sun, Boxuan Liu, Tailai Chen, Tianxu Guo, Jiabin Liu

机构 * 1Department of Information Management, Peking University, Beijing 100871, China 2School of Intelligence Science Technology, Peking University, Beijing 100871, China 3State Key Laboratory of General Artificial Intelligence, BIGAI, Beijing 100080, China 4Yuanpei College, Peking University, Beijing 100871, China 5China Agricultural University, Beijing, China 6CRSC Research \& Design Institute Group Co., Ltd., Beijing, China

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出Reason--Imagine--Act (RIA)闭环框架,结合LLM推理器与动作条件世界模型进行在线安全验证,在CARLA点目标协议下实现80.05%路线完成率、51.10%到达率和0.20%碰撞率。

Comments Accepted by the 2026 IEEE International Conference on Intelligent Transportation Systems (ITSC 2026). 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13840 2026-08-20 cs.RO cs.CV 版本更新 94%

Multi-Agent Embodied Autonomous Driving (MAEAD): From V2X Information Exchange to Shared World Models

多智能体具身自动驾驶:从V2X信息交换到共享世界模型

Senkang Hu, Zhengru Fang, Yihang Tao, Zihan Fang, Yiqin Deng, Yuguang Fang

机构 * Lingnan University, Hong Kong(岭南大学(香港))

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文综述了从单车智能向多智能体具身系统转变的自动驾驶技术,通过共享世界模型实现感知共享、意图推断和协同规划,并指出了在仿真评估、实时安全保证等方面的研究空白。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10386 2026-08-12 cs.LG cs.RO 新提交 94%

Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving

Dreamer-SAC:用于样本高效自动驾驶的潜世界模型离线策略学习

Jiazhuo Li, Linjiang Cao, Qi Liu, Xi Xiong

机构 * Tongji University(同济大学)

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出Dreamer-SAC框架,结合循环状态空间世界模型与离线策略SAC算法,在自动驾驶场景中优于DreamerV3、SAC等基线,且所需真实环境交互更少。

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14005 2026-07-16 cs.CV cs.RO 新提交 94%

M$^\text{4}$World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming

M$^\text{4}$World:用于交互式对象操纵和分钟级流的多视图多模态驾驶世界模型

Ke Cheng, Hanqiao Ye, Lei Shi, Yahui Liu, Yunhan Shen, Jingtao Dong, Zhenke Wang, Wenxuan Ao, Weixiang Xu, Kaining Huang, Shuhan Shen

专题命中 自动驾驶 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 针对现有驾驶世界生成方法局限,提出M$^\text{4}$World模型,通过灵活接口与多阶段训练实现对象操纵及长时流稳定,引入后训练与生成模型,并用新管道评估,实验证明其在驾驶模拟中有高质量、可控性与稳定性。

Comments 24 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01536 2026-02-03 cs.RO cs.CV 94%

UniDWM: Towards a Unified Driving World Model via Multifaceted Representation Learning

UniDWM: 通过多维表征学习实现统一的驾驶世界模型

Shuai Liu, Siheng Ren, Xiaoyao Zhu, Quanmin Liang, Zefeng Li, Qiang Li, Xin Hu, Kai Huang

机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家) School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家) Xpeng Motors Technology Co Ltd(小鹏汽车科技有限公司)

专题命中 自动驾驶 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 UniDWM通过多维表征学习实现统一驾驶世界模型,提升自动驾驶中的轨迹规划和4D重建能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15341 2026-06-16 cs.CV 新提交 93%

CausalDrive: Real-time Causal World Models for Autonomous Driving

CausalDrive: 用于自动驾驶的实时因果世界模型

Tianyi Yan, Huan Zheng, Dubing Chen, Meizhi Qu, Yingying Shen, Lijun Zhou, Mingfei Tu, Bing Wang, Guang Chen, Hangjun Ye, Haiyang Sun, Cheng-zhong Xu, Jianbing Shen

机构 * SKL-IOTSC, CIS, University of Macau(澳门大学协同创新研究院,科技学院) Xiaomi EV(小米汽车) CASIA(中国科学院自动化研究所)

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出CausalDrive,一种可控、实时的驾驶世界渲染器,通过因果预测和Context-Forced DMD架构实现交互式模拟,支持闭环评估、强化学习后训练和人在环仿真。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09701 2026-05-12 cs.CV 93%

DriveFuture: Future-Aware Latent World Models for Autonomous Driving

DriveFuture: 用于自动驾驶的面向未来的潜在世界模型

Yufeng Hong, Xiaotian Zhou, Yingyan Li, Xiangpo Zhou, Lin Liu, Yadan Luo, Shaoqing Xu, Lei Yang, Ziying Song

机构 * Beijing Institute of Technology(北京理工大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beihang University(北航) Beijing Jiaotong University(北京交通大学) The University of Queensland(昆士兰大学) University of Macau(澳门大学) Nanyang Technological University(南洋理工大学) School of Artificial Intelligence ( School of Software), Yanshan University(燕山大学人工智能学院(软件学院))

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出DriveFuture,一种面向未来的潜在世界建模框架,通过将未来世界状态条件化于当前潜在状态建模过程,提升自动驾驶轨迹规划性能。

Comments 24pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00969 2026-04-02 cs.CV 93%

DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving

DLWM:双潜在世界模型实现自动驾驶中的整体高斯中心预训练

Yiyao Zhu, Ying Xue, Haiming Zhang, Guangfeng Jiang, Wending Zhou, Xu Yan, Jiantao Gao, Yingjie Cai, Bingbing Liu, Zhen Li, Shaojie Shen

机构 * HKUST(香港科技大学) CUHK-SZ(香港中文大学(深圳)) USTC(中国科学技术大学) Huawei Foundation Model Department(华为基础模型部门)

专题命中 自动驾驶 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出DLWM,通过双潜在世界模型实现自动驾驶中的整体高斯中心预训练,提升3D占用感知、4D占用预测和运动规划性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12751 2025-12-16 cs.CV 93%

GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation

GenieDrive: 向具有物理意识的驾驶世界模型迈进:基于4D占用的视频生成

Zhenya Yang, Zhe Liu, Yuxiang Lu, Liping Hou, Chenxuan Miao, Siyi Peng, Bailan Feng, Xiang Bai, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Huazhong University of Science and Technology(华中科技大学)

专题命中 自动驾驶 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 GenieDrive通过4D占用引导的视频生成,实现物理意识的驾驶视频生成,提升预测精度和视频质量。

Comments The project page is available at https://huster-yzy.github.io/geniedrive_project_page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19631 2026-05-20 cs.RO cs.CV 93%

HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models

HEAT: 基于轨迹引导的世界模型实现异构端到端自动驾驶

Hoonhee Cho, Giwon Lee, Jae-Young Kang, Hyemin Yang, Heejun Park, Kuk-Jin Yoon

机构 * KAIST(韩国科学技术院)

专题命中 自动驾驶 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本文提出一种基于轨迹引导的学习方法,通过规划轨迹组织训练,使模型能够捕捉驾驶意图的领域不变表示,并结合预测未来潜在特征的世界模型,提高特征一致性并缓解领域偏见,从而在多个异构数据集上实现强性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16377 2026-03-31 cs.RO cs.AI 93%

VLM-SAFE: Vision-Language Model-Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving

VLM-SAFE: 基于世界模型的视觉-语言模型引导的安全强化学习用于自动驾驶

Yansong Qu, Zilin Huang, Zihao Sheng, Jiancong Chen, Yue Leng, Samuel Labi, Sikai Chen

机构 * Lyles School of Civil and Construction Engineering, Purdue University(普渡大学莱尔斯土木与建筑工程学院) Department of Civil and Environmental Engineering, University of Wisconsin-Madison(威斯康星大学麦迪逊分校土木与环境工程系) Google(谷歌)

专题命中 自动驾驶 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本文提出VLM-SAFE框架,通过观察-想象-评估-行动闭环,利用视觉语言模型提供语义安全信号,结合世界模型预测未来轨迹,优化策略以提升自动驾驶的安全性和效率。

Comments N/A

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16274 2026-06-16 cs.CV 新提交 92%

GraphWorld: Long-Horizon Planning with World Models for End-to-End Autonomous Driving

GraphWorld: 基于世界模型的长时域规划实现端到端自动驾驶

Ziying Song, Caiyan Jia, Lin Liu, Lei Yang, Shengkai Zhang, Feiyang Jia, Fengda Zhao, Peiliang Wu, Shaoqing Xu, Chen Lv, Yadan Luo

机构 * Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院,交通数据挖掘与具身智能北京市重点实验室) School of Artificial Intelligence (School of Software), Yanshan University(燕山大学人工智能学院(软件学院)) School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与航空航天工程学院) University of Macau(澳门大学) The University of Queensland(昆士兰大学)

专题命中 自动驾驶 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 提出GraphWorld框架,通过潜在世界建模增强长时域规划,利用自车中心交互图建模邻车关系,并基于世界状态条件规划实现安全轨迹生成,显著降低碰撞率。

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29031 2026-08-03 cs.RO cs.AI 新提交 92%

Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving

Auto-JEPA:面向端到端自动驾驶的连续意图隐式世界模型

Jiwei Yang, Zhengxian Chen, Chaosheng Huang, Jun Li

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);world models(abstract);driving world model(abstract)

AI总结 Auto-JEPA是面向端到端自动驾驶的连续意图隐式世界模型,通过联合嵌入预测学习未来驾驶意图,无需密集未来世界建模,在NAVSIM数据集上取得优异规划性能,可聚焦规划相关视觉特征。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19505 2026-08-18 cs.CV 版本更新 92%

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

DrivingWorld:通过视频GPT构建自动驾驶领域的世界模型

Xiaotao Hu, Mingkai Jia, Xiaoyang Guo, Qian Zhang, Xiao-xiao Long, Wei Yin

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);world models(abstract);driving world model(abstract)

AI总结 本文提出名为DrivingWorld的自动驾驶GPT风格世界模型,通过时空融合等策略提升视频生成质量与时长,实现更优的可控未来视频生成效果。

Journal ref ICPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00298 2026-08-04 cs.AI 新提交 92%

WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation

WM-Cov:交互式世界模型式自动驾驶仿真的测试充分性

Jianxun Cui, Ping Wu, Stanisa Peric, Marko Milojkovic, Vladan Devedzic

机构 * School of Transportation Science and Engineering, Harbin Institute of Technology(哈尔滨工业大学交通科学与工程学院) Chongqing Research Institute of HIT(哈尔滨工业大学重庆研究院) Chongqing Changan Automobile Co., Ltd.(重庆长安汽车股份有限公司) University of Nis(尼什大学) University of Belgrade(贝尔格莱德大学)

专题命中 自动驾驶 :world-model(title,abstract);world-model(title,abstract);world model(abstract);world models(abstract)

AI总结 本文针对交互式世界模型式自动驾驶仿真测试的充分性问题,提出与提供方无关的 WM-Cov 评估层,经多组实验验证其可通过有效交互式证据收敛性更科学地评估测试效果。

Comments 11 pages, 6 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00603 2025-07-02 cs.CV 92%

World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model

Yupeng Zheng, Pengxuan Yang, Zebin Xing, Qichao Zhang, Yuhang Zheng, Yinfeng Gao, Pengfei Li, Teng Zhang, Zhongpu Xia, Peng Jia, Dongbin Zhao

机构 * CASIA(中国科学院自动化研究所) Li Auto(力汽车) PCL(鹏城实验室) NUS(新加坡国立大学) Tsinghua(清华大学)

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);world models(abstract);driving world model(abstract)

Comments ICCV 2025, first version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12560 2026-06-16 cs.CV cs.LG cs.RO 版本更新 91%

CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving

CoIRL-AD:面向自动驾驶的潜在世界模型中的协作-竞争模仿-强化学习

Xiaoji Zheng, Ziyuan Yang, Yanhao Chen, Yuhang Peng, Yuanrong Tang, Gengyuan Liu, Bokui Chen, Jiangtao Gong

机构 * University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学)

专题命中 自动驾驶 :world model(title);world models(title);world model(title);world models(title)

AI总结 提出CoIRL-AD框架,通过解耦模仿学习与强化学习、利用潜在世界模型进行长时程奖励估计以及引入竞争机制,在离线训练中提升自动驾驶的鲁棒性,尤其在跨城市泛化和长尾场景中表现优异。

Comments 19 pages, 22 figures, ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16354 2026-08-18 cs.AI cs.CV 新提交 91%

DriveCache: Action-Aware Caching for Driving World Model Inference

DriveCache:面向驾驶世界模型推理的动作感知缓存

Jianchun Yang, Jian Liang, Xianda Guo, Pinhan Fu, Yanlun Peng, Conglang Zhang, Wenke Huang, Mang Ye

专题命中 自动驾驶 :world model(title);driving world model(title);world model(title);driving world model(title)

AI总结 针对扩散驾驶生成器吞吐量受限的问题,提出动作感知的DriveCache控制器,利用规划运动与动态规划优化缓存,提升保真度-效率权衡,代码将公开。

Comments 9 pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08719 2026-04-13 cs.CV cs.AI cs.RO 91%

LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving

LMGenDrive: 联合多模态理解与生成世界建模以实现端到端驾驶

Hao Shao, Letian Wang, Yang Zhou, Yuxuan Hu, Zhuofan Zong, Steven L. Waslander, Wei Zhan, Hongsheng Li

机构 * CUHK MMLab(香港中文大学多媒体实验室) University of Toronto(多伦多大学) UC Berkeley(加州大学伯克利分校)

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出LMGenDrive框架,结合LLM多模态理解与生成世界模型,提升自动驾驶的闭环性能,通过视频预测和控制信号生成,增强时空场景建模和指令遵循能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09627 2024-12-13 cs.CV cs.AI cs.LG 91%

Doe-1: Closed-Loop Autonomous Driving with Large World Model

Wenzhao Zheng, Zetian Xia, Yuanhui Huang, Sicheng Zuo, Jie Zhou, Jiwen Lu

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);driving world model(abstract);driving world model(abstract)

Comments Code is available at: https://github.com/wzzheng/Doe

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20988 2026-07-24 cs.CV cs.AI 新提交 90%

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving

HyWorldVLA:一种用于自动驾驶的具有混合世界建模的视觉-语言-动作模型

Quanfu Yu, Xian Wu, Hao Xu, Liulong Ma

机构 * Automotive New Technology Research Institute, BYD Company Limited(比亚迪汽车新技术研究院)

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 研究针对自动驾驶中视觉-语言-动作模型的不足,提出HyWorldVLA框架,统一像素级监督与潜在表征学习。预训练阶段预测视频潜在并重建帧,微调阶段预测潜在特征生成轨迹,实验表明其性能优于基线,还建立了世界模型噪声鲁棒性评估新基准。

Comments 20 pages with 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10564 2026-05-12 cs.CV cs.RO 90%

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving

DeepSight: 通过潜在状态预测实现长视距世界建模的端到端自动驾驶

Lingjun Zhang, Changjie Wu, Linzhe Shi, Jiangyang Li, Jiaxin Liu, Lei Yang, Hang Zhang, Mu Xu, Hong Wang

机构 * Tsinghua University(清华大学) Amap, Alibaba Group(阿里巴巴集团Amap) Nanyang Technological University(南洋理工大学)

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);driving world model(abstract);driving world model(abstract)

AI总结 本文提出通过鸟瞰图空间预测连续未来帧的潜在语义特征,实现长视距世界建模,并引入高效适应性文本推理机制提升复杂场景下的驾驶性能。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03454 2026-05-11 cs.CV cs.AI 90%

Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles

在驾驶前思考:基于世界模型的多模态接地用于自动驾驶车辆

Haicheng Liao, Huanming Shen, Bonan Wang, Yongkang Li, Yihong Tang, Chengyue Wang, Dingyi Zhuang, Kehua Chen, Hai Yang, Chengzhong Xu, Zhenning Li

机构 * University of Macau(澳门大学) UESTC(电子科技大学) Purdue University(普渡大学) McGill University(麦吉尔大学) Massachusetts Institute of Technology(麻省理工学院) University of Washington(华盛顿大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文提出ThinkDeeper框架,通过预测未来空间状态提升自动驾驶车辆的自然语言指令理解能力,结合超图引导解码器融合多模态输入,提出DrivePilot数据集并在多个基准测试中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24587 2026-04-02 cs.LG cs.RO 90%

DreamerAD: Efficient Reinforcement Learning via Latent World Model for Autonomous Driving

DreamerAD:通过潜在世界模型实现高效的强化学习用于自动驾驶

Pengxuan Yang, Yupeng Zheng, Deheng Qian, Zebin Xing, Qichao Zhang, Linbo Wang, Yichen Zhang, Shaoyu Guo, Zhongpu Xia, Qiang Chen, Junyu Han, Lingyun Xu, Yifeng Pan, Dongbin Zhao

机构 * Institute of Automation, CAS(中国科学院自动化研究所) Chongqing Chang’an Technology Co., Ltd(重庆长安科技有限公司) School of Advanced Interdisciplinary Sciences, UCAS(中国科学院大学先进交叉科学学院) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院)

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 DreamerAD通过压缩扩散采样将步骤从100步减少到1步,实现80倍加速并保持视觉可解释性,解决了自动驾驶中真实世界数据训练成本高和安全风险大的问题。

Comments authors update

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27287 2026-03-31 cs.RO cs.CV 90%

Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving

Uni-World VLA:自主驾驶中的交织世界建模与规划

Qiqi Liu, Huan Xu, Jingyu Li, Bin Sun, Zhihui Hao, Dangen She, Xiatian Zhu, Li Zhang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Li Auto Inc.(理想汽车) University of Surrey(萨里大学)

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 本文提出Uni-World VLA模型,通过交织未来帧预测与轨迹规划提升自主驾驶决策能力,结合单目深度信息增强场景预测。

Comments 22 pages, 8 figures. Submitted to ECCV 2026. Code will be released

详情

展开后加载摘要…

URL PDF HTML 收藏