arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-07-02 至 2026-07-02 共收录 18 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 8 篇

2602.23164 2026-07-02 cs.LG 版本更新 94%

MetaOthello: A Controlled Study of Multiple World Models in Transformers

MetaOthello: Transformer中多个世界模型的控制研究

Aviral Chawla, Galen Hall, Juniper Lovato

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 通过引入共享语法但规则或分词不同的Othello变体套件,训练小型GPT处理混合数据,发现Transformer不划分容量为孤立子模型,而是收敛于跨变体因果转移的共享棋盘状态表示。

Comments Camera-ready version. Accepted to the 43rd International Conference on Machine Learning (ICML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00148 2026-07-02 cs.RO cs.CV 新提交 94%

3D Point World Models: Point Completion Enables More Accurate Dynamics Learning

3D点云世界模型:点云补全实现更准确的动力学学习

Skand Peri, Hung Nguyen, Chanho Kim, Li Fuxin, Stefan Lee

机构 * Oregon State University(俄勒冈州立大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出3D点云世界模型(3DPWM),通过先补全部分点云再学习动作条件动力学,实现长时程可靠推演和精确成本评估,支持多种机器人操作任务。

Comments 21 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00917 2026-07-02 cs.LG cs.AI 新提交 94%

Valdi: Value Diffusion World Models

Valdi: 价值扩散世界模型

Christopher Lindenberg, Kashyap Chitta

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出Valdi框架,结合端到端在线训练与潜在扩散动力学模型,在CarRacing环境中使用单步扩散达到与确定性MLP基线相当的性能,揭示了预测多模态与控制性能之间的权衡。

Comments RLC 2026 WMW

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01896 2026-07-02 cs.CV 版本更新 93%

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models

分而治之:多模态世界模型的解耦表示对齐

Junyuan Xiao, Dingkang Liang, Xin Zhou, Yixuan Ye, Tongtong Su, Guangmo Yi, Bin Xia, Qiang Lyu, Shurui Shi, Jun Huang, Jianlou Si, Wenming Yang

机构 * Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学) CSU(中南大学) ZJU(浙江大学) CUHK(香港中文大学) UCAS(中国科学院大学) Alibaba Group(阿里巴巴集团)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出M²-REPA方法,通过解耦扩散模型中间表示中的模态特定特征并与对应的专家基础模型对齐,实现多模态视频生成中多种基础模型先验的充分利用,显著提升视觉质量和长期一致性。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00673 2026-07-02 cs.RO 新提交 93%

Path Planning in Physically Viable World Models

物理可行世界模型中的路径规划

Su Ann Low, Cheng-Hsi Hsiao, Xingjian Li, Adam J. Thorpe, Ufuk Topcu, Krishna Kumar

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出物理可行世界模型,通过物理仿真生成修改后的场景,使机器人能在部署前评估未来地形变化对路径可行性的影响。

Comments 18 pages, 7 figures, submitted to CORL

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00457 2026-07-02 cs.AI 新提交 93%

Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments

面向演化环境中具身智能体的多尺度世界模型混合

Jinwoo Jang, Daniel J. Rho, Sihyung Yoon, Hyunsuk Cho, Honguk Woo

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出MuSix框架,通过尺度感知的世界模型混合与演化,解决具身智能体在动态环境中的多尺度推理和知识适应问题,在EmbodiedBench和HAZARD上超越现有方法。

Comments Accepted at ECCV 2026. 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00889 2026-07-02 cs.CV cs.AI 新提交 88%

DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors

DeWorldSG: 基于世界模型先验的深度感知3D语义场景图生成

Seok-Young Kim, Abdelrahman Elskhawy, Taewook Ha, Dooyoung Kim, Eunjae Shin, Benjamin Busam, Woontack Woo

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院) Technical University of Munich(慕尼黑工业大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) La Trobe University(拉筹伯大学)

专题命中 通用世界模型 :world-model(title);world-model(title);world model(abstract);world model(abstract)

AI总结 提出DeWorldSG框架,通过深度引导滤波估计实例级3D高斯分布,并利用世界模型先验聚合时空证据,生成时空鲁棒的3D语义场景图,在对象和谓词预测上达到最先进性能。

Comments 19 pages, 6 figures, ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00189 2026-07-02 cs.CV 新提交 81%

VOCA: Visual Odometry with Codec Awareness

VOCA: 具有编解码感知的视觉里程计

Nouri Alexander Hilscher, Mateo de Mayo, Dominik Muhle, Christoph Otten genannt Hermes, Daniel Cremers

机构 * Technical University of Munich(慕尼黑工业大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出VOCA方法,利用视频编解码信息改进压缩流上的立体视觉里程计跟踪性能,在因果VO上达到最先进的相对轨迹误差和效率。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频世界模型 2 篇

2607.00310 2026-07-02 cs.CV cs.AI 新提交 95%

RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail

RetailSMV:零售场景中基础视频世界模型的外视角与内视角适应

Amirreza Rouhi, Rajat Aggarwal, Parikshit Sakurikar, Anoop M. Namboodiri, Sashi P. Reddi

机构 * DreamVu

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world models(title);world model(title,abstract)

AI总结 研究零售场景下基础视频世界模型的外视角与内视角适应,发现仅用外视角数据训练的模型在多项指标上优于或等同于联合训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20083 2026-07-02 cs.CV 新提交 94%

Holo-World: Unified Camera, Object and Weather Control for Video World Model

Holo-World: 视频世界模型的统一相机、物体和天气控制

Xiangchen Yin, Wenzhang Sun, Jiahui Yuan, Zijie Liu, Yinda Chen, Wei Li, Dachun Kai, Chunfeng Wang, Xiaoyan Sun

机构 * University of Science and Technology of China(中国科学技术大学) Li Auto Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合国家科学中心人工智能研究院)

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 提出Holo-World,一种从单张图像联合控制相机、物体运动和天气的统一视频世界模型,通过场景适配器和解耦CFG实现世界保持与天气迁移。

Comments Project Page: https://xiangchenyin.github.io/Holo-World Code: https://github.com/XiangchenYin/Holo-World

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 具身与机器人 1 篇

2607.01166 2026-07-02 cs.RO cs.CV 新提交 62%

Structured 4D Latent Predictive Model for Robot Planning

结构化4D潜在预测模型用于机器人规划

Zhiyi Li, Peilin Wu, Xiaoshen Han, Ruojin Cai, Yilun Du

专题命中 具身与机器人 :predictive model(title,abstract);分类 cs.CV、cs.RO;predictive models(abstract)

AI总结 提出结构化4D潜在预测模型,在潜在空间中预测场景3D结构演化,实现3D一致的场景理解与机器人规划,在复杂操作任务中优于现有视频基规划器。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 自动驾驶 2 篇

2601.04453 2026-07-02 cs.CV 版本更新 90%

UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

UniDrive-WM:面向自动驾驶的统一理解、规划和生成世界模型

Zhexiao Xiong, Xin Ye, Burhan Yaman, Sheng Cheng, Yiren Lu, Jingru Luo, Nathan Jacobs, Liu Ren

机构 * Bosch Research North America & Bosch Center for Artificial Intelligence (BCAI)(博世北美研究院与博世人工智能中心) Washington University in St. Louis(圣路易斯华盛顿大学) Arizona State University(亚利桑那州立大学) Case Western Reserve University(凯斯西储大学)

专题命中 自动驾驶 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 提出UniDrive-WM,一个基于VLM的统一世界模型,联合执行场景理解、轨迹规划和未来图像生成,在Bench2Drive上轨迹误差降低7.3%,碰撞率降低10.4%。

Comments Accepted to ECCV 2026. Project Page: https://unidrive-wm.github.io/UniDrive-WM

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04198 2026-07-02 cs.CV cs.RO 版本更新 87%

DriveVA: Video Action Models are Zero-Shot Drivers

DriveVA: 视频动作模型是零样本驱动器

Mengmeng Liu, Diankun Zhang, Jiuming Liu, Jianfeng Cui, Hongwei Xie, Guang Chen, Hangjun Ye, Michael Ying Yang, Francesco Nex, Hao Cheng

机构 * University of Twente(特温特大学) Xiaomi EV(小米电动汽车) University of Cambridge(剑桥大学) University of Bath(巴斯大学)

专题命中 自动驾驶 :world model(abstract);world-model(abstract);driving world model(abstract);world model(abstract)

AI总结 DriveVA提出了一种新的自动驾驶世界模型,通过共享潜在生成过程联合解码未来视觉预测和动作序列,提升跨数据集和传感器配置的泛化能力,减少碰撞率和误差。

Comments Accepted to ECCV 2026. 30 pages, 12 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 模型式强化学习 5 篇

2601.14232 2026-07-02 cs.LG cs.AI cs.CV 版本更新 60%

KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning

KAGE-Bench:面向强化学习的已知轴视觉泛化快速评估

Egor Cherepanov, Daniil Zelezetsky, Alexey K. Kovalev, Aleksandr I. Panov

机构 * AXXX, Moscow, Russia(AXXX,莫斯科,俄罗斯) MIRAI, Moscow, Russia(MIRAI,莫斯科,俄罗斯)

专题命中 模型式强化学习 :latent dynamics(abstract);分类 cs.AI、cs.LG、cs.CV

AI总结 提出KAGE-Bench基准,通过解耦视觉轴独立评估像素策略在视觉分布偏移下的泛化能力,发现背景和光度偏移严重影响性能,而智能体外观偏移影响较小。

Comments 41 pages, 47 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00808 2026-07-02 cs.LG 新提交 56%

Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos

局部运动至关重要:一种用于从视频中进行强化学习预训练的解构-重组范式

Jinwen Wang, Youfang Lin, Xiaobo Hu, Shuo Wang, Kai Lv

机构 * Beijing Jiaotong University(北京交通大学) Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence(北京交通数据挖掘与具身智能重点实验室)

专题命中 模型式强化学习 :latent dynamics(abstract);分类 cs.LG;dynamics model(abstract)

AI总结 提出解构-重组范式(DRP),通过解构全局运动为原子动作学习局部运动表示,再重组以加速下游策略学习,在机器人控制任务中显著提升样本效率。

Comments 20 pages, 16 figures

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pages 9859-9868

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17519 2026-07-02 cs.LG cs.AI 版本更新 56%

TANDEM: Temporal Attention-guided Neural Differential Equations for Missingness in Time Series Classification

TANDEM:基于缺失数据的时序分类的时序注意力引导神经微分方程

YongKyung Oh, Dong-Young Lim, Sungil Kim, Alex Bui

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Ulsan National Institute of Science and Technology(蔚山科学技术院)

专题命中 模型式强化学习 :latent dynamics(abstract);分类 cs.AI、cs.LG

AI总结 本文提出TANDEM框架,通过注意力机制整合观测数据、插值控制路径和连续潜在动态,提升时序分类中缺失数据的处理能力,实验显示其优于现有方法。

Comments CIKM '25: Proceedings of the 34th ACM International Conference on Information and Knowledge Management. https://doi.org/10.1145/3746252.3760996

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01201 2026-07-02 cs.RO 新提交 53%

Sensorless Four-Channel Control Architecture Using Inverse Dynamics Modeling for Human-Scale Bilateral Teleoperation

基于逆动力学建模的无传感器四通道控制架构用于人尺度双边遥操作

Amir Noohian, Dylan Miller, Justin Valentine, Alan Lynch, Martin Jagersand

机构 * University of Alberta(阿尔伯塔大学)

专题命中 模型式强化学习 :dynamics model(title,abstract);分类 cs.RO

AI总结 针对人尺度遥操作中高惯性、建模困难和力传感器依赖问题,提出基于逆动力学的无传感器四通道架构,在WAM平台上验证,优于传统方案,提升位置/力跟踪并降低操作力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01022 2026-07-02 cs.LG 新提交 50%

Seahorse: A Unified Benchmarking Framework for Spatiotemporal Event Modeling

Seahorse: 时空事件建模的统一基准框架

Yahya Aalaila, Gerrit Großmann, Sebastian Vollmer

机构 * German Research Center for Artificial Intelligence (DFKI), Data Science and its Applications Research Group, Kaiserslautern, Germany(德国人工智能研究中心(DFKI),数据科学及其应用研究组,凯泽斯劳滕,德国) Department of Computer Science, Rhineland-Palatinate Technical University of Kaiserslautern-Landau (RPTU), Kaiserslautern, Germany(计算机科学系,莱茵兰-普法尔茨凯泽斯劳滕-兰道工业大学(RPTU),凯泽斯劳滕,德国)

专题命中 模型式强化学习 :latent dynamics(abstract);分类 cs.LG

AI总结 提出Seahorse统一框架,通过编码-演化-解码接口标准化神经时空点过程,实现公平比较与诊断分析,并引入合成压力测试套件揭示各模型族的归纳偏差。

Comments 24 pages, 9 figures. Code: https://github.com/YahyaAalaila/seahorse

详情

展开后加载摘要…

URL PDF HTML 收藏