arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-08-06 至 2026-08-06 共收录 6 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 6 篇

2608.04653 2026-08-06 cs.CV cs.RO 新提交 94%

Overcoming Statistical Bias in Action-Controllable World Models

克服动作可控世界模型中的统计偏差

Yuhong Shi, Zhenhao Chu, Jie Wei, Jun Hao, Jianyi Liu, Jingwen Fu

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 该研究针对动作可控世界模型的统计偏差问题,提出CoCo反事实一致性框架,结合ARC、DE评估指标及Mini-SSMB数据集,在多项任务上提升了动作可控性与视频预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04378 2026-08-06 cs.SD cs.LG eess.AS 新提交 92%

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation

助力音乐协同创作智能体“良好聆听”:用于理解与生成的分层自监督世界模型

Scott H. Hawley

机构 * Belmont University(贝尔蒙特大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本研究提出分层自监督世界模型,通过Swin V2编码器与条件流匹配模型构建协同音乐创作智能体,提升了和弦与调式检测准确率,可快速生成音乐建议并支持交互演示。

Comments 20 pages, 14 figures. A 6-page version was submitted to the NeurIPS 2026 Creative AI Track. Supplemental website with listening examples: https://drscotthawley.github.io/midi-rae-jepa-son/. Live demo: https://drscotthawley-midi-rae-jepa-son.hf.space/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02150 2026-08-06 cs.CV cs.AI 版本更新 82%

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs

PhyCheck:面向视频大语言模型的物理规律理解细粒度证据驱动数据集

Zhongjie Ba, Shengwang Xu, Peng Cheng, Jinyang Zou, Ting Yu, Zhibo Wang, Zhan Qin

机构 * Zhejiang University(浙江大学) ZJU-Hangzhou Global Scientific and Technological Innovation Center(浙江大学杭州国际科技创新中心) Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出PhyCheck数据集,通过粗粒度、细粒度及诊断子集提升视频大语言模型的物理规律理解,实验证实其训练效果,同时指出当前模型难以整合因果条件的问题。

Comments 15pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04412 2026-08-06 cs.CV 新提交 81%

muSync-GS: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Hazards

muSync-GS:面向天气与几何道路危险场景的物理同步驾驶视频合成方法

Yang Chen, Yicheng Zhu, Tao Li, Zilin Bian

机构 * Rochester Institute of Technology(罗切斯特理工学院) City University of Hong Kong(香港城市大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 muSync-GS是一种物理同步驾驶视频合成框架,通过耦合天气、道路几何编辑与车辆动力学,在12个CarSim案例上实现了车辆状态的精准预测与场景同步。

Comments 42 pages, 14 figures; includes an appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04996 2026-08-06 cs.RO 新提交 69%

DreamWAM: Beyond RGB Future Prediction for World Action Models

DreamWAM:超越RGB未来预测的世界动作模型

Shanglin Yuan, Weiheng Zhao, Xin Shi, Haoyi Jiang, Xianda Guo, Liu Liu, Wenyu Liu, Wei Sui, Xinggang Wang

专题命中 通用世界模型 :world model(abstract);world model(abstract);分类 cs.RO

AI总结 DreamWAM通过结构化超越RGB的世界建模,在LIBERO及真实操作任务中提升了世界动作模型的鲁棒性与成功率,相关代码模型已公开。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21030 2026-08-06 eess.SY cs.AI cs.RO cs.SY math.OC 版本更新 50%

A Systematic Review and Taxonomy of Reinforcement Learning-Model Predictive Control Integration for Linear Systems

一种线性系统中强化学习-模型预测控制整合的系统综述与分类

Mohsen Jalaeian Farimani, Roya Khalili Amirabadi, Davoud Nikkhouy, Malihe Abdolbaghi, Mahshad Rastegarmoghaddam, Shima Samadzadeh, Mahdi Ghane

机构 * Department of Electronics, Information and Bioengineering (DEIB), Politecnico di Milano(电子信息生物工程系(DEIB),米兰理工大学) Department of Applied Mathematics, Ferdowsi University of Mashhad(应用数学系,法尔德大学) Department of Mechanical Engineering, Politecnico di Milano(机械工程系,米兰理工大学) Department of Applied Mathematics, Faculty of Mathematical Sciences, University of Guilan(应用数学系,加林大学) Department of Robotics and Control Engineering, Shahrood University of Technology(机器人与控制工程系,沙霍德技术大学)

专题命中 通用世界模型 :分类 cs.AI、cs.RO;predictive model(abstract);predictive models(abstract)

AI总结 本文系统综述了线性系统中强化学习与模型预测控制整合的研究,通过多维分类揭示了整合策略、挑战及方法学趋势,为相关研究提供结构化参考。

详情

展开后加载摘要…

URL PDF HTML 收藏