arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

2026-06-23 至 2026-06-23 共收录 39 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 39 篇

2606.23105 2026-06-23 cs.CV 新提交 96%

Compression and Retrieval: Implicit Memory Retrieval for Video World Models

压缩与检索:视频世界模型的隐式记忆检索

Zhan Peng, Jie Ma, Huiqiang Sun, Chong Gao, Zhijie Xue, Zhiyu Pan, Zhiguo Cao, Jun Liang, Jing Li

机构 * Huazhong University of Science and Technology(华中科技大学) HUJING Digital Media & Entertainment Group(虎鲸数字媒体与娱乐集团) Sun Yat-sen University(中山大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);video world model(title,abstract);world model(title,abstract)

AI总结 提出注意力驱动的隐式记忆检索机制CaR,通过位置编码注入视角信息实现灵活检索,并引入轻量级上下文压缩网络,在构建的SceneFly数据集上取得SOTA结果并展现强泛化性。

Comments Project page: 3DV-Team/CaR" target="_blank" rel="noopener">https://github.com/Orange-3DV-Team/CaR

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24354 2026-06-23 cs.CV 版本更新 94%

SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation

SparseWorld: 通过具有稀疏场景表示的世界模型增强端到端自动驾驶

Ruoyu Wang, Jingke Wang, Yukai Ma, Yuehao Huang, Shuangming Lei, Guanglin Xu, Aixue Ye, Yong Liu

机构 * Institute of Cyber-Systems and Control, Zhejiang University(浙江大学控制系统研究所) Labs, Huawei(华为2012实验室) State Key Laboratory of Industrial Control Technology(国家工业控制技术重点实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出SparseWorld,一种基于稀疏场景表示的轻量级世界模型,通过自回归预测未来地图元素和周围智能体,并利用预测结果优化下游运动预测和轨迹规划,在nuScenes数据集上实现0.05%的碰撞率,达到开放循环规划指标的最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18313 2026-06-23 cs.CV 版本更新 94%

OmniNWM: Omniscient Driving Navigation World Models

OmniNWM:全知驾驶导航世界模型

Bohan Li, Zhuang Ma, Dalong Du, Baorui Peng, Zhujin Liang, Zhenqiang Liu, Xianda Guo, Zheng Zhu, Chao Ma, Yueming Jin, Xin Jin, Hao Zhao, Wenjun Zeng

机构 * Shanghai Jiao Tong University(上海交通大学) Eastern Institute of Technology(东部技术研究所) PhiGent National University of Singapore(新加坡国立大学) Tsinghua University(清华大学) Wuhan University(武汉大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出OmniNWM全知全景导航世界模型,在统一概率框架下处理状态、动作和奖励三个维度,实现多模态全景视频生成、零样本轨迹控制及内在奖励机制,在生成保真度和控制精度上达到SOTA。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21775 2026-06-23 cs.LG cs.AI 新提交 94%

Beyond the Next Step: Variable-Length Latent World Models for Long-Horizon Planning

超越下一步:用于长时程规划的变长潜世界模型

Tianqi Du, Qi Zhang, Yifei Wang, Yisen Wang

机构 * State Key Lab of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院通用人工智能国家重点实验室) Amazon AGI SF Lab(亚马逊AGI旧金山实验室) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出变长潜世界模型(VLWM),通过学习变长动作序列的条件潜状态预测,解决递归一步预测在长时程规划中的累积误差问题,结合课程训练策略,在长时程控制任务上平均提升13%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23296 2026-06-23 cs.RO 新提交 93%

IOI: Decoupling Kinematics and Physics for Interactive World Models

IOI: 解耦运动学与物理学的交互式世界模型

Chengyu Bai, Peidong Jia, Tiecheng Guo, Yukai Wang, Rui Ma, Fangyuan Zhao, Chunkai Fan, Xiaobao Wei, Jintao Chen, Hao Wang, Ying Li, Xiaozhu Ju, Jian Tang, Shanghang Zhang

机构 * Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) Peking University(北京大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出IOI混合交互式世界模型,通过显式运动学先验与学习物理动力学解耦,实现精确控制对齐和物理合理视觉反馈,在RoboTwin基准上达到最优仿真性能与零样本泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23463 2026-06-23 econ.GN q-fin.EC 新提交 93%

Equilibrium World Models

均衡世界模型

Simon Scheidegger, Andreas Schaab

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出均衡世界模型(EWMs),一种深度学习方法,通过扩大训练分布并引入代理函数,全局求解包含罕见灾难、约束和反事实状态的动态随机模型,克服标准神经网络求解器的自我确认偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21672 2026-06-23 cs.RO cs.AI cs.LG 新提交 93%

Imitation from Heterogeneous Demonstrations using Grounded Latent-Action World Models

使用基于接地潜在动作世界模型的异构演示模仿学习

Tianyou Wang, Anson Lei, Joe Watson, Ingmar Posner

机构 * University of Oxford(牛津大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 提出GLAM方法,通过共享潜在动作空间对齐异构数据源,学习接地潜在动作世界模型,在数据稀缺时提升模仿学习性能,平均任务成功率提高48%。

Comments 17 pages, 8 figures. Project page: https://viccccciv.github.io/glam/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20764 2026-06-23 cs.CV cs.AI cs.GR cs.LG 新提交 93%

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception

一图足矣:基于文本世界模型的智能体单样本图像生成用于长尾空间感知

Keqin Zeng, Shuting Su, Shihao Lin, Ziyue Li, Rui Zhao

机构 * Tsinghua University(清华大学) SenseTime Research(商汤科技研究院) Sun Yat-Sen University(中山大学) Technical University of Munich(慕尼黑工业大学) Heilbronn Data Science Center(海尔布隆数据科学中心) Munich Data Institute(慕尼黑数据研究所)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 提出WMGen-v1框架,利用单张参考图像通过LVLM构建结构化场景表示,LLM进行物理合理的场景扩展,再由扩散模型生成多样化的长尾训练数据,缓解空间感知中的数据稀缺问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21173 2026-06-23 cs.LG cs.AI 新提交 93%

Inverting the Bellman Equation: From $Q$-Values to World Models

逆推贝尔曼方程:从 $Q$ 值到世界模型

Alistair Letcher, Mattie Fellows, Alexander D. Goldie, Jonathan Richens, Jakob N. Foerster, Oliver Richardson

机构 * FLAIR, University of Oxford(FLAIR,牛津大学) Google DeepMind(谷歌DeepMind) Mila, University of Montreal(Mila,蒙特利尔大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 本文证明基于值的智能体在丰富奖励函数上训练时隐式编码世界模型,提出 $P$-learning 从 $Q$ 值提取模型,并给出编码真实转移核的充分条件,实验验证了隐式模型的准确性和泛化能力。

Comments 48 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23280 2026-06-23 cs.RO 新提交 92%

Causal Reward World Models: Zero-shot Reward Design for Automated Skill Generation

因果奖励世界模型:面向自动化技能生成的零样本奖励设计

Yang Yang, Yuchuang Tong, Zhengtao Zhang, Xu Ding, Ning Yang, Yifan Zhang, Haipeng Li, Kehu Yang, Miao Xin

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Intelligent Manufacturing Institute, HFUT(合肥工业大学智能制造研究院) School of Electrical Engineering and Automation, Anhui University(安徽大学电气工程与自动化学院) School of Artificial Intelligence, China University of Mining and Technology (Beijing)(中国矿业大学(北京)人工智能学院) National Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institution of Automation, Chinese Academy of Sciences(中国科学院自动化研究所复杂系统认知与决策智能全国重点实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 提出因果奖励世界模型(CRWM),通过离线预训练学习候选奖励组件与任务目标物理变量间的因果拓扑关系,结合显式机制解耦与置信感知软融合的联合优化模块构建因果骨架,使LLM在零样本下生成可执行奖励函数,无需反馈迭代,显著降低新技能设计延迟并保持或超越现有性能。

Comments 22 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22729 2026-06-23 cs.RO 新提交 92%

Temporal Logic Guidance for Action-Only Diffusion Policies with World Models

基于时序逻辑的动作扩散策略与世界模型引导

Moritz Zoellner, Anastasios Manganaris, Rohan Paleja

机构 * German Academic Exchange Service (DAAD)(德国学术交流中心(DAAD))

专题命中 通用世界模型 :world model(title,abstract);world models(title);world model(title,abstract);world models(title)

AI总结 提出一种利用世界模型实现时序逻辑鲁棒性可微评估的引导方法,在不重新训练的情况下改善动作扩散策略的约束满足,在Robomimic任务中将违规率从80%降至4%。

Comments Accepted at the ICRA 2026 Workshop on Bridging the Gap between Robot Learning and Human-Robot Interaction. 3 pages, 2 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22449 2026-06-23 cs.AI cs.RO 新提交 92%

Self-Evolving Cognitive Framework via Causal World Modeling for Embodied Scientific Intelligence

基于因果世界建模的自进化认知框架用于具身科学智能

Yi Yu, Tetsunari Inamura

机构 * Graduate School of Advanced Science and Engineering, Hiroshima University(广岛大学先进科学与工程研究生院) Advanced Intelligence and Robotics Research Center, Brain Science Institute, Tamagawa University(玉川大学脑科学研究所先进智能与机器人研究中心)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);embodied world model(abstract)

AI总结 提出一种自进化认知框架,通过因果世界建模、干预驱动推理和持续认知精炼,使具身智能体在交互中不断构建和修正内部因果模型,实现从预测智能到认知智能的转变。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18697 2026-06-23 cs.LG cs.CR cs.RO 新提交 92%

Stealthy World Model Manipulation via Data Poisoning

通过数据投毒进行隐蔽的世界模型操纵

Yibin Hu, Xiaolin Sun, Zizhan Zheng

机构 * Department of Computer Science(计算机科学系)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world-model(abstract)

AI总结 提出SWAAP框架,通过两阶段数据投毒(双层级优化寻找有害目标模型+梯度匹配隐蔽实现)操纵学习到的世界模型,导致规划性能显著下降,且能规避多种防御检测。

Comments 41 pages, 8 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22966 2026-06-23 cs.LG cs.AI cs.CR 新提交 91%

Attacking the Trusted Imagination: Oracle-Level Integrity Attacks on Imagine-then-Act World Models

攻击可信想象:对“先想象后行动”世界模型的预言机级完整性攻击

Linghan Chen, Kaiyan Ji, Minyu Guo

机构 * Adelaide University(阿德莱德大学)

专题命中 通用世界模型 :world model(title);world models(title);world model(title);world models(title)

AI总结 针对“先想象后行动”的视觉-语言-动作策略,发现世界模型中的想象轨迹是暴露的攻击面,通过有界观测扰动破坏想象,并设计无参数去噪检测器,实验显示非定向破坏效果显著,但定向控制受限。

Comments 13 pages, 5 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23079 2026-06-23 cs.RO cs.AI 新提交 90%

AdaReP:Adaptive Re-Planning under Model Mismatch for Neural World-Model Predictive Control

AdaReP:模型失配下的自适应重规划用于神经世界模型预测控制

Yutian Cheng, Xiaojian Ma, Xianhao Wang, Min Yang, Rongpeng Su, Hangxin Liu, Xi Chen, Shuai Li, Qing Li

机构 * Shanghai Jiao Tong University(上海交通大学) Beijing Institute for General Artificial Intelligence (BIGAI)(北京通用人工智能研究院) University of Science and Technology of China(中国科学技术大学)

专题命中 通用世界模型 :world-model(title);world-model(title);world model(abstract);world models(abstract)

AI总结 针对神经世界模型预测控制中重规划计算开销大的问题,提出AdaReP方法,通过在线自适应调整重规划容忍度,在保持任务性能的同时大幅减少规划器计算量。

Comments Accepted at ICANN 2026. This arXiv version contains supplementary materials and appendices that are omitted from the conference version due to space limitations

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23085 2026-06-23 cs.RO 新提交 90%

Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents

Foresight: 基于动作条件世界模型潜在变量的长时程机器人操作故障检测

Haoran Zhang, Yifu Lu, Boyang Wang, Xuhui Kang, Yen-Ling Kuo, Zezhou Cheng, Mengdi Wang, Odest Chadwicke Jenkins

机构 * University of Michigan(密歇根大学) Princeton University(普林斯顿大学) University of Virginia(弗吉尼亚大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world-model(abstract);world-model(abstract)

AI总结 提出Foresight框架,利用动作条件世界模型的潜在表征监测操作轨迹,仅用最终任务级标签训练,结合函数共形预测自适应校准阈值,在仿真和真实机器人长时程任务中实现高效故障检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16689 2026-06-23 eess.SP cs.NI 版本更新 90%

Against the Monolithic Wireless World Model: Why NextG Needs Composable and Agentic Intelligence

反对单一无线世界模型:为什么NextG需要可组合和代理智能

Aladin Djuhera, Farhan Ahmed, Vlad C. Andrei, Swanand Ravindra Kadhe, Alecio Binotto, Haris Gacanin, Holger Boche

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);world models(abstract);world models(abstract)

AI总结 本文指出无线领域缺乏等同于LLM的数据基础,主张采用可组合和代理智能架构以实现可部署的AI原生网络。

Journal ref ICML 2026 Workshop on AI and ML for NextG

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21961 2026-06-23 cs.LG 新提交 89%

VegSim: A Geospatial World Model for Scenario-Conditioned Vegetation Simulation

VegSim:一种用于情景条件植被模拟的地理空间世界模型

Irene Iele, Elena Mulero Ayllón, Paolo Soda, Matteo Tortora

机构 * Università Campus Bio-Medico di Roma(罗马生物医学大学) Umeå University(于默奥大学) University of Genoa(热那亚大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);latent dynamics(abstract);分类 cs.LG

AI总结 提出VegSim,一种地理空间世界模型,通过可控制的未来气象输入,实现情景条件植被模拟,在分布内和分布外数据上均优于基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11061 2026-06-23 cs.CV 版本更新 89%

VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation

VDAWorld: 通过VLM导向的抽象与模拟进行世界建模

Felix O'Mahony, Roberto Cipolla, Ayush Tewari

机构 * University of Cambridge(剑桥大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);latent dynamics(abstract);分类 cs.CV

AI总结 提出VDAWorld框架,利用视觉语言模型自主构建场景表示并选择物理模拟器,通过抽象与自适应模拟实现高质量世界建模,在交互控制、反事实生成和物理逻辑推理任务上取得最优结果。

Comments Website: https://felixomahony.github.io/vdaworld/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12373 2026-06-23 cs.LG cs.AI cs.SI 版本更新 88%

Policy4OOD: A Knowledge-Guided World Model for Policy Intervention Simulation against the Opioid Overdose Crisis

Policy4OOD:一种知识引导的世界模型,用于针对阿片类药物过量危机的政策干预模拟

Yijun Ma, Zehong Wang, Weixiang Sun, Zheyuan Zhang, Kaiwen Shi, Nitesh Chawla, Yanfang Ye

机构 * University of Notre Dame, South Bend, USA(南达科他大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出Policy4OOD,一种知识引导的时空世界模型,通过联合编码政策知识图谱、空间依赖和时间序列,实现政策干预的预测、反事实推理和优化,并在2019-2024年美国州级数据上验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21886 2026-06-23 cs.RO cs.AI 版本更新 88%

From Discrete Plans to Real-World Execution: A World-Model-Driven Framework for Execution-Aware Multi-Agent Path Finding

从离散规划到真实世界执行:一种面向执行感知的多智能体路径规划的世界模型驱动框架

Jingtian Yan, Shuai Zhou, He Jiang, Stephen F. Smith, Jiaoyang Li

机构 * IEEE Publication Technology Group(IEEE出版技术组)

专题命中 通用世界模型 :world-model(title);world-model(title);world model(abstract);world model(abstract)

AI总结 提出ExecTimeNet世界模型预测离散MAPF方案在物理机器人上的执行状态,并基于此构建REMAP框架和ESADG优化方法,在仿真和实物实验中分别减少高达21%和15.3%的执行延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22261 2026-06-23 cs.LG cs.GT 新提交 88%

Learning a Normal World Model for Few-Shot Boundary-Calibrated Abnormality Detection

学习正常世界模型用于少样本边界校准的异常检测

Weizhi Nie, Weichao Liu, Weijie Wang, Yuting Su

机构 * Tianjin University(天津大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.LG

AI总结 提出超图熵正常世界模型,通过正常事件学习正常世界并用少量异常样本校准边界,在NASA C-MAPSS数据集上实现强零样本和少样本异常检测性能。

Comments 23 pages,8 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21501 2026-06-23 cs.RO 新提交 88%

UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling

UniviewVLA:统一的多视图视觉-语言-动作模型与世界建模

Tao Xu, Runhao Zhang, Zhijian Huang, Jiayi Guan, Jiaxin Wang, Yifan Ding, Yong-Lu Li, Long Chen, Guang Chen, Jinghui Lu

机构 * Tongji University(同济大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Jiao Tong University(上海交通大学) Xiaomi EV(小米汽车)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.RO

AI总结 提出UniviewVLA,利用世界模型生成多视图未来帧,从标准双摄像头观测中预测动作,解决遮挡问题,无需额外硬件或显式重建,并通过运动信息令牌压缩和动作熵视图选择提升效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21315 2026-06-23 cs.AI 新提交 88%

Social World Model for Lifelong Social Intelligence

社会世界模型:面向终身社会智能

Yu Luo

机构 * Central South University(中南大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI

AI总结 提出社会世界模型,将社交互动分解为五个维度构建闭环学习框架,并配套数据合成机制与终身学习基准,使小模型可持续获取社会协调能力。

Comments 13 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20867 2026-06-23 cs.CV cs.AI 新提交 87%

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation

FOCA: 面向未来的条件化用于数据高效的视觉-语言-动作适应

Duc Minh Nguyen, Nghiem Tuong Diep, Binh Gia Nguyen, Trong-Bao Ho, Doanh Le, Tan Q. Nguyen, Thien-Loc Ha, Nhiem Tran, Bao Thach, Nhat X. Tran, Tuan A. Tran, Artur Habuda, Philip Lund Møller, Tran Nguyen Le, Daniel Sonntag, Matthias Niepert, Khoa D. Doan, Vu Duong, Hung Ngo, Minh N. Vu, Duy M. H. Nguyen, An Thai Le, Ngo Anh Vien

机构 * Center for AI Research, VinUniversity, Vietnam University of Utah, USA German Research Center for Artificial Intelligence (DFKI) Technical University of Denmark, Denmark University of Oldenburg, Germany University of Stuttgart, Germany Max Planck Research School for Intelligent Systems (IMPRS-IS), Germany

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 提出FOCA框架,结合未来交互嵌入预测与目标观测隐式对齐,实现数据高效的VLA少样本适应,在LIBERO、RoboCasa和真实机器人上取得新最优结果。

Comments Accepted at ICML 2026. Project page: https://focavla.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20781 2026-06-23 cs.RO cs.CV 新提交 87%

World Action Models: A Survey

世界行动模型:综述

Qiuhong Shen, Shihua Zhang, Yue Liao, Qi Li, Zhenxiong Tan, Shizun Wang, Shuicheng Yan, Xinchao Wang

机构 * National University of Singapore(新加坡国立大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 本文综述世界行动模型(WAMs),通过两种互补视角组织现有工作,揭示其设计权衡与未来趋势。

Comments 57 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22363 2026-06-23 cs.AI cs.LG cs.RO 新提交 85%

Reference-Free Assessment of Physical Consistency in World Model-based Video Generation

基于世界模型的视频生成中物理一致性的无参考评估

Yun Oh, Sukmin Yun

机构 * Hanyang University ERICA(汉阳大学ERICA校区)

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.AI、cs.LG、cs.RO

AI总结 提出结合相对与绝对方法的无参考度量,用于评估生成视频的物理一致性,无需人工投票或真实参考,通过DROID-SLAM和SEA-RAFT量化不一致性,过滤后视频任务成功率提升超8%,并实现时空定位。

Comments Accepted to the 2nd 3D-LLM/VLA Workshop, CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24121 2026-06-23 cs.RO 版本更新 85%

Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation

在线世界建模实现真实世界逆强化学习从观察中学习

Tyler Han, Bat Nemekhbold, Siyang Shen, Rohan Baijal, Richard Ebock, Harine Ravichandiran, Sanghun Jung, Kevin Huang, Byron Boots

机构 * University of Washington(华盛顿大学)

专题命中 通用世界模型 :world model(title);world model(title);分类 cs.RO;simulation model(abstract)

AI总结 提出MPAIL2方法,通过在线世界建模实现从观察中逆强化学习,首次在真实世界从零学习视觉操作任务,40分钟内达到82%成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23455 2026-06-23 cs.CV 新提交 81%

MeGAS: Thermomechanical Dynamic Gaussian Splatting for Thermophysical Scene Editing

MeGAS: 用于热物理场景编辑的热力学动态高斯泼溅

Zesong Yang, Yuanhang Lei, Liyuan Cui, Yihang Chen, Jiaer Huang, Boming Zhao, Peter Yichen Chen, Hujun Bao, Zhaopeng Cui

机构 * State Key Laboratory of CAD\&CG, Zhejiang University(浙江大学计算机辅助设计与图形学国家重点实验室) University of British Columbia(不列颠哥伦比亚大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出MeGAS框架,将热力学相变动力学融入3D高斯泼溅,通过温度属性、热对流-扩散求解器和MPM动力学实现热物理现象的物理真实合成,并采用拓扑自适应渲染策略处理极端变形。

Comments Accepted by ECCV 2026. Project page: http://zju3dv.github.io/MeGAS

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22918 2026-06-23 cs.CV cs.GT 新提交 81%

Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation

各有所尺:发现用于物理视频评估的每VLM分类体系

Yu Cao, Ziquan Liu, Zhensong Zhang, Jiankang Deng, Shaogang Gong, Jifei Song

机构 * Queen Mary University of London(伦敦玛丽女王大学) Huawei Darwin Research Center(华为达尔文研究中心) Imperial College London(伦敦帝国理工学院)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出JudgeFit方法,通过迭代优化为每个视觉语言模型发现其专属的物理视频评估分类体系,在16个VLM上平均相对提升约32%,并揭示模型特定盲点。

详情

展开后加载摘要…

URL PDF HTML 收藏