arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 6450 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 4308 篇

2608.07981 2026-08-11 cs.CV 新提交 93%

Distilling Physical Priors into Streaming World Models

将物理先验知识蒸馏为流式世界模型

Liangliang Zhao, Junying Wang, Danni Yang, Yifan Chang, Bin Fu, Yu Qiao, Bowen Zhou, Yihao Liu

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出PhyS三阶段框架,构建120K物理交互数据集,经微调、蒸馏及在线强化学习结合TCR方法,提升流式世界模型的物理一致性,在PhysicsIQ等基准上取得显著性能提升。

Comments 9 pages, 7 figures. Project page: https://lyongo.github.io/PhyS/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00559 2026-08-10 cs.AI 版本更新 93%

Social World Models

社交世界模型

Xuhui Zhou, Jiarui Liu, Akhila Yerukola, Hyunwoo Kim, Maarten Sap

机构 * Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA, USA(卡内基梅隆大学语言技术研究所) NVIDIA, Santa Clara, CA, USA(英伟达)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出社交世界模型(SWMs)及结构化社交世界表示形式(S3AP),通过显式建模隐藏心理状态提升AI在社交推理任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05720 2026-08-07 cs.CV 新提交 93%

PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models

PhyLatent:为JEPA世界模型学习与动力学相关的表征

Xi Zeng, Haojie Ren, Ziying Song

机构 * School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与航天工程学院) School of Artificial Intelligence (School of Software), Yanshan University(燕山大学人工智能学院(软件学院)) The University of Sheffield(谢菲尔德大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 PhyLatent为JEPA世界模型设计训练目标,解决其三类失效模式,在OGBench-Cube等数据集上显著降低失效率、提升MPC成功率,证明全局非坍塌不足以学习可靠的JEPA状态空间。

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05706 2026-08-07 cs.CV 新提交 93%

LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models

LAWM-3D:从人类视频中学习具有3D感知能力的潜在动作以构建可泛化的机器人世界模型

Jiarui Yang, Jiale Zhange, Jiawei Li, Hang Guo, Wen Huang, Jinpeng Wang, Peidong Liu, Shu-Tao Xia

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本研究针对现有潜在动作模型缺乏3D感知能力的问题,提出LAWM-3D模型,通过多视图动作 token 化、几何对齐约束和RGB-D联合重建目标,实现了性能SOTA的机器人世界模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05523 2026-08-07 cs.CV 新提交 93%

HERA: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models

HERA:用于潜在世界模型中物理预测的历史证据路由适配器

Yuanruyi, Yue Cao, Haojia Gao, Guanqiu Guo, Ziyuezhang, Shangqin, Junbo Tan, Bokui Chen, Zhuo Zou, Xueqian Wang

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本研究提出 HERA 框架,通过 RRPM 适配器为冻结潜在预测器路由历史证据,在 IntPhys2 数据集上显著提升 V-JEPA 2-G 的物理预测准确率,验证了该策略的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21910 2026-08-07 cs.AI cs.DB 版本更新 93%

TRW: TRACE-RealWorld---An Auditable Consistency Contract for World Models as Materialized Views

TRW: TRACE-RealWorld——作为物化视图的世界模型的可审计一致性契约

Edward Y. Chang

机构 * Stanford University(斯坦福大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 TRACE-RealWorld解决在变化物理世界中维护物化视图的问题,提出承诺级有效性抽象等数据管理方法,通过端到端评估测量多方面指标,贡献是为部署世界表示提供一致性、恢复和问责契约。

Comments v2 propagates declared point and extended hazards through the composition theorem and separates point-event from interval-risk budgets. It aligns Flood-SAR adjudication with each hazard class, sharpens calibration-slack and drift claims, adds bootstrap Monte Carlo error and joint-coverage requirements, revises terminology, and expands related-work context

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13460 2026-08-07 cs.CV 版本更新 93%

VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models

VISA: VLM引导的实例语义审计用于3D占据世界模型

Ruiqi Xian, Yuehan Xian, Jing Liang, Xuewei Qi, Dinesh Manocha

机构 * University of Maryland College Park(马里兰大学帕克分校) Nanjing University of Posts and Telecommunications(南京邮电大学) Stanford University(斯坦福大学) Motional AD Inc.(Motional AD公司)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出VISA方法,利用离线VLM对每个物理对象实例进行结构化语义审计,并通过可靠性加权损失蒸馏到3D占据模型中,无需VLM推理即可提升封闭集占据mIoU。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00883 2026-08-05 cs.MM cs.CV cs.SD 版本更新 93%

Audio-Visual World Models: Learning Physically Grounded Multisensory Dynamics

视听世界模型:为具身智能体奠定多感官想象的基础

Jiahua Wang, Leqi Zheng, Jialong Wu, Yaoxin Mao, Shijie Cheng

机构 * Tsinghua University(清华大学) Beijing Institute of Technology(北京理工大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出视听世界模型(AVWM)统一框架,通过条件扩散Transformer(AV-CDiT)联合预测双耳音频与视觉动态,在30小时基准AVW-4k上实现高保真多模态预测,并验证其在具身导航中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02603 2026-08-04 cs.CV 新提交 93%

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

WorldExam:从表观外观到内在反应性对世界模型进行基准测试

Yuxue Yang, Shuyao Shang, Jiahe Wang, Zitong Zhou, Liang Tan, Junhan Zeng, Ruizhi Li, Junyan Li, Yu Liu, Xiao Yang, Yong Li, Jun Zhu, Hongsheng Li, Tieniu Tan, Lue Fan, Zhaoxiang Zhang

机构 * CASIA(中国科学院自动化研究所) SLAI(智能科学与技术实验室) CUHK(香港中文大学) AMAP(中国农业科学院) THU(清华大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 该研究推出WorldExam基准,评估20个模型在视觉质量等四层级的表现,发现不同驱动模型能力有明显分化,无模型兼具广泛任务覆盖与稳定性能,高视觉质量不代表内在反应性强。

Comments Project Website: https://WorldExam.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01926 2026-08-04 cs.AI 新提交 93%

ProWorld: Progress-Aware Hyperbolic World Models for Long-Horizon Visual Goal Reaching

ProWorld:用于长程视觉目标到达的感知进展双曲世界模型

Zihan Liu, Yuzhe Zhuang, Yuanzu Li, Wanshuang Gou, Jiahong Liu, Min Zhou, Menglin Yang

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文针对长程视觉目标到达任务中视觉世界模型的进展感知与轨迹区分难题,提出双曲视觉世界模型ProWorld,经实验相较LeWM实现9.67的平均绝对成功率提升。

Comments 24 pages, 14 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05238 2026-08-04 cs.AI 版本更新 93%

Branch-JEPA: Finite-Support Predictive Distributions for JEPA World Models

MoP-JEPA:用于随机JEPA世界模型的硬分配预测器混合架构

Zhi Song, Ximing Xing, Zhenchao Tang, hanbo Huang, Jiehui Huang, Weilong Yan, Tianxu Lv, Minghao Yang, Zhongzheng Niu, Bing He, Lusheng Wang, Jianhua Yao

机构 * City University of Hong Kong, China(香港城市大学) Tencent, China(腾讯)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 针对传统JEPA在随机环境下的状态坍缩问题,提出硬分配预测器混合的MoP-JEPA,可收敛到转移分布量化器,在OGBench测试中规划性能远超基线方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17102 2026-08-04 physics.pop-ph cs.AI cs.ET cs.HC quant-ph 版本更新 93%

Quantum Cinema: An Interactive Cinematic Exploration of Quantum Computing Hardware via Generative World Models

量子影院:通过生成世界模型对量子计算硬件进行交互式电影探索

Aoyu Zhang, Dongping Liu, Luyao Zhang

机构 * Amazon Web Services(亚马逊网络服务) Duke Kunshan University(杜克昆山大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出量子影院,一个基于生成世界模型的开源交互式应用,通过四幕叙事将不可见的量子硬件转化为可探索的电影体验,旨在弥合量子计算与公众之间的想象鸿沟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26754 2026-07-30 cs.CV 新提交 93%

StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation

StatePlay:用于机制一致性生成的状态感知游戏世界模型

Zijun Lin, Zeqing Wang, Cheston Tan, Bihan Wen, Yeying Jin

机构 * Tencent(腾讯)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出StatePlay,一种状态感知游戏世界模型,采用混合Transformer架构联合预测视觉内容与游戏状态,经实验验证可提升游戏生成的机制保真度。

Comments Project Page: https://jimntu.github.io/stateplay_page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26752 2026-07-30 cs.LG 新提交 93%

CalTwin: Towards Calibrated, Shift-Robust Medical World Models via Fisher-Information Regularisation

CalTwin:基于Fisher信息正则化的校准、分布偏移鲁棒医学世界模型研究

Behraj Khan, Shabir Ahmad, Syed Ahmad Chan Bukhari, Tahir Qasim Syed

机构 * Institute of Business Administration Karachi(卡拉奇商业管理学院) Gachon University(嘉泉大学) St. John’s University(圣约翰大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本研究提出CalTwin方法,结合Fisher信息偏移惩罚与置信度失配惩罚,应用于GRU医学世界模型,在PhysioNet脓毒症挑战赛上显著降低了分布外隐状态预测误差,同时改善了置信度校准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15284 2026-07-30 cs.CV 版本更新 93%

Walk through Paintings: Egocentric World Models from Internet Priors

穿越绘画:来自互联网先验的自我中心世界模型

Anurag Bagchi, Zhipeng Bao, Homanga Bharadhwaj, Yu-Xiong Wang, Pavel Tokmakov, Martial Hebert

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Toyota Research Institute(丰田研究机构)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 EgoWM通过利用互联网先验和轻量级条件层,将预训练视频扩散模型转换为自我中心世界模型,实现可控的未来预测,提升结构一致性评分并降低推理延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26056 2026-07-29 cs.RO 新提交 93%

INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models

INTACT:用于无搜索世界模型的同构意图到动作学习

Junhan Sun, Hao Zhao, Guofeng Zhang

机构 * State Key Laboratory of CAD&CG, Zhejiang University(浙江大学计算机辅助设计与图形学国家重点实验室) Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) InSpatio

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究针对前向潜在世界模型恢复动作需昂贵搜索的问题,提出INTACT方法,通过特定架构和技术实现意图到动作学习,在多任务上取得高成功率,减少采样并提升性能,还具有快速推理能力。

Comments 28 pages, 11 figures, including appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25236 2026-07-29 cs.CL cs.RO 新提交 93%

VisualPatchWorld: Code World Models as Latent Structured Representations for Planning

视觉补丁世界:作为规划潜在结构化表示的代码世界模型

Jiaxin Bai, Jiaxuan Xiong

机构 * Hong Kong Baptist University(香港浸会大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究旨在构建用于规划的代码世界模型。提出VisualPatchWorld方法,先选定性动力学形式,再拟合参数。实验表明其平均规划成功率69.0%,超基线23.5分,在多方面接近真实引擎成功率,为自动构建有用的代码世界模型提供实用途径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23147 2026-07-28 cs.CR cs.AI 新提交 93%

False Prophets: On the Security of World Models in Agentic Systems

虚假预言:论智能体系统中世界模型的安全性

Erik Imgrund, Anna Wimbauer, Klim Kireev, Konrad Rieck

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究智能体系统中世界模型的安全性,发现其存在特定漏洞,引入安全基准数据集,指出攻击者能诱导错误预测,成功率达95%,并为从业者提供减轻危害、强化系统的实用建议。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22430 2026-07-28 cs.LG 版本更新 93%

On the Identifiability of Controlled World Models

关于受控世界模型的可识别性

Xiangteng Zhang, Yang Guan, Bo Zhang, Hongyang Li, Ya-Qin Zhang, Shengbo Eben Li

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究受控世界模型的可识别性问题,建立在状态依赖高斯行为策略下的联合可识别性理论,识别出两个条件,证明满足条件时JEPA目标的全局极小值可识别潜在状态和受控转移,推导定量界限并通过实验验证理论。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18715 2026-07-22 cs.AI 新提交 93%

DWM: Separating World Effects from Actions in Latent World Models

DWM:在潜在世界模型中分离世界效应与动作

Yi-Ge Zhang, Tianqi Du, Qi Zhang, Yisen Wang

机构 * Peking University(北京大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究在潜在世界模型中分离世界效应与动作的问题,提出DWM框架,通过辅助世界头和正交性约束实现预测转换的显式加法分解,在构建的W变体基准测试中取得更好效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01177 2026-07-22 cs.RO 版本更新 93%

Scaling Cross-Embodiment World Models for Dexterous Manipulation

为灵巧操作扩展跨实体世界模型

Zihao He, Bo Ai, Tongzhou Mu, Yulin Liu, Weikang Wan, Jiawei Fu, Yilun Du, Henrik I. Christensen, Hao Su

机构 * UC San Diego(UC圣迭戈大学) Harvard University(哈佛大学) Shanghai Jiao Tong University(上海交通大学) Hillbot

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究跨实体学习中不同形态机器人数据共享和控制转移问题,提出将人类和机器人手表示为3D粒子集,在随机交互数据上训练基于图的世界模型并与模型预测控制集成,实验表明该方法能提高泛化能力、有效结合数据并实现不同机器人手的控制。

Comments Accepted to IROS 2026, Project Page: https://alan-heoooh.github.io/dexwm.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05558 2026-07-21 cs.LG 版本更新 93%

Autoregressive Diffusion World Models for Off-Policy Evaluation of LLM Agents

自回归扩散世界模型用于LLM智能体的离线评估

Kaixuan Liu, Guojun Xiong, Weinan Zhang, Shengpu Tang

机构 * Department of Computer Science, Emory University(埃默里大学计算机科学系) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出ADWM框架,通过自回归扩散世界模型从预收集轨迹中模拟环境响应,实现无需在线交互的LLM智能体策略离线评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13560 2026-07-16 q-bio.NC cs.AI 新提交 93%

Grounded world models in biological organisms and future embodied AI

生物有机体和未来具身人工智能中的基础世界模型

Giovanni Pezzulo, Davide Nuzzi, Marco D'Alessandro, Riccardo Proietti, Roberto Bottini, Paul Cisek

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究探讨生物有机体中基础世界模型,通过五个神经回路例子揭示当前具身人工智能缺失的特征,如内在动力学作用等,还讨论了生物系统原则对未来具身人工智能的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12474 2026-07-16 cs.AI 版本更新 93%

From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery

从观察到洞察:机制世界模型与自主发现探索

Ingmar Posner, Anson Lei, Bernhard Schölkopf

机构 * MPI for Intelligent Systems & ELLIS Institute(马克斯·普朗克智能系统研究所及埃利斯研究所)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文探讨科学发现问题,提出机制世界模型这一新设计范式,将可重复使用机制置于核心,推导其计算能力、设计原则等,指出虽有不同研究方向捕捉该范式要素但缺统一框架,为推动AI走向自主科学发现提供基础和蓝图。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08312 2026-07-10 cs.LG 新提交 93%

Write-Protected Discrete Bottlenecks for Language-Grounded World Models: A Structural Limitation and Sufficient Fix

用于语言基础世界模型的写保护离散瓶颈:一种结构限制及充分修复

Jiayi Fang

机构 * Shanghai University of Finance and Economics(上海金融学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 研究语言与世界模型离散符号系统的交互问题,提出防止基于Gumbel-softmax的离散符号瓶颈失败的三个约束,经实验验证该方法能实现高基础准确率,且修复参数少、无需微调。

Comments 20 pages, 7 figures. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06445 2026-07-09 cs.CV 版本更新 93%

What if? Emulative Simulation with World Models for Situated Reasoning

如果呢?基于世界模型的仿真模拟用于情境推理

Ruiping Liu, Yufan Chen, Yuheng Zhang, Junwei Zheng, Kunyu Peng, Chengzhi Wu, Chenguang Huang, Di Wen, Jiaming Zhang, Kailun Yang, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Hunan University(湖南大学) ETH Zürich(苏黎世联邦理工学院) INSAIT, Sofia University ``St. Kliment Ohridski''(INSAIT,索菲亚大学『圣克莱门特·奥赫里迪斯』) RAI Institute(RAI研究所)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出WanderDream数据集,利用世界模型进行心理探索的仿真模拟,使代理无需主动探索即可回答空间假设问题,实验证明心理探索对情境推理至关重要。

Comments Accepted at ECCV 2026. The data and code are available at: https://github.com/RuipingL/WanderDream

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31422 2026-07-07 cs.AI 新提交 93%

Ask the World Before Acting: Environment Probing for Calibrated Agent World Models

行动前先询问世界:预算受限的环境探测用于世界模型校准

Xinyuan Song, Zekun Cai

机构 * Emory University(埃默里大学) The University of Tokyo(东京大学) LocationMind

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出预算受限的环境探测算子,通过结构化信念表在行动前校准世界模型,减少长期任务中的终端误差。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27780 2026-07-07 cs.AI 新提交 93%

Understanding Rollout Error in Graph World Models

理解图世界模型中的展开误差

Xinyuan Song, Zekun Cai

机构 * Emory University(埃默里大学) The University of Tokyo(东京大学) LocationMind

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文研究图世界模型中的长时域展开误差,提出统一框架分析拓扑与模型引起的误差放大,并设计误差感知图世界模型,通过谱正则化、展开一致性和关键节点加权防止长时域发散。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01537 2026-07-03 cs.LG 新提交 93%

Certified World Models as Sensing Clocks: Drift-Aware Deadlines for Active Perception

认证世界模型作为感知时钟:主动感知的漂移感知截止时间

Hongbo Wang

机构 * Hongbo Wang(王Hongbo)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出一种基于认证世界模型的感知时钟,通过漂移感知的截止时间规则决定主动感知时机,在合成基准和3D VN-JEPA模型上验证了有效性。

Comments 15 pages, 3 figures, 6 tables. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01896 2026-07-02 cs.CV 版本更新 93%

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models

分而治之:多模态世界模型的解耦表示对齐

Junyuan Xiao, Dingkang Liang, Xin Zhou, Yixuan Ye, Tongtong Su, Guangmo Yi, Bin Xia, Qiang Lyu, Shurui Shi, Jun Huang, Jianlou Si, Wenming Yang

机构 * Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学) CSU(中南大学) ZJU(浙江大学) CUHK(香港中文大学) UCAS(中国科学院大学) Alibaba Group(阿里巴巴集团)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出M²-REPA方法,通过解耦扩散模型中间表示中的模态特定特征并与对应的专家基础模型对齐,实现多模态视频生成中多种基础模型先验的充分利用,显著提升视觉质量和长期一致性。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏