arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 592 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 421 篇

2607.15550 2026-07-20 cs.AI 新提交 88%

SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction

SeerGuard:一种通过世界模型预测实现的移动 GUI 代理安全框架

Xue Yu, Bo Yuan, Pengshuai Yang, Kailin Zhao, Hong Hu, Junlan Feng

机构 * JIUTIAN Research(九天研究)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI

AI总结 研究针对移动 GUI 代理安全风险问题,提出 SeerGuard 框架,通过执行前指令级筛选和行动级风险评估减轻风险。构建安全增强世界模型,经实验验证其在不同代理上有效泛化,提升了安全效用分数并降低风险成本分数。

Comments 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13651 2026-07-16 cs.CV 新提交 88%

From Surface Forecasting to Observability Forecasting: A Latent World Model for Cloud-Aware EO Monitoring

从地表预测到可观测性预测:用于云感知地球观测监测的潜在世界模型

Mohanad Albughdadi

机构 * European Centre for Medium-Range Weather Forecasts(欧洲中期天气预报中心)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.CV

AI总结 研究地球观测处理链中地表可观测性预测问题,采用LeWorldModel模型并将其应用于云感知地球观测序列,经训练和评估,该模型在可观测性基准上优于持久性方法,在多方面表现良好且能产生异常信号。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05577 2026-07-08 cs.AI cs.CL cs.IR 新提交 88%

Narrative World Model: Narratology-Grounded Writer Memory for Long-Form Fiction

叙事世界模型:基于叙事学的长篇小说作者记忆

Mohammad Saifullah, Thomas Kornmaier, Taaha Kazi, Vasu Sharma, Aditya Sanjiv Kanade, Aanand Kumar Yadav

机构 * PocketFM(口袋FM)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI

AI总结 研究针对长篇小说作者记忆问题,提出叙事世界模型(NWM),结合叙事学时间状态图与查询条件混合检索,在多跳叙事学问答上显著优于基线,优势源于其独特表征,非提取方式。

Comments 23 pages, 4 figures; 9-page main text plus appendix. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02045 2026-07-03 cs.CV 新提交 88%

PWM-ArtGen: Part World Model for Articulated Object Generation

PWM-ArtGen: 面向铰接物体生成的部分世界模型

Wentao Zheng, Ancong Wu

机构 * School of Computer Science and Engineering, Sun Yat-sen University, China(计算机科学与工程学院,中山大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.CV

AI总结 提出部分世界模型PWM-ArtGen,通过联合学习视觉动态和运动学参数,利用无标注数据协同训练,实现从单张图像生成铰接3D物体,在静止状态和零样本泛化上显著优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01595 2026-07-03 cs.AI cs.CL 新提交 88%

Safe and Adaptive Cloud Healing: Verifying LLM-Generated Recovery Plans with a Neural-Symbolic World Model

安全且自适应的云修复:使用神经符号世界模型验证LLM生成的恢复计划

Junyan Tan, Haoran Lin, Siyuan Guo, Yichen Fang, Xinyue Luo, Tianyu Shen, Zeyu Qiao

机构 * Zhejiang University(浙江大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI

AI总结 提出PASE框架,将LLM作为规划引擎生成恢复计划,通过神经符号世界模型验证可行性,并利用DRL优化提示,实现动态自适应修复,显著降低恢复时间并提高未知故障检测精度。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10044 2026-06-30 cs.AI 新提交 88%

Business World Model

商业世界模型

Cecil Pang, Hiroki Sayama

机构 * AI Engineering, USA TODAY Co., Inc.(AI工程,USA TODAY公司) School of Systems Science and Industrial Engineering, Binghamton University, State University of New York(系统科学与工业工程学院,宾夕法尼亚州立大学宾夕法尼亚州立大学) Binghamton Center of Complex Systems, Binghamton University, State University of New York(宾夕法尼亚州立大学复杂系统中心) Waseda Innovation Lab, Waseda University, Tokyo, Japan(早稻田大学创新实验室)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI

AI总结 提出商业世界模型(BWM)架构,将世界模型思想应用于商业环境,通过编码状态、动态、约束和目标,支持自主决策与规划。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27575 2026-06-29 cs.CV 新提交 88%

Perceptual 3D Simulation With Physical World Modeling

基于物理世界建模的感知3D仿真

Wanhee Lee, Klemen Kotar, Rahul Mysore Venkatesh, Jared Watrous, Daniel L. K. Yamins

机构 * Stanford University(斯坦福大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.CV

AI总结 提出P3Sim系统,结合学习型物理世界模型、几何条件模块和持久场景记忆,在部分观测和不完整3D变换信号下模拟未来场景状态,实现多任务泛化。

Comments Published as a conference paper at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26713 2026-06-26 cs.AI 新提交 88%

LithoDreamer: A Physics-Informed World Model for Multi-Stage Computational Lithography

LithoDreamer: 一种用于多阶段计算光刻的物理信息世界模型

Yuqi Jiang, Yumeng Liu, Zimu Li, Jinyuan Deng, Qian Jin, Yucheng Cui, Yu Li, Xunzhao Yin, Qi Sun, Cheng Zhuo

机构 * Zhejiang University(浙江大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI

AI总结 提出首个物理信息世界模型LithoDreamer,将光刻流程建模为决策驱动的多步演化系统,通过对比变分优化实现可解释干预,在正向演化和逆向规划中达到最优性能。

Comments Correspondence to: Qi Sun <qisunchn \at zju \dot edu \dot cn>

Journal ref Proceedings of the 43rd International Conference on Machine Learning (ICML), Jul. 6-11, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22261 2026-06-23 cs.LG cs.GT 新提交 88%

Learning a Normal World Model for Few-Shot Boundary-Calibrated Abnormality Detection

学习正常世界模型用于少样本边界校准的异常检测

Weizhi Nie, Weichao Liu, Weijie Wang, Yuting Su

机构 * Tianjin University(天津大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.LG

AI总结 提出超图熵正常世界模型,通过正常事件学习正常世界并用少量异常样本校准边界,在NASA C-MAPSS数据集上实现强零样本和少样本异常检测性能。

Comments 23 pages,8 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21501 2026-06-23 cs.RO 新提交 88%

UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling

UniviewVLA:统一的多视图视觉-语言-动作模型与世界建模

Tao Xu, Runhao Zhang, Zhijian Huang, Jiayi Guan, Jiaxin Wang, Yifan Ding, Yong-Lu Li, Long Chen, Guang Chen, Jinghui Lu

机构 * Tongji University(同济大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Jiao Tong University(上海交通大学) Xiaomi EV(小米汽车)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.RO

AI总结 提出UniviewVLA,利用世界模型生成多视图未来帧,从标准双摄像头观测中预测动作,解决遮挡问题,无需额外硬件或显式重建,并通过运动信息令牌压缩和动作熵视图选择提升效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21315 2026-06-23 cs.AI 新提交 88%

Social World Model for Lifelong Social Intelligence

社会世界模型:面向终身社会智能

Yu Luo

机构 * Central South University(中南大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI

AI总结 提出社会世界模型,将社交互动分解为五个维度构建闭环学习框架,并配套数据合成机制与终身学习基准,使小模型可持续获取社会协调能力。

Comments 13 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18589 2026-06-18 cs.RO 新提交 88%

DREAM-Chunk: Reactive Action Chunking with Latent World Model

DREAM-Chunk:基于潜在世界模型的反应式动作分块

Wenxi Chen, Kaidi Zhang, Chi Lin, Zhiyuan Zhang, Yu She, Yuejiang Liu, Raymond A. Yeh, Shaoshuai Mou, Yan Gu

机构 * Purdue University(普渡大学) Stanford University(斯坦福大学)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.RO

AI总结 提出DREAM-Chunk方法,通过轻量级潜在世界模型在测试时采样多个候选动作分块并选择最优执行,提升动作分块策略在随机动态下的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11187 2026-06-10 cs.CV 新提交 88%

Next Forcing: Causal World Modeling with Multi-Chunk Prediction

Next Forcing: 基于多块预测的因果世界建模

Gangwei Xu, Qihang Zhang, Jiaming Zhou, Xing Zhu, Yujun Shen, Xin Yang, Yinghao Xu

机构 * Robbyant HUST(华中科技大学) HKUST(香港科技大学) HKUST (GZ)(香港科技大学(广州))

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.CV

AI总结 提出Next Forcing框架,通过多块预测训练目标加速视频生成模型收敛、提升精度并实现推理加速,在多个基准上取得最优结果。

Comments Project page: https://gangweix.github.io/next-forcing/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10934 2026-06-10 cs.AI 新提交 88%

WorldKernel: A World Model is the Coupling Kernel of Admissible Possible Worlds

WorldKernel:世界模型是可能世界的耦合核

Fabio Rovai

机构 * The Tesseract Academy(泰塞克特学院)

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract);分类 cs.AI

AI总结 本文发现强预测器在反事实耦合上失效,提出将世界模型建模为可能世界上的半正定耦合核,其非对角元编码反事实信息,并通过半正定性约束和逻辑公理实现高效推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07127 2026-06-08 cs.LG 新提交 88%

Learning Explicit Behavioral Models with Adaptive Questions and World-Model Probes

通过自适应问题和世界模型探针学习显式行为模型

Hikaru Shindo, Yu Deng, Teng Cao, Quentin Delfosse, Christopher Tauchmann, Jannis Blüml, Gopika Sudhakaran, Kristian Kersting

机构 * Artificial Intelligence and Machine Learning Lab(人工智能与机器学习实验室) Technical University of Darmstadt(德累斯顿技术大学) Hessian Center for Artificial Intelligence (hessian.AI)(黑森人工智能中心) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) Department of Computer Science(计算机科学系) Centre for Cognitive Science(认知科学中心)

专题命中 通用世界模型 :world-model(title,abstract);world-model(title,abstract);分类 cs.LG

AI总结 提出显式符号行为模型(ESBM),通过自适应问题和世界模型探针将任务性能与可解释机制结合,在Atari任务中学习高分策略并生成显式答案和机制预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19709 2026-08-21 cs.NI 新提交 88%

RFWM: Physics-Guided World Model for Dynamic Wireless Radiance Field Generation

RFWM:用于生成动态无线辐射场的物理引导世界模型

Zijiu Yang, Qianqian Yang

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract)

AI总结 本研究针对动态未知环境下RF场建模泛化难的问题,提出物理引导的RF世界模型RFWM,采用两阶段训练策略,在新构建的基准上实现了优于现有最优方法的RF场生成性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10222 2026-08-12 eess.SP 新提交 88%

A JEPA-Based Field-Layer World Model for Bridging Channel Prediction and Estimation

一种基于JEPA的场层世界模型,用于衔接信道预测与估计

Yuzhi Yang, Brahim Mefgouda, Hang Zou, Lina Bariah, Anis Bara, Yuhuan Lu, Hao Zhang, MérouaneDebbah

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract)

AI总结 该研究针对MIMO-OFDM无线系统中CSI预测鲁棒性差的问题,提出基于JEPA的场层世界模型,通过多尺度对齐策略实现跨频段重构,显著提升了波束成形增益。

Comments submitted to IEEE trans

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17560 2026-07-21 cs.AI 新提交 88%

Reinforcement Learning: From Algorithms To Foundation Models

强化学习:从算法到基础模型

Zihan Ding

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 本文从游戏算法和基础模型时代的RL两个视角展开研究,前者聚焦多智能体RL中相关概念的相互作用,后者利用生成和基础模型丰富序列决策,开发多种模型并进行相关研究,呈现RL在复杂序列域的统一视图,突出其连接多方面的作用。

Comments Princeton University PhD Thesis 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03942 2026-07-07 eess.SY cs.SY 新提交 88%

ThermoForce: A Physics-Structured Interventional World Model for Building HVAC Control

ThermoForce:用于构建暖通空调控制的物理结构干预世界模型

Yifan Wang

专题命中 通用世界模型 :world model(title,abstract);world model(title,abstract)

AI总结 研究建筑暖通空调系统模型预测控制所需的热模型问题。提出ThermoForce,将时间序列基础模型冻结为被动自由响应,学习物理结构的强制响应算子,在相关干预中表现出色,嵌入MPC可降低热不适与能耗。

Comments 23 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02846 2026-07-07 cs.AI 新提交 87%

Object-Centric Environment Modeling for Agentic Tasks

用于智能任务的以对象为中心的环境建模

Yiyang Li, Tianyi Ma, Zehong Wang, Yijun Ma, Yanfang Ye

机构 * University of Notre Dame(圣母大学)

专题命中 通用世界模型 :environment model(title,abstract);world model(abstract);world models(abstract);world model(abstract)

AI总结 研究大语言模型智能体经验维护难题,提出以对象为中心的环境建模(OCM),将经验组织成可执行模型,含对象知识与过程知识,在线更新并验证,实验表明其能提升智能体性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02365 2026-08-04 cs.AI cs.LG cs.RO 新提交 87%

Faster-WAM: Do World Action Models Need Deep Action Modules?

Faster-WAM:世界动作模型需要深度动作模块吗?

Liheng Ma, Rui Heng Yang, Zhanguang Zhang, Mateo Clemente, Ziwen Hu, Tongtong Cao, Yingxue Zhang

机构 * Huawei Noah’s Ark Lab(华为诺亚方舟实验室) Huawei Celia Team(华为Celia团队) Labs(2012实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 针对现有WAMs动作模块深度绑定视频骨干网络导致延迟高的问题,提出以视频为中心的DoT架构,构建Faster-WAM,实现低延迟、高性能与强泛化,较Fast-WAM提速3.2倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27036 2026-07-30 cs.CV cs.LG 新提交 87%

Mitigating Compounding Error via Video Representation Regularization

通过视频表示正则化缓解复合误差

Taiye Chen, Qi Zhang, Yisen Wang

机构 * Peking University(北京大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 针对视频扩散世界模型自回归生成的复合误差问题,研究发现其与表示维度崩溃相关,提出视频表示正则化方法,在VBench指标上显著优于Diffusion Forcing,提升了长视频生成的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20867 2026-06-23 cs.CV cs.AI 新提交 87%

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation

FOCA: 面向未来的条件化用于数据高效的视觉-语言-动作适应

Duc Minh Nguyen, Nghiem Tuong Diep, Binh Gia Nguyen, Trong-Bao Ho, Doanh Le, Tan Q. Nguyen, Thien-Loc Ha, Nhiem Tran, Bao Thach, Nhat X. Tran, Tuan A. Tran, Artur Habuda, Philip Lund Møller, Tran Nguyen Le, Daniel Sonntag, Matthias Niepert, Khoa D. Doan, Vu Duong, Hung Ngo, Minh N. Vu, Duy M. H. Nguyen, An Thai Le, Ngo Anh Vien

机构 * Center for AI Research, VinUniversity, Vietnam University of Utah, USA German Research Center for Artificial Intelligence (DFKI) Technical University of Denmark, Denmark University of Oldenburg, Germany University of Stuttgart, Germany Max Planck Research School for Intelligent Systems (IMPRS-IS), Germany

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 提出FOCA框架,结合未来交互嵌入预测与目标观测隐式对齐,实现数据高效的VLA少样本适应,在LIBERO、RoboCasa和真实机器人上取得新最优结果。

Comments Accepted at ICML 2026. Project page: https://focavla.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20781 2026-06-23 cs.RO cs.CV 新提交 87%

World Action Models: A Survey

世界行动模型:综述

Qiuhong Shen, Shihua Zhang, Yue Liao, Qi Li, Zhenxiong Tan, Shizun Wang, Shuicheng Yan, Xinchao Wang

机构 * National University of Singapore(新加坡国立大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 本文综述世界行动模型(WAMs),通过两种互补视角组织现有工作,揭示其设计权衡与未来趋势。

Comments 57 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12688 2026-06-16 cs.LG cs.AI cs.DC 新提交 87%

M*: A Modular, Extensible, Serving System for Multimodal Models

M*: 一个模块化、可扩展的多模态模型服务系统

Atindra Jha, Naomi Sagan, Keisuke Kamahori, Irmak Sivgin, Rohan Sanda, Steven Gao, Mark Horowitz, Luke Zettlemoyer, Olivia Hsu, Jure Leskovec, Baris Kasikci, Stephanie Wang

机构 * Stanford University(斯坦福大学) University of Washington(华盛顿大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 提出M*系统,通过将模型表示为数据流图并引入Walk Graph抽象,支持多模态复合模型的高效服务,在多个任务上降低延迟并提升吞吐量。

Comments The codebase is available at https://github.com/mstar-project/mstar

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14490 2026-08-17 cs.AI 新提交 86%

Twin: Playing an Unknown Game with a Test-Time Digital Twin

Twin:利用测试时数字孪生玩未知游戏

Alexy Skoutnev, Kirill Acharya, Gaston Longhitano, Madeleine Udell, Kevin Ellis, Iddo Drori

专题命中 通用世界模型 :world-model(abstract,abstract_cn);world-model(abstract,abstract_cn);world model(abstract);world model(abstract)

AI总结 该研究提出Twin系统,通过测试时数字孪生构建可执行世界模型,在ARC-AGI-3等未知游戏中通关率达97.8%,效率优于人类,核心是通过模拟交互和反例修复推断游戏规则与目标。

Comments Project website with action-by-action replays of all 25 runs: https://arc-agi-3-twin.vercel.app/ Code: AGI-3" target="_blank" rel="noopener">https://github.com/Alexyskoutnev/TWIN-ARC-AGI-3

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06291 2026-07-08 cs.CV cs.HC 新提交 86%

AlayaWorld: Long-Horizon and Playable Video World Generation

AlayaWorld:长期可玩视频世界生成

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao

机构 * Alaya Lab(阿莱亚实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 研究旨在解决游戏世界构建难题,核心方法是提出AlayaWorld全栈开源框架,可实现开放式实时交互,统一开发流程,主要贡献是为生成世界模型研究和应用奠定实践基础。

Comments Authors are listed alphabetically by the first name and their role. See the contribution section for details

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02517 2026-07-03 cs.CV 新提交 86%

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

WorldDirector: 构建具有持久动态记忆的可控世界模拟器

Hanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang, Qingyan Bai, Ka Leong Cheng, Yue Yu, Yixuan Li, Yihao Meng, Zichen Liu, Yanhong Zeng, Yujun Shen, Qifeng Chen

机构 * HKUST(香港科技大学) Ant Group(蚂蚁集团) ZJU(浙江大学) CUHK(香港中文大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 提出WorldDirector框架,通过LLM协调3D轨迹与相机运动作为视频生成控制信号,解耦语义运动编排与视觉生成,实现持久动态对象记忆和自由视角探索。

Comments Project Page: https://worlddirector.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27964 2026-06-29 cs.CV 新提交 86%

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control

Directing the World: 具有组合式人体-相机控制的快速自回归视频生成

Haoyuan Wang, Yabo Chen, Haibin Huang, Chi Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence, China Telecom (TeleAI)(中国电信人工智能研究院(TeleAI))

专题命中 通用世界模型 :world model(abstract);world models(abstract);world-model(abstract);world model(abstract)

AI总结 提出一种快速自回归框架,通过解耦控制学习并保持统一视频先验,实现人体运动和相机轨迹的组合控制,支持稳定长程生成与高视觉质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27364 2026-06-26 cs.CV 新提交 86%

PhysiFormer: Learning to Simulate Mechanics in World Space

PhysiFormer: 在世界空间中学习模拟力学

Yiming Chen, Yushi Lan, Andrea Vedaldi

专题命中 通用世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 提出PhysiFormer,一种直接在世界坐标下通过去噪扩散过程预测3D物体顶点轨迹的扩散Transformer,无需归纳偏置即可生成物理合理的刚性和弹性运动,并泛化到混合材料、未见几何和更多物体。

Comments Project page: https://yimingc9.github.io/physiformer

详情

展开后加载摘要…

URL PDF HTML 收藏