arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 53 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 视频世界模型 53 篇

2606.09507 2026-06-09 cs.CV 新提交 94%

Prisma-World: Camera-Controllable Multi-Agent Video World Model

Prisma-World: 相机可控的多智能体视频世界模型

Huiqiang Sun, Zhan Peng, Size Wu, Kun Wang, Kang Liao, Dianyi Wang, Xingyu Zeng, Sheng Jin, Yangguang Li, Zhiguo Cao, Ziwei Liu, Wei Li

机构 * School of AIA, HUST(华中科技大学人工智能与自动化学院) S-Lab, NTU(南洋理工大学S-Lab) SenseTime Research(商汤科技研究院) FDU(复旦大学) SUAT(深圳大学) HKU(香港大学) CUHK(香港中文大学)

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 提出Prisma-World,通过联合几何感知去噪过程实现多智能体视频生成中的跨视角一致性,支持灵活智能体数量和相机控制。

Comments Project page: https://huiqiang-sun.github.io/prisma-world/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06508 2026-05-26 cs.RO 94%

World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy

World-VLA-Loop: 视频世界模型与VLA策略的闭环学习

Xiaokang Liu, Zechen Bai, Hai Ci, Kevin Yuchen Ma, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室)

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 提出World-VLA-Loop框架,通过状态感知视频世界模型联合预测未来帧和二元奖励,并采用协同进化范式迭代优化VLA策略,减少对真实环境交互的依赖。

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05138 2026-03-31 cs.CV 94%

VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control

VerseCrafter:基于4D几何控制的动态真实视频世界模型

Sixiao Zheng, Minghao Yin, Wenbo Hu, Xiaoyu Li, Ying Shan, Yanwei Fu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) HKU(香港大学) ARC Lab, Tencent PCG(腾讯PCG ARC实验室)

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 本文提出VerseCrafter,通过4D几何控制生成动态真实视频,相比传统方法更精确控制相机和多物体运动。

Comments Project Page: https://sixiaozheng.github.io/VerseCrafter_page/, Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26037 2026-07-29 cs.CV cs.GR 新提交 94%

Wonder: Video World Model Done Better

Wonder:改进的视频世界模型

Jiacong Xu, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli, Vishal M. Patel, Yiqun Mei

机构 * Adobe Research(Adobe研究院) Johns Hopkins University(约翰·霍普金斯大学)

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 Wonder是用于实时相机可控世界探索的通用视频世界模型,通过系统级协同设计,包括新型相机条件设定、高效内存机制等,能合成多样视频,支持视频条件生成,保持长时间的连贯视觉效果。

Comments Project Page: https://wonder-world-model.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13489 2026-08-14 cs.CV cs.RO 新提交 94%

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

DreamX-Phi 1.0:面向机器人操控的动作条件视频世界模型

DreamX Team, Rui Chen, Xiangxiang Chu, Geng Li, Jifan Li, Qingfeng Shi, Datao Tang, Jing Tang, Jun Wang, Pengfei Zhang

机构 * DreamX

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 DreamX-Phi 1.0是面向机器人操控的动作条件视频世界模型,通过几何编码、深度分支等优化,在WorldArena 2.0挑战赛获Track1第一、Track2第二,模型与代码将公开。

Comments Code: https://github.com/AMAP-ML/DreamX-Phi

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15185 2026-05-15 cs.CV cs.AI 94%

Quantitative Video World Model Evaluation for Geometric-Consistency

几何一致性定量视频世界模型评估

Jiaxin Wu, Yihao Pi, Yinling Zhang, Yuheng Li, Xueyan Zou

机构 * Tsinghua University - IEI Lab(清华大学-IEI实验室) UW-Madison(威斯康星大学麦迪逊分校) Adobe Research(Adobe研究院)

专题命中 视频世界模型 :world model(title,abstract);video world model(title);world model(title,abstract);video world model(title)

AI总结 本文提出PDI-Bench框架,通过分割与点跟踪获取物体中心观测,利用单目重建获取3D坐标,计算投影几何残差以评估生成视频的几何一致性,揭示视频生成器的特定失败模式。

Comments 12 pages, 5 figures. Project page : https://pdi-bench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13376 2026-06-18 cs.CV 新提交 93%

MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold

MoVerse: 基于全景高斯支架的实时视频世界建模

Yang Zhou, Ziheng Wang, Yuqin Lu, Haofeng Liu, Jun Liang, Shengfeng He, Jing Li

机构 * South China University of Technology Columbia University Orange Team, Youku Moku-Lab, HUJING Digital Media \& Entertainment Group Singapore Management University

专题命中 视频世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 提出MoVerse,从单张窄视场图像实时构建可交互漫游的360度全景世界,通过拓扑感知扩散补全视场、全景几何残差预测生成3D高斯支架,并结合双向扩散教师蒸馏为因果自回归学生实现低延迟视频渲染。

Comments Project Page: https://orange-3dv-team.github.io/MoVerse/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20698 2026-06-23 cs.RO 新提交 91%

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model

SafeDojo:基于交互世界模型的安全强化学习用于视觉-语言-动作模型

Kai Tang, Peidong Jia, Zhong Chu, Jixian Wu, Rui Ma, Jiajun Cao, Fangyuan Zhao, Sixiang Chen, Yichen Guo, Xiaowei Chi, Chun-Kai Fan, Kevin Zhang, Jinchang Xu, Fubing Yang, Weishi Mi, Xiaozhu Ju, Jian Tang, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室) Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) Nanyang Technological University(南洋理工大学) Hong Kong University of Science and Technology(香港科技大学) University of Electronic Science and Technology of China(电子科技大学)

专题命中 视频世界模型 :world model(title,abstract);world model(title,abstract);video world model(abstract);video world model(abstract)

AI总结 提出SafeDojo,首个基于模型的安全强化学习框架,通过交互式视频世界模型进行想象学习安全动作,结合解耦的任务奖励和安全代价信号,在SafeLIBERO和真实机器人上取得最佳安全成功率。

Comments 20 pages, 5 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01060 2026-07-16 cs.RO 版本更新 89%

RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation

RoboWorld: 用于通用机器人策略评估的快速可靠神经模拟器

Byeongguk Jeon, Seonghyeon Ye, JaeHyeok Doo, Sungdong Kim, Minjoon Seo, Hyungmok Son, Kimin Lee

机构 * KAIST(韩国科学技术院) Config

专题命中 视频世界模型 :world model(abstract);world models(abstract);world-model(abstract);video world model(abstract)

AI总结 提出RoboWorld自动化评估流程,结合快速自回归视频世界模型和任务进度感知视觉语言模型评分,通过Step Forcing减少训练-测试不匹配,实现与真实世界评估高度一致。

Comments Project page: https://byeongguks.github.io/RoboWorld/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02330 2026-08-03 cs.CV cs.AI cs.LG 版本更新 87%

ActionParty: Multi-Subject Action Binding in Generative Video Games

ActionParty:生成视频游戏中的多主体动作绑定

Alexander Pondaven, Ziyi Wu, Igor Gilitschenski, Philip Torr, Sergey Tulyakov, Fabio Pizzati, Aliaksandr Siarohin

机构 * Snap Research(Snap研究院) University of Oxford(牛津大学) University of Toronto(多伦多大学) MBZUAI(穆罕默德·本·扎耶德人工智能大学)

专题命中 视频世界模型 :world model(abstract);world models(abstract);video world model(abstract);world model(abstract)

AI总结 本文提出ActionParty,一种可控制多主体的生成视频游戏世界模型,通过引入主体状态标记和空间偏置机制,提升动作关联准确性与身份一致性。

Comments ECCV 2026 - Project page: https://action-party.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19191 2026-07-22 cs.CV cs.AI cs.LG 新提交 83%

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

ABot-World-0:在单台桌面GPU上进行无限交互式世界展开

Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu, Zheng Zhou, Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo

专题命中 视频世界模型 :world model(abstract);video world model(abstract);world model(abstract);video world model(abstract)

AI总结 介绍ABot-World-0这一用于实时长视野闭环交互的动作条件视频世界模型,利用多源数据学习世界动态。通过多种技术提炼模型,设计控制界面与部署堆栈,在单台桌面GPU上实现高效视频流传输,实验验证其有竞争力的可控性和世界演变能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09038 2025-02-28 cs.CV cs.AI cs.GR cs.LG 83%

Do generative video models understand physical principles?

Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini, Robert Geirhos

专题命中 视频世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07600 2025-05-22 cs.CV cs.RO 81%

PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning

Angel Villar-Corrales, Sven Behnke

机构 * Autonomous Intelligent Systems, Computer Science Institute VI – Intelligent Systems(智能系统) Robotics, Center for Robotics(机器人学) the Lamarr Institute for Machine Learning(拉马尔机器学习研究所) Artificial Intelligence, University of Bonn, Germany(人工智能,波恩大学,德国)

专题命中 视频世界模型 :latent dynamics(title);world model(abstract);world model(abstract);分类 cs.CV、cs.RO

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.14718 2025-07-15 cs.CV 81%

A Survey on Future Frame Synthesis: Bridging Deterministic and Generative Approaches

Ruibo Ming, Zhewei Huang, Jingwei Wu, Zhuoxuan Ju, Daxin Jiang, Jianming Hu, Lihui Peng, Shuchang Zhou

机构 * Tsinghua University(清华大学) StepFun Peking University(北京大学) Megvii Technology(旷视科技)

专题命中 视频世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments TMLR 2025/07

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14948 2025-05-22 cs.CV cs.AI cs.LG 73%

Programmatic Video Prediction Using Large Language Models

Hao Tang, Kevin Ellis, Suhas Lohit, Michael J. Jones, Moitreya Chatterjee

机构 * Mitsubishi Electric Research Laboratories (MERL)(三菱电机研究实验室(MERL))

专题命中 视频世界模型 :world model(abstract);world model(abstract);分类 cs.AI、cs.LG、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09082 2025-12-01 cs.CV 69%

Taming generative video models for zero-shot optical flow extraction

驯服生成视频模型以实现零样本光流提取

Seungwoo Kim, Khai Loong Aw, Klemen Kotar, Cristobal Eyzaguirre, Wanhee Lee, Yunong Liu, Jared Watrous, Stefan Stojanov, Juan Carlos Niebles, Jiajun Wu, Daniel L. K. Yamins

机构 * Stanford University(斯坦福大学)

专题命中 视频世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 本文提出KL-tracing方法,通过反事实提示实现生成视频模型的零样本光流提取,无需微调即可在现实和合成数据集上竞争现有最佳模型。

Comments Project webpage: https://neuroailab.github.io/projects/kl_tracing

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16484 2025-11-21 cs.CV 69%

Flow and Depth Assisted Video Prediction with Latent Transformer

基于流和深度的视频预测与潜在变换器

Eliyas Suleyman, Paul Henderson, Eksan Firkat, Nicolas Pugeault

专题命中 视频世界模型 :world model(abstract);world model(abstract);分类 cs.CV

AI总结 本文提出基于流和深度的视频预测方法,通过整合点流和深度信息提升遮挡场景下的预测性能和背景运动准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17359 2025-05-30 cs.CV 69%

Position: Interactive Generative Video as Next-Generation Game Engine

Jiwen Yu, Yiran Qin, Haoxuan Che, Quande Liu, Xintao Wang, Pengfei Wan, Di Zhang, Xihui Liu

机构 * HKU(香港大学) HKUST(香港理工大学) Kuaishou Technology(快手科技)

专题命中 视频世界模型 :world model(abstract);world model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.16084 2022-03-31 cs.CV cs.LG 62%

STRPM: A Spatiotemporal Residual Predictive Model for High-Resolution Video Prediction

Zheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma, Wen Gao

专题命中 视频世界模型 :predictive model(title,abstract);分类 cs.LG、cs.CV;predictive models(abstract)

Comments This work has been accepted by CVPR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.11894 2022-08-09 cs.CV cs.LG cs.RO 56%

MaskViT: Masked Visual Pre-Training for Video Prediction

Agrim Gupta, Stephen Tian, Yunzhi Zhang, Jiajun Wu, Roberto Martín-Martín, Li Fei-Fei

专题命中 视频世界模型 :分类 cs.LG、cs.CV、cs.RO;predictive model(abstract);predictive models(abstract)

Comments Project page: https://maskedvit.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.11062 2021-05-25 cs.CV cs.LG 56%

Taylor saves for later: disentanglement for video prediction using Taylor representation

Ting Pan, Zhuqing Jiang, Jianan Han, Shiping Wen, Aidong Men, Haiying Wang

专题命中 视频世界模型 :latent dynamics(abstract);分类 cs.LG、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.06701 2024-11-20 cs.CV cs.AI cs.LG 50%

S-HR-VQVAE: Sequential Hierarchical Residual Learning Vector Quantized Variational Autoencoder for Video Prediction

Mohammad Adiban, Kalin Stefanov, Sabato Marco Siniscalchi, Giampiero Salvi

专题命中 视频世界模型 :分类 cs.AI、cs.LG、cs.CV;predictive model(abstract)

Comments 12 pages, 6 figures, 5 tables. Accepted for publication on IEEE Transactions on Multimedia on 2024-11-19

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18964 2024-09-30 cs.CV cs.AI cs.LG 50%

PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation

Shaowei Liu, Zhongzheng Ren, Saurabh Gupta, Shenlong Wang

专题命中 视频世界模型 :分类 cs.AI、cs.LG、cs.CV;simulation model(abstract)

Comments Accepted to ECCV 2024. Project page: https://stevenlsw.github.io/physgen/

详情

展开后加载摘要…

URL PDF HTML 收藏