arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 6450 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 4308 篇

2512.04537 2025-12-05 cs.CV 81%

X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale

X-Humanoid:将人类视频机器人化以生成大规模人形视频

Pei Yang, Hai Ci, Yiren Song, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 X-Humanoid通过生成式视频编辑方法,将人类视频机器人化,生成大规模人形视频数据集,提升人形机器人研究的训练效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04515 2025-12-05 cs.CV 81%

EgoLCD: Egocentric Video Generation with Long Context Diffusion

EgoLCD:基于长上下文扩散的自体视频生成

Liuzhou Zhang, Jiarui Ye, Yuanlei Wang, Ming Zhong, Mingju Cao, Wanke Xia, Bowen Zeng, Zeyu Zhang, Hao Tang

机构 * Peking University(北京大学) Sun Yat-sen University(中山大学) Zhejiang University(浙江大学) Chinese Academy of Sciences(中国科学院) Tsinghua University(清华大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 EgoLCD通过结合长时稀疏KV缓存和LoRA扩展的注意力机制,实现了高效稳定的自体长上下文视频生成,提升了感知质量和时间一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03058 2025-12-04 cs.LG 81%

Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding

自注意力机制中令牌的动力学特性及位置编码的影响

Duy-Tung Pham, An The Nguyen, Viet-Hoang Tran, Nhan-Phu Chung, Xin T. Tong, Tan M. Nguyen, Thieu N. Vo

机构 * FPT Software AI Center(FPT软件AI中心) National University of Singapore(国立新加坡大学) Ho Chi Minh University of Economics(胡志明经济大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文研究了Transformer中令牌的动力学特性,揭示了位置编码对模型性能的影响,并提出了改进方法以缓解收敛问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02457 2025-12-04 cs.CV 81%

Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation

听觉有助于视觉吗?研究音频视频联合去噪用于视频生成

Jianzong Wu, Hao Lian, Dachao Hao, Ye Tian, Qingyu Shi, Biaolong Chen, Hao Jiang, Yunhai Tong

机构 * Peking University(北京大学) Alibaba Group(阿里巴巴集团)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出通过音频视频联合去噪提升视频生成质量,验证了跨模态训练对构建更物理基础的世界模型的潜力。

Comments Project page at https://jianzongwu.github.io/projects/does-hearing-help-seeing/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02016 2025-12-02 cs.CV 81%

Objects in Generated Videos Are Slower Than They Appear: Models Suffer Sub-Earth Gravity and Don't Know Galileo's Principle...for now

生成视频中的物体比看起来更慢:模型遭受次地球重力并目前还不知道伽利略原理...

Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan, Anand Bhattad

机构 * Indian Institute of Science(印度科学研究院) Johns Hopkins University(约翰霍普金斯大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 研究发现视频生成器在重力表示上存在偏差,通过针对性适配可提升其重力模拟精度。

Comments https://gravity-eval.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19836 2025-11-26 cs.CV 81%

4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models

4DWorldBench: 一种用于3D/4D世界生成模型的综合评估框架

Yiting Lu, Wei Luo, Peiyan Tu, Haoran Li, Hanxin Zhu, Zihao Yu, Xingrui Wang, Xinyi Chen, Xinge Peng, Xin Li, Zhibo Chen

机构 * University of Science and Technology of China(中国科学技术大学) Zhejiang University(浙江大学) Beijing Zhongguancun Academy(北京中关村学院)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 4DWorldBench提出了一种用于评估3D/4D世界生成模型的综合框架,通过四个维度评估模型的感知质量、条件对齐、物理真实性和一致性,支持多模态输入并整合多种评估方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11773 2025-11-12 cs.CV cs.HC 81%

AgentSense: Virtual Sensor Data Generation Using LLM Agents in Simulated Home Environments

Zikang Leng, Megha Thukral, Yaqi Liu, Hrudhai Rajasekhar, Shruthi K. Hiremath, Jiaman He, Thomas Plötz

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22200 2025-10-29 cs.CV 81%

LongCat-Video Technical Report

Meituan LongCat Team, Xunliang Cai, Qilong Huang, Zhuoliang Kang, Hongyu Li, Shijun Liang, Liya Ma, Siyu Ren, Xiaoming Wei, Rixu Xie, Tong Zhang

机构 * Meituan LongCat Team(美团LongCat团队)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21682 2025-10-27 cs.CV cs.GR 81%

WorldGrow: Generating Infinite 3D World

Sikuang Li, Chen Yang, Jiemin Fang, Taoran Yi, Jia Lu, Jiazhong Cen, Lingxi Xie, Wei Shen, Qi Tian

机构 * MoE Key Lab of Artificial Intelligence, School of Computer Science, SJTU(人工智能前沿实验室,计算机科学学院,上海交通大学) Huawei Inc.(华为公司) Huazhong University of Science and Technology(华中科技大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments Project page: https://world-grow.github.io/ Code: https://github.com/world-grow/WorldGrow

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13809 2025-10-16 cs.CV 81%

PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning

Sihui Ji, Xi Chen, Xin Tao, Pengfei Wan, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments Project Page: https://sihuiji.github.io/PhysMaster-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06209 2025-10-08 cs.CV 81%

Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models

Jiahao Wang, Zhenpei Yang, Yijing Bai, Yingwei Li, Yuliang Zou, Bo Sun, Abhijit Kundu, Jose Lezama, Luna Yue Huang, Zehao Zhu, Jyh-Jing Hwang, Dragomir Anguelov, Mingxing Tan, Chiyu Max Jiang

机构 * Johns Hopkins University(约翰霍普金斯大学) Waymo Google DeepMind(谷歌DeepMind)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02110 2025-10-03 cs.SD cs.LG eess.AS 81%

SoundReactor: Frame-level Online Video-to-Audio Generation

Koichi Saito, Julian Tanke, Christian Simon, Masato Ishii, Kazuki Shimada, Zachary Novack, Zhi Zhong, Akio Hayakawa, Takashi Shibuya, Yuki Mitsufuji

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25161 2025-09-30 cs.CV 81%

Rolling Forcing: Autoregressive Long Video Diffusion in Real Time

Kunhao Liu, Wenbo Hu, Jiale Xu, Ying Shan, Shijian Lu

机构 * Nanyang Technological University(南洋理工大学) ARC Lab, Tencent PCG(腾讯PCG ARC实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments Project page: https://kunhao-liu.github.io/Rolling_Forcing_Webpage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24441 2025-09-30 cs.CV 81%

NeoWorld: Neural Simulation of Explorable Virtual Worlds via Progressive 3D Unfolding

Yanpeng Zhao, Shanyan Guan, Yunbo Wang, Yanhao Ge, Wei Li, Xiaokang Yang

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能联合实验室、人工智能研究院、上海交通大学) vivo Mobile Communication Co., Ltd.(vivo移动通信有限公司)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20998 2025-09-26 cs.AI 81%

CORE: Full-Path Evaluation of LLM Agents Beyond Final State

Panagiotis Michelakis, Yiannis Hadjiyiannis, Dimitrios Stamoulis

机构 * Synkrasis Labs(Synkrasis实验室) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments Accepted: LAW 2025 Workshop NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13798 2025-09-15 cs.CL cs.AI cs.IT math.IT 81%

Slaves to the Law of Large Numbers: An Asymptotic Equipartition Property for Perplexity in Generative Language Models

Tyler Bell, Avinash Mudireddy, Ivan Johnson-Eversoll, Soura Dasgupta, Raghu Mudumbai

机构 * University of Iowa(爱荷华大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11299 2025-09-09 cs.CL cs.AI 81%

BriLLM: Brain-inspired Large Language Model

Hai Zhao, Hongqiu Wu, Dongjie Yang, Anni Zou, Jiale Hong

机构 * AGI Institute(AGI研究院) Computer School(计算机学院) Shanghai Jiao Tong University(上海交通大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04600 2025-09-08 cs.CV 81%

WATCH: World-aware Allied Trajectory and pose reconstruction for Camera and Human

Qijun Ying, Zhongyuan Hu, Rui Zhang, Ronghui Li, Yu Lu, Zijiao Zeng

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00258 2025-08-29 cs.AI q-bio.NC 81%

Possible Principles for Aligned Structure Learning Agents

Lancelot Da Costa, Tomáš Gavenčiak, David Hyland, Mandana Samiei, Cristian Dragos-Manta, Candice Pattisapu, Adeel Razi, Karl Friston

机构 * VERSES AI Research Lab(VERSES AI研究实验室) Charles University(查尔斯大学) University of Oxford(牛津大学) Mila, Quebec AI Institute(魁北克人工智能研究院) McGill University(麦吉尔大学) University of Montreal(蒙特利尔大学) University College London(伦敦大学学院) Monash University(莫纳什大学) CIFAR Azrieli Global Scholars Program(CIFAR阿兹里埃利全球学者计划)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments 24 pages of content, 33 with references; accepted version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16512 2025-08-25 cs.CV 81%

Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation

Chun-Peng Chang, Chen-Yu Wang, Julian Schmidt, Holger Caesar, Alain Pagani

机构 * Delft University of Technology(代尔夫特理工大学) German Research Center for Artificial Intelligence(德国人工智能研究中心) Mercedes-Benz(梅赛德斯-奔驰)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15859 2025-08-25 q-bio.NC cs.AI cs.CL 81%

Beyond Individuals: Collective Predictive Coding for Memory, Attention, and the Emergence of Language

Tadahiro Taniguchi

机构 * Graduate School of Informatics(信息学研究科)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Journal ref Cognitive Neuroscience, 1-2 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15013 2025-08-22 cs.AI q-bio.NC 81%

Goals and the Structure of Experience

Nadav Amir, Stas Tiomkin, Angela Langdon

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21755 2025-08-21 cs.CV 81%

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Dian Zheng, Ziqi Huang, Hongbo Liu, Kai Zou, Yinan He, Fan Zhang, Lulu Gu, Yuanhan Zhang, Jingwen He, Wei-Shi Zheng, Yu Qiao, Ziwei Liu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) S-Lab, Nanyang Technological University(南洋理工大学S实验室) Sun Yat-sen University(中山大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments Equal contributions from first two authors. Project page: https://vchitect.github.io/VBench-2.0-project/ Code: https://github.com/Vchitect/VBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13154 2025-08-19 cs.CV 81%

4DNeX: Feed-Forward 4D Generative Modeling Made Easy

Zhaoxi Chen, Tianqi Liu, Long Zhuo, Jiawei Ren, Zeng Tao, He Zhu, Fangzhou Hong, Liang Pan, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments Project Page: https://4dnex.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10770 2025-08-15 cs.CV 81%

From Diagnosis to Improvement: Probing Spatio-Physical Reasoning in Vision Language Models

Tiancheng Han, Yunfei Gao, Yong Li, Wuzhou Yu, Qiaosheng Zhang, Wenqi Shao

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10489 2025-08-15 cs.LG 81%

Learning State-Space Models of Dynamic Systems from Arbitrary Data using Joint Embedding Predictive Architectures

Jonas Ulmen, Ganesh Sundaram, Daniel Görges

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments 6 Pages, Published in IFAC Joint Symposia on Mechatronics & Robotics 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05619 2025-08-08 cs.AI nlin.AO physics.bio-ph physics.comp-ph physics.hist-ph 81%

The Missing Reward: Active Inference in the Era of Experience

Bo Wen

机构 * IBM T.J. Watson Research Center(IBM TJ沃森研究中心)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09144 2025-08-05 cs.CV 81%

$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting

Zhimin Liao, Ping Wei, Ruijie Zhang, Shuaijia Chen, Haoxuan Wang, Ziyang Ren

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家级重点实验室) Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人工智能与机器人学院,西安交通大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19272 2025-07-28 cs.CV 81%

Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception

Marcel Simon, Tae-Ho Kim, Seul-Ki Yeom

机构 * Nota AI GmbH

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments 4 pages, 2 figures, 2 tables

Journal ref 2025 International Conference on Machine Learning Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12547 2025-07-21 cs.CL cs.AI cs.PL 81%

Modeling Open-World Cognition as On-Demand Synthesis of Probabilistic Models

Lionel Wong, Katherine M. Collins, Lance Ying, Cedegao E. Zhang, Adrian Weller, Tobias Gerstenberg, Timothy O'Donnell, Alexander K. Lew, Jacob D. Andreas, Joshua B. Tenenbaum, Tyler Brooke-Wilson

机构 * Stanford University(斯坦福大学) MIT(麻省理工学院) University of Cambridge(剑桥大学) Harvard University(哈佛大学) McGill University(麦吉尔大学) Yale University(耶鲁大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

Comments Presented at CogSci 2025

详情

展开后加载摘要…

URL PDF HTML 收藏