arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 6441 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 4300 篇

2606.27806 2026-07-07 cs.AI 新提交 94%

Agent vs. Parametric World Models: Hybrid Planning for Reliable Language Agents

基于基础迭代语言规划:参数化世界模型如何减少LLM代理中的幻觉传播

Xinyuan Song, Zekun Cai

机构 * Emory University(埃默里大学) The University of Tokyo(东京大学) LocationMind

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出GILP方法,结合小参数化世界模型与LLM推理,通过一致性门控减少幻觉,在规划基准上将幻觉率从0.176降至0.035,成功率从0.668提升至0.838。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23164 2026-07-02 cs.LG 版本更新 94%

MetaOthello: A Controlled Study of Multiple World Models in Transformers

MetaOthello: Transformer中多个世界模型的控制研究

Aviral Chawla, Galen Hall, Juniper Lovato

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 通过引入共享语法但规则或分词不同的Othello变体套件,训练小型GPT处理混合数据,发现Transformer不划分容量为孤立子模型,而是收敛于跨变体因果转移的共享棋盘状态表示。

Comments Camera-ready version. Accepted to the 43rd International Conference on Machine Learning (ICML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22136 2026-06-24 cs.RO 新提交 94%

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

Wh0: 生成式世界模型作为自我中心人类手部操作数据的可扩展来源

Yangtao Chen, Zixuan Chen, Peiyang Wang, Yong-Lu Li, Jing Huo, Jieqi Shi, Yang Gao

机构 * Shanghai Innovation Institute(上海创新研究院) Nanjing University(南京大学) Shanghai Jiaotong University(上海交通大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出Wh0框架,利用生成式视频世界模型产生自我中心人-物交互数据集WM-H,通过手部运动重建和视觉编辑转化为机器人可训练监督,提升预训练灵巧VLA模型在真实任务上的零样本成功率。

Comments Under review. The first three authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24354 2026-06-23 cs.CV 版本更新 94%

SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation

SparseWorld: 通过具有稀疏场景表示的世界模型增强端到端自动驾驶

Ruoyu Wang, Jingke Wang, Yukai Ma, Yuehao Huang, Shuangming Lei, Guanglin Xu, Aixue Ye, Yong Liu

机构 * Institute of Cyber-Systems and Control, Zhejiang University(浙江大学控制系统研究所) Labs, Huawei(华为2012实验室) State Key Laboratory of Industrial Control Technology(国家工业控制技术重点实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出SparseWorld,一种基于稀疏场景表示的轻量级世界模型,通过自回归预测未来地图元素和周围智能体,并利用预测结果优化下游运动预测和轨迹规划,在nuScenes数据集上实现0.05%的碰撞率,达到开放循环规划指标的最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18313 2026-06-23 cs.CV 版本更新 94%

OmniNWM: Omniscient Driving Navigation World Models

OmniNWM:全知驾驶导航世界模型

Bohan Li, Zhuang Ma, Dalong Du, Baorui Peng, Zhujin Liang, Zhenqiang Liu, Xianda Guo, Zheng Zhu, Chao Ma, Yueming Jin, Xin Jin, Hao Zhao, Wenjun Zeng

机构 * Shanghai Jiao Tong University(上海交通大学) Eastern Institute of Technology(东部技术研究所) PhiGent National University of Singapore(新加坡国立大学) Tsinghua University(清华大学) Wuhan University(武汉大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出OmniNWM全知全景导航世界模型,在统一概率框架下处理状态、动作和奖励三个维度,实现多模态全景视频生成、零样本轨迹控制及内在奖励机制,在生成保真度和控制精度上达到SOTA。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20545 2026-06-19 cs.CV 新提交 94%

Current World Models Lack a Persistent State Core

当前世界模型缺乏持久状态核心

Jinpeng Lu, Dexu Zhu, Haoyuan Shi, Linghan Cai, Guo Tang, Yinda Chen, Jie Cao, Duyu Tang, Yi Zhang, Yong Dai, Xiaozhu Ju

机构 * University of Science and Technology of China(中国科学技术大学) Beijing Innovation Center of Humanoid Robotics (X-Humanoid)(北京人形机器人创新中心) NLPR, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室) Independent Researcher(独立研究者) Dresden University of Technology(德累斯顿工业大学) Peking University(北京大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出WRBench基准测试,发现现有世界模型在观测中断时无法维持世界状态演化,强调物理状态核稳定性应成为世界模型设计首要目标。

Comments 39 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16070 2026-06-17 cs.AI 新提交 94%

Mind-Studio: Executable World Models with Lookahead Evaluation for Partially Observable Games

Mind-Studio: 针对部分可观测游戏的可执行世界模型与前向评估

Yifei Dong, Mingen Zheng, Linquan Wu, Jeff Z. Pan, Jiaxin Bai

机构 * Hong Kong University of Science and Technology(香港科技大学) City University of Hong Kong(香港城市大学) University of Edinburgh(爱丁堡大学) Hong Kong Baptist University(香港浸会大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出Mind-Studio框架,利用大语言模型从轨迹合成可执行的pygame风格世界模型,通过K步前向保真度协议评估,在Montezuma's Revenge等游戏中显著提升预测准确性和子目标验证。

Comments 12 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07967 2026-06-09 cs.CV 新提交 94%

DisCo: World Models with Discrete Camera Motion Control

DisCo: 具有离散相机运动控制的世界模型

Hongrui Huang, Junke Wang, Quanhao Li, Yu-Gang Jiang, Zuxuan Wu

机构 * Fudan University(复旦大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出DisCo,通过离散动作原语替代连续相机轨迹作为条件,解决可控视频生成中动作表示纠缠问题,提升动作跟随可靠性,并引入DisCoBench基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05138 2026-06-09 cs.AI 版本更新 94%

Executable World Models for ARC-AGI-3 in the Era of Coding Agents

可执行世界模型在编码智能体时代的ARC-AGI-3应用

Sergey Rodionov

机构 * SingularityNET

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 提出一种编码智能体系统,通过维护可执行Python世界模型、验证观察、重构简化抽象和模型内规划,在ARC-AGI-3游戏中取得初步成果,GPT-5.5高推理下完全解决15个游戏。

Comments 13 pages. Accepted for publication at AGI-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06556 2026-06-08 cs.RO 新提交 94%

Robots Need More than VLA and World Models

机器人需要的不仅仅是VLA和世界模型

Elis Karcini, Faisal Mehrban, Quang Nguyen, Mac Schwager, Arash Ajoudani, Cesar Cadena, Jan Peters, Marco Hutter, Haitham Bou-Ammar

机构 * Motoniq.ai Stanford University(斯坦福大学) Istituto Italiano di Tecnologia(意大利技术研究院) ETH Zurich(苏黎世联邦理工学院) Technical University of Darmstadt(德累斯顿技术大学) UCL Centre for AI(伦敦大学学院人工智能中心)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文认为机器人通用智能的关键瓶颈不仅是策略学习,还缺乏将非结构化行为数据转化为机器人可用监督的机制,并提出了四种缺失的接口组件。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00133 2026-06-02 cs.LG cs.ET 94%

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications

世界模型:架构、方法论、推理范式与应用的全面综述

Arif Hassan Zidan, Yi Pan, Hanqi Jiang, Ruiyu Yan, Wei Ruan, Zihao Wu, Lifeng Chen, Weihang You, Xinliang Li, Bowen Chen, Huawen Hu, Peilong Wang, Sizhuang Liu, Jing Zhang, Siyuan Li, Zhengliang Liu, Yu Bao, Lin Zhao, Lichao Sun, Dajiang Zhu, Xiang Li, Jinglei Lv, Quanzheng Li, Wei Liu, Tianming Liu, Wei Zhang

机构 * School of Computer and Cyber Sciences, Augusta University(奥古斯塔大学计算机与网络科学学院) School of Computing, University of Georgia(佐治亚大学计算机学院) Department of Biomedical Engineering, New Jersey Institute of Technology(新泽西理工学院生物医学工程系) Department of Radiology, Massachusetts General Hospital, Harvard Medical School(麻省总医院放射科,哈佛医学院) Department of Computer Science and Engineering, University of Texas at Arlington(德克萨斯大学阿灵顿分校计算机科学与工程系) Department of Graduate Psychology, James Madison University(詹姆斯麦迪逊大学研究生心理学系) Computer Science and Engineering, Lehigh University(莱斯大学计算机科学与工程系) School of Biomedical Engineering, The University of Sydney(悉尼大学生物医学工程学院) Tandon School of Engineering, New York University(纽约大学泰坦工程学院) Department of Radiation Oncology, City of Hope National Medical Center(城市希望国家医学中心放射肿瘤科) Department of Mayo Clinic Comprehensive Cancer Center, Mayo Clinics(梅奥诊所综合癌症中心,梅奥诊所) Savannah River Ecology Laboratory (SREL), University of Georgia(萨凡纳河生态实验室(SREL),佐治亚大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出一个多轴分类法,从架构、方法论、推理策略和应用领域四个维度系统综述世界模型,涵盖从早期认知科学基础到PlaNet、Dreamer系列、MuZero、Sora等里程碑系统,并指出未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06331 2026-06-02 cs.CV 94%

WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching

WorldCache: 通过异构令牌缓存免费加速世界模型

Weilun Feng, Guoxin Fan, Haotong Qin, Mingqiang Wu, Yuqi Li, Xiangqi Li, Zhulin An, Libo Huang, Dingrui Wang, Longlong Liao, Michele Magno, Yongjun Xu, Chuanguang Yang

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 针对扩散世界模型中令牌异质性和非均匀时间动态导致的推理慢问题,提出基于曲率引导的异构令牌预测和混沌优先自适应跳过的缓存框架WorldCache,实现高达3.7倍加速并保持98%的推出质量。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24578 2026-05-26 cs.CV 94%

World Models as Group Actions

世界模型作为群作用

Zijie Wang, Wei Zhang, Weiming Zhang, Fanqi Zhang, Xiao Tan, Yipeng Qin, Guanbin Li

机构 * Sun Yat-sen University(中山大学) Shenzhen Loop Area Institute(深圳河套学院) Baidu Inc.(百度公司) Cardiff University(卡迪夫大学) Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出将动作条件世界建模形式化为状态空间上的群作用,通过潜在空间正则化强制执行恒等、逆和组合一致性,并引入群作用一致性(GAC)和群作用鲁棒性(GAR)指标来评估结构正确性和展开稳定性。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16725 2026-05-19 cs.AI 94%

Baba in Wonderland: Online Self-Supervised Dynamics Discovery for Executable World Models

《白兔在奇幻世界:在线自监督动态发现用于可执行世界模型》

SeungWon Seo, DongHeun Han, SeongRae Noh, HyeongYeop Kang

机构 * Korea University(韩国大学) Kyung Hee University(庆熙大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 该研究探讨了在先验错配情况下,如何通过交互证据自监督学习可执行世界模型,引入了Alice系统,通过失败的候选更新作为结构信号,发现并改进动态,从而提升可执行世界模型的学习效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23180 2026-05-19 cs.CV 94%

GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation

GaussianDWM: 基于3D高斯场景表示的统一场景理解和多模态生成驱动世界模型

Tianchen Deng, Xuefeng Chen, Yi Chen, Qu Chen, Yuyao Xu, Lijin Yang, Le Xu, Yu Zhang, Bo Zhang, Wuxiong Huang, Hesheng Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) MEGVII Technology(商汤科技) Mach Drive

专题命中 通用世界模型 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 本文提出基于3D高斯表示的统一驱动世界模型框架,实现3D场景理解和多模态生成,并通过语言引导采样策略和双条件生成模型提升生成效果,实验验证其在nuScenes和NuInteract数据集上的优越性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13740 2026-05-14 cs.LG 94%

Learning POMDP World Models from Observations with Language-Model Priors

从观测中学习POMDP世界模型:利用语言模型先验

Valentin Six, Frederik Panse, Mathis Fajeau, Lancelot Da Costa, Mridul Sharma, Alfonso Amayuelas, Tim Z. Xiao, David Hyland, Philipp Hennig, Bernhard Schölkopf

机构 * Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) IRIIS University of California, Santa Barbara(加州大学圣芭芭拉分校) University of Tübingen(图宾根大学) University of Oxford(牛津大学) ELLIS Institute Tübingen(图宾根ELLIS研究所)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出Pinductor,利用语言模型先验从少量观测-动作轨迹学习POMDP模型,实现高效世界模型学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20289 2026-04-23 cs.CV 94%

X-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models Inference

X-Cache:跨块块缓存用于少步自回归世界模型推理

Yixiao Zeng, Jianlei Zheng, Chaoda Zheng, Shijia Chen, Mingdian Liu, Tongping Liu, Tengwei Luo, Yu Zhang, Boyang Wang, Linkun Xu, Siyuan Lu, Bo Tian, Xianming Liu

机构 * AI Infra Team, XPeng Inc.(AI基础设施团队,小鹏汽车)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出X-Cache,一种无需训练的加速方法,通过跨连续生成块缓存而非去噪步骤来提升少步自回归世界模型的推理效率,实现71%的块跳过率和2.6倍的速度提升,同时保持最小的性能下降。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13824 2026-04-16 cs.LG 94%

Beyond State Consistency: Behavior Consistency in Text-Based World Models

超越状态一致性:文本基础世界模型中的行为一致性

Youling Huang, Guanqiao Chen, Junchi Yao, Lu Wang, Fangkai Yang, Chao Du, ChenZhuo Zhao, Pu Zhao, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang

机构 * Dalian University of Technology(大连理工大学) MBZUAI Peking University(北京大学) Microsoft(微软)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出行为一致性训练方法,通过改进世界模型与真实环境的功能一致性,提升长期对齐效果,实验显示在WebShop和TextWorld中表现优异,同时保持单步预测质量。

Comments 20 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23149 2026-03-25 cs.AI 94%

Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models

描述-然后-行动:通过蒸馏的语言-动作世界模型实现前瞻性智能体导航

Massimiliano Pappa, Luca Romani, Valentino Sacco, Alessio Palma, Stéphane Lathuilière, Fabio Galasso, Xavier Alameda-Pineda, Indro Spinelli

机构 * Sapienza University of Rome, Italy(罗马大学萨皮恩扎) Inria, Univ. Grenoble Alpes, CNRS, LJK(法国国家信息与自动化技术研究所)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出DILLO模型,通过蒸馏方法实现无需视觉模拟的前瞻性智能体导航,提升任务成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17420 2026-03-19 cs.AI 94%

From Digital Twins to World Models:Opportunities, Challenges, and Applications for Mobile Edge General Intelligence

从数字孪生到世界模型:移动边缘通用智能的机会、挑战与应用

Jie Zheng, Dusit Niyato, Changyuan Zhao, Jiawen Kang, Jiacheng Wang

机构 * State-Province Joint Engineering and Research Center of Advanced Networking and Intelligent Information Services, College of Computer, Northwest University(高级网络与智能信息服务省州联合工程研究中心,计算机学院,西北大学) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) Automation of School, Guangdong University of Technology(自动化学院,广东工业大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文探讨了从数字孪生到世界模型的转变,分析了其在移动边缘通用智能中的作用,涵盖概念差异、设计原则、关键组件及应用挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01960 2026-03-17 cs.LG 94%

Grounding Generated Videos in Feasible Plans via World Models

通过世界模型将生成视频接地到可行计划中

Christos Ziakas, Amir Bar, Alessandra Russo

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出GVP-WM方法,通过学习动作条件的世界模型,将生成视频计划接地为可行动作序列,解决视频生成计划在时间一致性和物理约束上的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13587 2026-02-27 cs.CV 94%

UniFuture: A 4D Driving World Model for Future Generation and Perception

UniFuture: 一个面向未来生成与感知的4D驾驶世界模型

Dingkang Liang, Dingyuan Zhang, Xin Zhou, Sifan Tu, Tianrui Feng, Xiaofan Li, Yumeng Zhang, Mingyang Du, Xiao Tan, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Baidu Inc.(百度公司)

专题命中 通用世界模型 :world model(title,abstract);driving world model(title,abstract);world model(title,abstract);driving world model(title,abstract)

AI总结 UniFuture通过统一4D建模提升自动驾驶中的未来生成与几何感知能力。

Comments Accepted by ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22208 2026-02-27 cs.CV 94%

Solaris: Building a Multiplayer Video World Model in Minecraft

Solaris:在Minecraft中构建多玩家视频世界模型

Georgy Savva, Oscar Michel, Daohan Lu, Suppakit Waiwitlikhit, Timothy Meehan, Dhairya Mishra, Srivats Poddar, Jack Lu, Saining Xie

机构 * New York University(纽约大学)

专题命中 通用世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 Solaris通过多玩家数据系统和分阶段训练方法,在Minecraft中构建了能模拟多视角观测的视频世界模型,提升了多代理交互的建模能力。

Comments Project website: https://solaris-wm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03747 2026-02-04 cs.CV 94%

LIVE: Long-horizon Interactive Video World Modeling

LIVE: 长时距交互视频世界建模

Junchao Huang, Ziyang Ye, Xinting Hu, Tianyu He, Guiyu Zhang, Shaoshuai Shi, Jiang Bian, Li Jiang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Loop Area Institute(深圳河套学院) Microsoft Research(微软研究院) The University of Hong Kong(香港大学) Voyager Research, Didi Chuxing Project(Voyager研究,滴滴出行项目)

专题命中 通用世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 LIVE通过循环一致性目标限制误差累积,无需教师蒸馏,实现长时距交互视频世界建模,取得最优性能。

Comments 18 pages, 22 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01644 2026-02-03 cs.LG cs.AI cs.CV cs.MA cs.RO 94%

From Perception to Action: Spatial AI Agents and World Models

从感知到行动:空间AI代理与世界模型

Gloria Felicia, Nolan Bryant, Handi Putra, Ayaan Gazali, Eliel Lobo, Esteban Rojas

机构 * AtlasPro AI

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出了一种统一的三轴分类法,将代理能力与空间任务联系起来,强调空间定位与符号定位的区别,并指出世界模型对跨尺度安全部署的重要性。

Comments 61 pages, 742 citations, 1 figure, 3 tables. Survey paper on spatial AI agents, embodied AI, graph neural networks, and world models

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19834 2026-01-28 cs.AI 94%

Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models

视觉生成通过多模态世界模型解锁类人推理

Jialong Wu, Xiaoying Zhang, Hongyi Yuan, Xiangcheng Zhang, Tianhao Huang, Changjing He, Chaoyi Deng, Renrui Zhang, Youbin Wu, Mingsheng Long

机构 * Tsinghua University(清华大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 本文提出视觉生成在特定任务中优于纯语言推理,通过构建VisWorld-Eval评估套件验证了多模态世界模型提升类人推理的能力。

Comments Project page: https://thuml.github.io/Reasoning-Visual-World

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22336 2025-12-30 cs.AI cs.CL 94%

Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback

Agent2World: 通过自适应多智能体反馈学习生成符号世界模型

Mengkang Hu, Bowei Xia, Yuran Wu, Ailing Yu, Yude Zou, Qiguang Chen, Shijian Wang, Jiarui Jin, Kexin Li, Wenxiang Jiao, Yuan Lu, Ping Luo

机构 * The University of Hong Kong(香港大学) Xiaohongshu Inc.(小红书公司) UESTC(电子科技大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

AI总结 Agent2World通过自适应多智能体反馈学习生成符号世界模型,提升推理和微调性能,实现30.95%的相对提升。

Comments 48 pages, 15 tables, 7 figures, Project page: https://agent2world.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13604 2025-12-16 cs.CV 94%

LongVie 2: Multimodal Controllable Ultra-Long Video World Model

LongVie 2: 多模态可控超长视频世界模型

Jianxiong Gao, Zhaoxi Chen, Xian Liu, Junhao Zhuang, Chengming Xu, Jianfeng Feng, Yu Qiao, Yanwei Fu, Chenyang Si, Ziwei Liu

机构 * FDU(福建大学) NJU(南京大学) NTU(National University of Taiwan) NVIDIA(英伟达) THU(清华大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 通用世界模型 :world model(title,abstract);video world model(title,abstract);world model(title,abstract);video world model(title,abstract)

AI总结 LongVie 2通过多模态引导、退化感知训练和历史-上下文引导,实现了超长视频的可控生成,支持连续视频生成长达五分钟,推动视频世界建模的统一发展。

Comments Project Page: https://vchitect.github.io/LongVie2-project/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03520 2025-10-29 cs.CV 94%

Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond

Zheng Zhu, Xiaofeng Wang, Wangbo Zhao, Chen Min, Bohan Li, Nianchen Deng, Min Dou, Yuqi Wang, Botian Shi, Kai Wang, Chi Zhang, Yang You, Zhaoxiang Zhang, Dawei Zhao, Liang Xiao, Jian Zhao, Jiwen Lu, Guan Huang

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

Comments This survey will be regularly updated at: https://github.com/GigaAI-research/General-World-Models-Survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17600 2025-09-18 cs.RO cs.AI cs.CV cs.LG 94%

GWM: Towards Scalable Gaussian World Models for Robotic Manipulation

Guanxing Lu, Baoxiong Jia, Puhao Li, Yixin Chen, Ziwei Wang, Yansong Tang, Siyuan Huang

机构 * Tsinghua University(清华大学) State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电子与电气工程学院)

专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)

Comments Published at ICCV 2025. Project page: https://gaussian-world-model.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏