arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2026-06-29 至 2026-06-29 共收录 7 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身推理 7 篇

2510.16732 2026-06-29 cs.CV 版本更新 89%

A Comprehensive Survey on World Models for Embodied AI

具身AI世界模型综述

Xinqing Li, Xin He, Le Zhang, Min Wu, Xiaoli Li, Yun Liu

机构 * College of Computer Science and the Academy for Advanced Interdisciplinary Studies, Nankai University(南开大学计算机科学学院与前沿交叉学科研究院) School of Computer Science and Engineering, Tianjin University of Technology(天津理工大学计算机科学与工程学院) School of Information and Communication Engineering, University of Electronic Science and Technology of China(电子科技大学信息与通信工程学院) Institute for Infocomm Research (I2R), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局资讯通信研究院) Information Systems Technology and Design (ISTD) Pillar, Singapore University of Technology and Design (SUTD)(新加坡科技设计大学信息系统科技与设计系)

专题命中 具身推理 :embodied AI(title,abstract);world model(title,abstract);robotics(abstract);分类 cs.CV

AI总结 本文系统综述了具身AI中的世界模型,提出了功能、时间建模和空间表示的三轴分类法,并总结了数据资源、评估指标及开放挑战。

Comments https://github.com/Li-Zn-H/AwesomeWorldModels

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27575 2026-06-29 cs.CV 新提交 85%

Perceptual 3D Simulation With Physical World Modeling

基于物理世界建模的感知3D仿真

Wanhee Lee, Klemen Kotar, Rahul Mysore Venkatesh, Jared Watrous, Daniel L. K. Yamins

机构 * Stanford University(斯坦福大学)

专题命中 具身推理 :world model(title,abstract);robotics(abstract);manipulation(abstract);分类 cs.CV

AI总结 提出P3Sim系统,结合学习型物理世界模型、几何条件模块和持久场景记忆,在部分观测和不完整3D变换信号下模拟未来场景状态,实现多任务泛化。

Comments Published as a conference paper at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05092 2026-06-29 cs.RO cs.AI cs.CV 版本更新 82%

Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout

Driver-WM:以驾驶员为中心的交通条件潜在世界模型用于车内动态展开

Haozhuang Chi, Daosheng Qiu, Hao Su, Haochen Liu, Zirui Li, Haoruo Zhang, Chen Lv

机构 * Nanyang Technological University, Singapore(南洋理工大学,新加坡) Hubei University, Wuhan, China(湖北大学,武汉,中国) Osaka University, Osaka, Japan(大阪大学,大阪,日本)

专题命中 具身推理 :world model(title,abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 提出Driver-WM,一种以驾驶员为中心的潜在世界模型,通过门控因果注入机制在紧凑潜在空间中联合预测驾驶员物理运动、行为和情感,实现外部交通到内部动态的因果展开。

Comments Accepted to the 19th European Conference on Computer Vision (ECCV 2026). This version includes the supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28127 2026-06-29 cs.CL cs.AI cs.LG 新提交 81%

From Tokens to States: LLMs as a Special Case of World Models and the Continuous Path Beyond

从令牌到状态:LLM作为世界模型的特例及其连续路径

Paul Dubois

机构 * Paul Dubois(保罗·杜博伊斯)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文论证LLM是世界模型的退化特例,并提出从NTP到JEPA的连续谱系,逐步放松LLM约束,同时探讨其可扩展性挑战。

Comments 10 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27681 2026-06-29 cs.LG cs.CL 新提交 79%

Textual Belief States for World Models: Identifiable Representation Learning Under Strict Mediation

用于世界模型的文本信念状态:严格中介下的可识别表示学习

Xiang Gao, Kaiwen Dong, Yuguang Yao, Padmaja Jonnalagedda, Kamalika Das

机构 * Intuit AI Research(Intuit AI研究)

专题命中 具身推理 :world model(title,abstract);分类 cs.LG

AI总结 提出严格中介原则解决LLM中历史旁路导致的潜在状态不可识别问题,引入离散文本潜在状态和因子化GRPO方法,在TextWorld和ScienceWorld上实现表示质量和滚动性能显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27644 2026-06-29 cs.CV 新提交 79%

CascadeOcc: Rethinking 3D Occupancy World Models with Cascaded VQ Representations

CascadeOcc: 用级联VQ表示重新思考3D占用世界模型

Kyumin Hwang, Wonhyeok Choi, Jaeyeul Kim, Jihun Park, Daehee Park, Sunghoon Im

机构 * Daegu Gyeongbuk Institute of Science and Technology (DGIST)(大邱庆北科学技术院)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 提出CascadeOcc,一种通过级联向量量化机制在自回归框架中利用占用表示内在结构层次,实现从粗到细的3D场景预测和运动规划,在4D占用预测和运动规划基准上取得优越性能。

Comments Accepted to IEEE Signal Processing Letters (SPL), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26964 2026-06-29 cs.AI cs.CV 新提交 73%

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds

Look-Before-Move:动态3D故事世界中的叙事驱动世界视觉注意力

Jiaming Bian, Bingliang Li, Yuehao Wu, Pichao Wang, Zhi Wang, Hailan Ma, Huadong Mo, Zhenhong Sun

专题命中 具身推理 :embodied AI(abstract);world model(abstract);分类 cs.AI、cs.CV

AI总结 提出Look-Before-Move框架,通过语义观察合约、蒙特卡洛视点搜索和语义轨迹接地,在动态3D故事世界中实现叙事驱动的视觉注意力规划,提升主体感知、意图一致性和轨迹质量。

Comments 25 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏