arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 1471 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身导航 307 篇

2511.08292 2026-07-07 q-bio.NC 版本更新 50%

Distance by de-correlation: Computing distance with heterogeneous grid cells

通过去相关计算距离:利用异构网格细胞计算距离

Pritipriya Dasbehera, Akshunna S. Dogra, William T. Redman

专题命中 具身导航 :navigation(abstract)

AI总结 研究基于网格细胞特性的空间距离编码,引入群体活动去相关理论,通过数学理论和模拟揭示距离编码规律,为理解网格细胞编码距离及特性差异提供新见解。

Comments 29 pages, 11 figures, comments welcome!

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01044 2026-07-07 eess.SY cs.SY 版本更新 50%

Nonlinear receding-horizon differential game for drone racing along a three-dimensional path

用于三维路径无人机竞速的非线性滚动时域微分博弈

Kijin Sung, Kenta Hoshino, Akihiko Honda, Takeya Shima, Toshiyuki Ohtsuka

专题命中 具身导航 :navigation(abstract)

AI总结 针对无人机竞速控制挑战,研究提出非线性滚动时域微分博弈框架(NRHDG),通过预测对手行为改进标准非线性模型预测控制,给出新路径跟踪公式、势函数及性能指标,仿真显示其在超越和阻挡性能上优于NMPC。

Comments 20 pages, 14 figures; simulations and figures added to show the superiority of the proposed method; license changed to CC BY 4.0

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11435 2026-07-01 cs.SE 版本更新 50%

A Grounded Theory of Debugging in Professional Software Engineering Practice

面向专业软件工程实践的调试 grounded 理论

Haolin Li, Michael Coblenz

专题命中 具身导航 :navigation(abstract)

AI总结 本文通过 grounded 理论方法研究专业开发者在真实代码库中调试的策略与过程,揭示调试作为结构化迭代诊断过程的本质,为工具设计和软件工程教育提供启示。

Comments Accepted by FSE'26

Journal ref J. ACM 3, FSE, Article 139.382 (February 2026), 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13840 2026-07-01 physics.bio-ph quant-ph 版本更新 50%

Quantum modeling of radical pair magnetic sensor based on electric dipole moment

基于电偶极矩的自由基对磁传感器的量子建模

Mahboobe Sehati, Ali Soltanmanesh, Shabnam Abutalebi, Abolfazl Bahrampour, Naser Haeri, Sareh Rostami, Alireza Bahrampour

专题命中 具身导航 :navigation(abstract)

AI总结 通过量子力学建模研究隐花色素蛋白中自由基对的电子自旋动力学,发现电偶极矩对外磁场敏感,表明其可作为磁生物传感器。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26492 2026-06-23 math.DG math-ph math.MP 版本更新 50%

Generalized Fermat's principle and Snell's law for cone structures and applications

锥结构的广义费马原理与斯涅尔定律及其应用

Miguel Ángel Javaloyes, Steen Markvorsen, Enrique Pendás-Recondo, Miguel Sánchez

专题命中 具身导航 :navigation(abstract)

AI总结 将费马原理推广至光滑界面分隔两个锥结构(洛伦兹-芬斯勒光锥)的情形,推导出广义斯涅尔折射定律和反射定律,并应用于Zermelo导航问题和离散时空测地线确定。

Comments Slight reformulation of the main results in Sections 4 and 5 to increase their generality. 49 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06670 2026-06-19 eess.SY cs.SY 版本更新 50%

A Geometric Analysis-Based Safety Assessment Framework for Marine Vehicle Route Decision-Making

基于几何分析的船舶航线决策安全评估框架

Zilong Xu, Zihao Wang, He Li, Dingli Yu, Zaili Yang, Jin Wang

专题命中 具身导航 :navigation(abstract)

AI总结 提出基于几何分析的航线安全评估框架(GARSA),利用线和点几何元素定义航道边界,构建动态宽度表征函数量化受限水域空间安全性,并建立考虑整体及局部约束的航行安全指数,为船舶选择最安全航道提供定量依据。

Comments Accepted for publication in Reliability Engineering & System Safety (RESS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18815 2026-06-18 math.DS math.CA math.DG math.FA 版本更新 50%

Bifurcations for Lagrangian systems and geodesics II

拉格朗日系统与测地线的分岔 II

Guangcun Lu

专题命中 具身导航 :navigation(abstract)

AI总结 研究自治拉格朗日系统及Finsler/Riemann流形上测地线分岔,利用Morse指标和零化度技术给出广义周期解分岔的充要条件,并精化经典Gauss引理。

Comments 63 pages, LaTeX; matches published version. The article arXiv:2404.18815v2 [math.DS] has been split into two or more articles. This is one of this split. Another part of this split has already appeared as arXiv:2603.20551

Journal ref Calc. Var. Partial Differ. Equ., 65(2026), no.7, Art. no. 206

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19499 2026-06-16 eess.SP 版本更新 50%

Geometric Performance Analysis of Doppler-Based Positioning with a Single LEO Satellite

基于单颗低轨卫星的多普勒定位几何性能分析

Qi Liu, Marc Fernandez-Temprado, Antoni Reus-Bergas, Gonzalo Seco-Granados, Jose A. Lopez-Salcedo

专题命中 具身导航 :navigation(abstract)

AI总结 本文通过理论分析和数值模拟,研究了单颗低轨卫星多普勒定位的几何特性与精度因子,揭示了跨轨方向误差大、沿轨方向误差小的固有几何限制,为单星定位可行性提供了深入见解。

Comments 15 pages, 15 figures, submitted to IEEE Internet of Things Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07300 2026-06-16 eess.SY cs.SY 版本更新 50%

Distributed Omniscient Observers for Multi-Agent Systems: Design and Applications

多智能体系统的分布式全知观测器:设计与应用

Ganghui Cao, Xunyuan Yin

专题命中 具身导航 :navigation(abstract)

AI总结 针对异构和同构线性多智能体系统,提出基于局部输入输出信息的分布式全知观测器,使每个智能体正确估计所有智能体状态,无需全局通信图知识,应用于分布式纳什均衡求解和人工蜂群社会行为模拟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16377 2026-06-16 physics.geo-ph gr-qc physics.space-ph 版本更新 50%

Frequency Differences between Clocks on the Earth and the Moon

地球与月球时钟之间的频率差异

Mingyue Zhang, Jürgen Müller, Sergei M. Kopeikin

专题命中 具身导航 :navigation(abstract)

AI总结 基于广义相对论,建立了地球-月球时钟频率比较的完整模型,分析了重力势、坐标时间比及信号传播效应对频率差异的影响,并指出需要多链路策略抑制多普勒效应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16413 2026-06-15 eess.SY cs.SY 版本更新 50%

Hierarchical Distributed Architecture for the Least Allan Variance Atomic Timing

最小Allan方差原子定时的分层分布式架构

Jiayu Chen, Takahiro Kawaguchi, Yuichiro Yano, Yuko Hanado, Takayuki Ishizaki

专题命中 具身导航 :navigation(abstract)

AI总结 提出一种基于微型原子钟组的分层分布式定时架构,通过下层分布式控制实现短期同步,上层监督器在正常/应急模式下分别锚定标准时间或最优浮动控制,使生成时标的Allan方差在1秒至数天内达10^{-23}量级。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04397 2026-06-10 cs.SE cs.IR 版本更新 50%

Context-as-AI-Service: Surfacing Cross-File Dependency Chains for LLM-Generated Developer Documentation

上下文即服务:为LLM生成的开发者文档揭示跨文件依赖链

Ameya Gawde, Vyzantinos Repantis, Harshvardhan Singh, Lucy Moys

专题命中 具身导航 :navigation(abstract)

AI总结 提出Context-as-a-Service (CaaS)检索层,通过索引代码库并支持关键词与语义搜索,帮助LLM代理在生成或审查文档时追踪跨文件依赖链,实验表明CaaS能发现基线遗漏的错误并减少时间和输入令牌消耗。

Comments 8 pages, 2 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07529 2026-06-10 eess.SY cs.SY math.OC 版本更新 50%

Stochastic Differential Dynamic Programming for Trajectory Optimization under Partial Observability

部分可观测下轨迹优化的随机微分动态规划

Masahiro Fujiwara, Naoya Ozaki

专题命中 具身导航 :navigation(abstract)

AI总结 提出一种随机微分动态规划算法,用于在部分可观测条件下联合优化标称控制序列和反馈增益,显式考虑协方差传播对标称轨迹的依赖,生成导航感知且鲁棒的轨迹。

Comments Revised version; 38 pages, 13 figures; submitted to the Journal of Guidance, Control, and Dynamics

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04217 2026-06-10 eess.SP 版本更新 50%

Towards 6G Single-Anchor Vehicle Localization Exploiting Radio-Reflective Road Markings in Tunnel Environments

面向隧道环境中利用无线电反射路面标记的6G单锚车辆定位

Lorenzo Italiano, Mattia Brambilla, Monica Nicoli

专题命中 具身导航 :navigation(abstract)

AI总结 针对隧道等无GNSS区域,提出一种利用近场传播和被动反射结构的单锚车辆定位方法JAVELIN,通过张量参数估计和递归贝叶斯跟踪实现亚米级定位。

Comments Currently submitted to IET-ITS

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02383 2026-06-09 cs.MA cs.GT cs.SY eess.SY 版本更新 50%

A Game-Theoretic Decision Framework for Optimal Selection of Coordination Detection Methods in Multi-UAV Fleet Operations

多无人机编队操作中协调检测方法最优选择的博弈论决策框架

Christian Manasseh, Savana Ammons

专题命中 具身导航 :navigation(abstract)

AI总结 提出一个博弈论决策框架,通过将方法选择建模为监控者与自然之间的零和博弈,解决多无人机编队协调检测中速度与准确性的权衡问题,并利用混合策略保证最坏情况下的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 具身推理 148 篇

2607.00836 2026-08-13 cs.RO cs.AI cs.SY eess.SY 版本更新 90%

From World Models to World Action Models: A Concise Tutorial for Robotics

从世界模型到世界动作模型:面向机器人学的简明教程

Xiaoxiong Zhang, Xiong Zeng, Wei Zhang

专题命中 具身推理 :robotics(title,abstract);world model(title,abstract);robotic(abstract);分类 cs.RO、cs.AI

AI总结 本教程提出世界模型作为动作条件预测模型的设计空间视图,分类为观测空间和状态空间模型,并引入世界动作模型,总结四种代表性范式,为具身预测与控制提供结构化分类。

Comments Project page: https://clearlab-sustech.github.io/WorldModelSurvey/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21918 2026-08-04 cs.RO 版本更新 89%

Action-Conditioned World Model for Goal Plane Probe Guidance in Robotic Ultrasound

用于机器人超声中目标平面探头引导的动作条件世界模型

Siqi Fan, Mingcong Chen, Ran Liu, Zixuan Yang, Xiaoyu Fu, Xiaoqing Gao, Yunhui Liu, Hongbin Liu

机构 * Department of Mechanical and Automation Engineering, Chinese University of Hong Kong(香港中文大学机械与自动化工程学系) Department of Biomedical Engineering, City University of Hong Kong(香港城市大学生物医学工程学系) Centre for Artificial Intelligence and Robotics Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences(中国科学院香港创新研究院人工智能与机器人中心) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Department of Surgical and Interventional Engineering, King’s College London(伦敦国王大学外科与介入工程学系) Department of Vascular Ultrasound, Xuanwu Hospital, the Capital Medical University(首都医科大学宣武医院血管超声科)

专题命中 具身推理 :world model(title,abstract);robotic(title,abstract);navigation(abstract);分类 cs.RO

AI总结 研究针对机器人超声目标平面探头引导问题,提出基于模型的两阶段学习管道,包括潜在条件扩散世界模型和目标条件时间变换器,在自收集数据集及实际闭环实验中取得较好成果,证明学习到的超声动力学对训练目标导向机器人探头导航有潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16732 2026-06-29 cs.CV 版本更新 89%

A Comprehensive Survey on World Models for Embodied AI

具身AI世界模型综述

Xinqing Li, Xin He, Le Zhang, Min Wu, Xiaoli Li, Yun Liu

机构 * College of Computer Science and the Academy for Advanced Interdisciplinary Studies, Nankai University(南开大学计算机科学学院与前沿交叉学科研究院) School of Computer Science and Engineering, Tianjin University of Technology(天津理工大学计算机科学与工程学院) School of Information and Communication Engineering, University of Electronic Science and Technology of China(电子科技大学信息与通信工程学院) Institute for Infocomm Research (I2R), Agency for Science, Technology and Research (A*STAR)(新加坡科技研究局资讯通信研究院) Information Systems Technology and Design (ISTD) Pillar, Singapore University of Technology and Design (SUTD)(新加坡科技设计大学信息系统科技与设计系)

专题命中 具身推理 :embodied AI(title,abstract);world model(title,abstract);robotics(abstract);分类 cs.CV

AI总结 本文系统综述了具身AI中的世界模型,提出了功能、时间建模和空间表示的三轴分类法,并总结了数据资源、评估指标及开放挑战。

Comments https://github.com/Li-Zn-H/AwesomeWorldModels

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03609 2026-06-17 cs.RO cs.LG 版本更新 89%

A 3D Isovist World Model -- Revealing a City's Unseen Geometry and Its Emergent Cross-City Signature

3D 等视域世界模型——揭示城市不可见几何及其涌现的跨城市特征

Xuhui Lin, Stephen Law, Nanjiang Chen, Kunyao Li, Tao Yang

机构 * The Bartlett School of Sustainable Construction University College London, UK(可持续建设学院伦敦大学学院,英国) Department of Geography University College London, UK(地理系伦敦大学学院,英国) School of Project Management, Faculty of Engineering The University of Sydney, AU(工程学院项目管理学院悉尼大学,澳大利亚) School of Engineering Cardiff University, UK(工程学院卡迪夫大学,英国) School of Architecture Tsinghua University, Beijing, CN(建筑学院清华大学,北京,中国)

专题命中 具身推理 :world model(title,abstract);robotics(abstract);embodied AI(abstract);embodied agent(abstract)

AI总结 提出一种预测3D等视域(球形可见性深度图)的具身世界模型,通过深度残差和自滚动调度采样训练,发现跨城市空间特征可从时间潜变量中线性解码。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15284 2026-07-30 cs.CV 版本更新 89%

Walk through Paintings: Egocentric World Models from Internet Priors

穿越绘画:来自互联网先验的自我中心世界模型

Anurag Bagchi, Zhipeng Bao, Homanga Bharadhwaj, Yu-Xiong Wang, Pavel Tokmakov, Martial Hebert

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Toyota Research Institute(丰田研究机构)

专题命中 具身推理 :world model(title,abstract);manipulation(abstract,abstract_cn);navigation(abstract);robotic(abstract)

AI总结 EgoWM通过利用互联网先验和轻量级条件层,将预训练视频扩散模型转换为自我中心世界模型,实现可控的未来预测,提升结构一致性评分并降低推理延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19190 2026-07-27 cs.RO cs.AI 版本更新 88%

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

智能体真实到模拟:使用视觉语言智能体进行基于物理的世界建模

Guanxiong Chen, Qianjun Xia, Jiawei Peng, Heng Zhang, Bole Ma, Justin Qian, Ziyi Jiao, Bingyang Zhou, Luoxin Ye, Kaifeng Zhang, Kunyi Wang, Weijia Zeng, Yunuo Chen, Pengzhi Yang, Ziqiu Zeng, Siyuan Luo, Huamin Wang, Chao Liu, Alan Yuille, Fan Shi, Changxi Zheng, Yunzhu Li, Chenfanfu Jiang, Peter Yichen Chen

机构 * University of British Columbia(英属哥伦比亚大学) Johns Hopkins University(约翰·霍普金斯大学) National University of Singapore(新加坡国立大学) Columbia University(哥伦比亚大学) University of California, Los Angeles(加利福尼亚大学洛杉矶分校) Style3D(无合适对应中文名)

专题命中 具身推理 :world model(title,abstract);robotics(abstract);manipulation(abstract);robotic(abstract)

AI总结 研究针对机器人与物体交互的真实到模拟转换难题,提出智能体真实到模拟框架,利用视觉语言智能体进行广义物理世界建模,能转换真实记录为可模拟孪生体,在多场景评估效果良好,成本低,可用于下游机器人任务。

Comments Authorship change

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01799 2026-06-09 cs.CV 版本更新 87%

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling

Embody4D: 面向具身4D世界建模的通用数据引擎

Peiyan Tu, Hanxin Zhu, Jingwen Sun, Shaojie Ren, Cong Wang, Yuyan Xu, Jiayi Luo, Xiaoqian Cheng, Zhibo Chen

机构 * Zhejiang University(浙江大学) Beijing Zhongguancun Academy(北京中关村学院) University of Science and Technology of China(中国科学技术大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Shanghai Jiao Tong University(上海交通大学) Beihang University(北京航空航天大学)

专题命中 具身推理 :world model(title,abstract);embodied agent(abstract);manipulation(abstract);robotic(abstract)

AI总结 提出Embody4D视频到视频世界模型,通过3D感知合成管道、潜在置信度专家调制和交互注意力机制,将单目机器人视频转换为多视角视频,解决具身智能中视角稀疏问题,提升下游规划与学习性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00412 2026-07-03 cs.AI cs.RO 版本更新 86%

Physically Native World Models: A Hamiltonian Perspective on Generative World Modeling

物理原生世界模型:生成式世界建模的哈密顿视角

Sen Cui, Jingheng Ma

机构 * Tsinghua University(清华大学)

专题命中 具身推理 :world model(title,abstract);robotics(abstract);robotic(abstract);分类 cs.RO、cs.AI

AI总结 提出哈密顿世界模型,通过结构化潜相空间和哈密顿动力学演化实现物理可靠、动作可控且长期稳定的未来预测,用于具身决策。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31158 2026-06-19 cs.CV cs.LG 版本更新 86%

Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models

光交互:交互式视频世界模型的免训练推理加速

Jiacheng Lu, Haoyi Zhu, Sipei Yi, Enze Xie, Yu Li, Cheng Zhuo

机构 * Zhejiang University(浙江大学) NVIDIA

专题命中 具身推理 :world model(title,abstract);embodied AI(abstract);navigation(abstract);分类 cs.CV、cs.LG

AI总结 针对交互式视频世界模型推理成本高的问题,提出免训练加速框架Light Interaction,通过自适应上下文管理、去噪缓存加速和3D块稀疏注意力实现最高2.59倍加速。

Comments 13 pages, 6 figures, 3 tables. Project page: https://2843721358l-del.github.io/Light-Interaction-Project/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08732 2026-06-08 cs.RO cs.LG 版本更新 86%

Latent Geometry Beyond Search: Amortizing Planning in World Models

超越搜索的潜在几何:在世界模型中摊销规划

Hoang Nguyen, Xiaohao Xu, Xiaonan Huang

机构 * Department of Robotics, University of Michigan, Ann Arbor(密歇根大学机器人系,安阿伯)

专题命中 具身推理 :world model(title,abstract);manipulation(abstract);navigation(abstract);分类 cs.RO、cs.LG

AI总结 提出在正则化潜在几何下,将规划摊销为潜在逆动力学映射,以轻量级GC-IDM替代在线搜索,在七个环境协议中匹配或超越CEM,决策成本降低100-130倍。

Comments 31 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12920 2026-07-13 cs.MA cs.AI cs.CL 版本更新 85%

Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue

通过对话对齐世界模型实现具身多智能体协调

Vardhan Dongre, Dilek Hakkani-Tür

机构 * Siebel School of Computing & Data Science(计算机与数据科学学院)

专题命中 具身推理 :world model(title,abstract);robotics(abstract);embodied agent(abstract);分类 cs.AI

AI总结 研究通过对话机制探索具身智能体的世界模型对齐,发现对话能减少冲突但降低任务成功率,提出评估世界模型对齐的框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02800 2026-06-24 cs.CV cs.AI cs.LG cs.MM cs.RO 版本更新 85%

Cosmos 3: Omnimodal World Models for Physical AI

Cosmos 3:面向物理AI的全模态世界模型

NVIDIA, :, Aditi, Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, Aarti Basant, Mukesh Beladiya, Mohammad Qazim Bhat, Zaid Pervaiz Bhat, Dan Blick, Vanni Brighella, Han Cai, Tiffany Cai, Eric Cameracci, Jiaxin Cao, Yulong Cao, Mark Carlson, Carlos Casanova, Ting-Yun Chang, Yan Chang, Yu-Wei Chao, Prithvijit Chattopadhyay, Roshan Chaudhari, Chieh-Yun Chen, Junyu Chen, Ke Chen, Qizhi Chen, Wenkai Chen, Xiaotong Chen, Yu Chen, An-Chieh Cheng, Click Cheng, Xiu Chia, Jeana Choi, Chaeyeon Chung, Wenyan Cong, Yin Cui, Magdalena Dadela, Nalin Dadhich, Wenliang Dai, Joyjit Daw, Alperen Degirmenci, Rodrigo Vieira Del Monte, Robert Denomme, Sameer Dharur, Marco Di Lucca, Ke Ding, Wenhao Ding, Yifan Ding, Yuzhu Dong, Nicole Drumheller, Yilun Du, Aigul Dzhumamuratova, Aleksandr Efitorov, Hamid Eghbalzadeh, Naomi Eigbe, Imad El Hanafi, Hassan Eslami, Benedikt Falk, Jiaojiao Fan, Jim Fan, Amol Fasale, Sergiy Fefilatyev, Liang Feng, Francesco Ferroni, Sanja Fidler, Xiao Fu, Vikram Fugro, Prashant Gaikwad, TJ Galda, Katelyn Gao, Yihuai Gao, Wenhang Ge, Sreyan Ghosh, Arushi Goel, Vivek Goel, Akash Gokul, Rama Govindaraju, Jinwei Gu, Miguel Guerrero, Elfie Guo, Aryaman Gupta, Siddharth Gururani, Hugo Hadfield, Song Han, Ankur Handa, Zekun Hao, Mohammad Harrim, Ali Hassani, Nathan Hayes-Roth, Yufan He, Chris Helvig, Cyrus Hogg, Madison Huang, Michael Huang, Sophia Huang, Yufan Huang, Jacob Huffman, DeLesley Hutchins, Suneel Indupuru, Boris Ivanovic, Arihant Jain, Joel Jang, Ryan Ji, Yanan Jian, Dongfu Jiang, Jingyi Jin, Atharva Joshi, Nikhilesh Joshi, Pranjali Joshi, Andy Ju, Jaehun Jung, Weiwei Kang, Scott Kassekert, Jan Kautz, Ashna Khetan, Julia Kiczka, Slawek Kierat, Gwanghyun Kim, Kuno Kim, Sunny Kim, Kezhi Kong, Xin Kong, Zhifeng Kong, Tomasz Kornuta, Egor Krivov, Hui Kuang, Saurav Kumar, Chia-Wen Kuo, George Kurian, Wojciech Kutak, JF Lafleche, Himangshu Lahkar, Omar Laymoun, Jayjun Lee, Sanggil Lee, Gabriele Leone, Boyi Li, Freya Li, Jiajun Li, Jinfeng Li, Ling Li, Pengcheng Li, Shangru Li, Tingle Li, Xiaolong Li, Xuan Li, Zhaoshuo Li, Zhiqi Li, Hao Liang, Maosheng Liao, Chen-Hsuan Lin, Tsung-Yi Lin, Ming-Yu Liu, Sifei Liu, Zihan Liu, Hai Loc Lu, Xiangyu Lu, Alice Luo, Ruipu Luo, Wenjie Luo, Jiangran Lyu, Martin Ding Ma, Nic Ma, Qianli Ma, Dawid Majchrowski, Louis Marcoux, Miguel Martin, Qing Miao, Ashkan Mirzaei, Shreyas Misra, Kaichun Mo, Durra Mohsin, Hyejin Moon, Pawel Morkisz, Saeid Motiian, Kirill Motkov, Seungjun Nah, Yashraj Narang, Deepak Narayanan, Thabang Ngazimbi, Julian Ouyang, Shubham Pachori, David Page, Yatian Pang, Sehwi Park, Mahesh Patekar, Mostofa Patwary, Marco Pavone, Trung Pham, Wei Ping, Soha Pouya, Shrimai Prabhumoye, Varun Praveen, Delin Qu, Hesam Rabeti, Morteza Ramezanali, Marilyn Reeb, Xuanchi Ren, Kristen Rumley, Wojciech Rymer, Jun Saito, Yeongho Seol, John Shao, Piyush Shekdar, Tianwei Shen, Humphrey Shi, Min Shi, Stella Shi, Kevin Shih, Mohammad Shoeybi, Mateusz Sieniawski, Shuran Song, Alexander Sotelo, Amir Sotoodeh, Sunil Srinivasa, Vignesh Srinivasakumar, Bartosz Stefaniak, Rahul Heinrich Steiger, Shangkun Sun, Jiaxiang Tang, Shitao Tang, Yangyang Tang, Yue Tang, Tolou Tavakkoli, Kayley Ting, Krzysztof Tomala, Wei-Cheng Tseng, Jibin Varghese, Sergei Vasilev, Thomas Volk, Raju Wagwani, Roger Waleffe, Andrew Z. Wang, Boxiang Wang, Haoxiang Wang, Qiao Wang, Shihao Wang, Shijie Wang, Ting-Chun Wang, Yan Wang, Yu Wang, Rohit Watve, David Wehr, Fangyin Wei, Xinshuo Weng, Jay Zhangjie Wu, Kedi Wu, Hongchi Xia, Summer Xiao, Tianjun Xiao, Kevin Xie, Daguang Xu, Jiashu Xu, Mengyao Xu, Ruqing Xu, Xingqian Xu, Yao Xu, Dinghao Yang, Dong Yang, Hans Yang, Xiaodong Yang, Xuning Yang, Yichu Yang, Yurong You, Zhiding Yu, Hao Yuan, Simon Yuen, Xiaohui Zeng, Pengcuo Zeren, Cindy Zha, Haotian Zhang, Jenny Zhang, Jing Zhang, Liangkai Zhang, Paris Zhang, Shun Zhang, Xuanmeng Zhang, Zhizheng Zhang, Ann Zhao, Yilin Zhao, Yuliya Zhautouskaya, Charles Zhou, Fengzhe Zhou, Shilin Zhu, Yuke Zhu, Dima Zhylko, Artur Zolkowski

机构 * NVIDIA

专题命中 具身推理 :world model(title,abstract);embodied agent(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 提出基于统一混合Transformer架构的全模态世界模型Cosmos 3,联合处理语言、图像、视频、音频和动作序列,在理解和生成任务上达到新最优,为具身智能体提供可扩展的通用骨干。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03208 2026-06-18 cs.LG 版本更新 85%

Hierarchical Planning with Latent World Models

基于潜在世界模型的分层规划

Wancong Zhang, Basile Terver, Artem Zholus, Soham Chitnis, Harsh Sutaria, Mido Assran, Randall Balestriero, Amir Bar, Adrien Bardes, Yann LeCun, Nicolas Ballas

机构 * FAIR at Meta(Meta旗下的FAIR) New York University(纽约大学) Mila - Québec AI Institute(魁北克AI研究院) Brown University(布朗大学)

专题命中 具身推理 :world model(title,abstract);manipulation(abstract);navigation(abstract);分类 cs.LG

AI总结 提出HWM架构,通过多时间尺度潜在世界模型和潜在匹配实现分层模型预测控制,解决长时域任务中单层规划失败和计算爆炸问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22281 2026-06-17 cs.CV cs.AI cs.CL cs.LG cs.RO 版本更新 85%

ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model

ThinkJEPA:赋予潜在世界模型大型视觉-语言推理能力

Haichao Zhang, Yijiang Li, Shwai He, Tushar Nagarajan, Mingfei Chen, Jianglin Lu, Ang Li, Yun Fu

机构 * Northeastern University(东北大学) University of California San Diego(加州大学圣地亚哥分校) University of Maryland(马里兰大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Washington(华盛顿大学)

专题命中 具身推理 :world model(title,abstract);manipulation(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 提出ThinkJEPA框架,结合密集JEPA分支与稀疏VLM思考者分支,通过分层金字塔表示提取模块,实现细粒度运动建模与长程语义引导,在手部操作轨迹预测任务上超越基线。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26182 2026-07-22 cs.CV cs.AI cs.LG 版本更新 85%

Lifting Embodied World Models for Planning and Control

提升具身世界模型用于规划与控制

Alex N. Wang, Trevor Darrell, Pavel Izmailov, Yutong Bai, Amir Bar

机构 * Computer Science, New York University(纽约大学计算机科学系) BAIR, UC Berkeley(伯克利大学BAIR)

专题命中 具身推理 :world model(title,abstract);embodied agent(abstract);分类 cs.AI、cs.CV、cs.LG

AI总结 本文提出一种轻量级策略,将高层动作映射到低层关节动作序列,结合冻结的世界模型,实现预测未来观察的提升世界模型,有效降低规划复杂度。

Comments Accepted to ECCV2026. Edited policy masking

详情

展开后加载摘要…

URL PDF HTML 收藏