arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 3100 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身推理 3100 篇

2606.03003 2026-07-03 cs.LG cs.AI cs.RO 版本更新 67%

Exact equivariance, kept through training, buys zero-shot generalisation across the symmetry group

精确等变性在训练中保持,实现跨对称群的零样本泛化

Hongbo Wang

机构 * Department of Mathematics, Stony Brook University(石溪大学数学系)

专题命中 具身推理 :world model(abstract);分类 cs.RO、cs.AI、cs.LG

AI总结 通过等变编码器和预测器构建的潜世界模型,其训练损失具有可证明的对称性,从而在仅拟合部分方向动力学时,数学上确定整个轨道上的行为,实现跨对称群的零样本泛化。

Comments 112 pages, 19 figures. v2 adds programme lineage to companion papers (arXiv:2606.13092, 2606.24945, 2606.24946), engages the equivariance-at-scale debate (arXiv:2410.23179), and adds experimental hardening: 5-seed CIs, frame-averaging/canonicalization baselines, a real-robot DROID anchor, a scale-vs-exactness curve. Core claims unchanged. Code: https://github.com/TimothyWang418/se3-ejepa

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02491 2026-06-30 cs.LG cs.AI cs.RO q-bio.NC stat.ML 67%

What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty

什么是有能力的智能体必须知道:在不确定性下鲁棒决策制定的选取定理

Aran Nayebi

机构 * Machine Learning Department and Neuroscience & Robotics Institutes, Carnegie Mellon University(机器学习系和神经科学与机器人研究院,卡内基梅隆大学)

专题命中 具身推理 :world model(abstract);分类 cs.RO、cs.AI、cs.LG

AI总结 本文探讨了在不确定性下智能体必需的内部结构,证明了强任务表现迫使世界模型、信念记忆及持久变量的必要性,提出了选取定理以解决部分可观测性和任务分布下的决策问题。

Comments 23 pages, 1 figure. To appear in Uncertainty in Artificial Intelligence (UAI) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26217 2026-06-26 cs.LG cs.CV cs.RO 新提交 67%

Fast LeWorldModel

快速LeWorldModel

Yuntian Gao, Xiangyu Xu

机构 * Xi’an Jiaotong University(西安交通大学)

专题命中 具身推理 :world model(abstract);分类 cs.RO、cs.CV、cs.LG

AI总结 提出Fast-LeWM,通过动作前缀并行预测未来潜在状态,替代LeWM的自回归滚动,降低规划时间并减少潜在误差累积。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23991 2026-06-24 cs.AI cs.LG cs.MA cs.RO 新提交 67%

Critique of Agent Model

智能体模型批判

Eric Xing, Mingkai Deng, Jinyu Hou

机构 * Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学基础模型研究所) School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院)

专题命中 具身推理 :world model(abstract);分类 cs.RO、cs.AI、cs.LG

AI总结 本文区分了自动化与智能体,提出真正智能体需内化目标、身份、决策、自我调节和学习等结构,并设计了GIC通用智能体架构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24631 2026-05-26 cs.LG cs.AI cs.CV 67%

Beyond Generative Priors: Minority Sampling with JEPA-Guided Diffusion

超越生成先验:JEPA引导扩散的少数采样

Sol Park, Soobin Um

机构 * Department of Artificial Intelligence, Kookmin University, Seoul, South Korea(人工智能系,韩国全州大学,首尔)

专题命中 具身推理 :world model(abstract);分类 cs.AI、cs.CV、cs.LG

AI总结 提出一种基于世界模型JEPA引导的扩散采样框架,通过近似策略实现高效计算,在无条件、类别条件和文本到图像生成中提升少数样本的保真度和语义有效性。

Comments ICML 2026, 21 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16848 2026-05-19 cs.CV cs.AI cs.CL cs.LG 67%

Thinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern Induction

基于模式的思考:通过模式诱导突破视觉规划中的感知瓶颈

Yichang Jian, Boyuan Xiao, Zhenyuan Huang, Yifei Peng, Yao-Xiang Ding

机构 * State Key Lab of CAD& CG(CAD与CG国家重点实验室)

专题命中 具身推理 :world model(abstract);分类 cs.AI、cs.CV、cs.LG

AI总结 本文提出通过模式诱导的方法,利用模式推理和模式诱导策略,使视觉语言模型在视觉规划任务中实现更高效和准确的感知与推理,解决传统模型在复杂输入下的感知瓶颈问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09693 2026-05-12 cs.CV cs.AI cs.LG 67%

Do multimodal models imagine electric sheep?

多模态模型是否想象出电羊?

Santhosh Kumar Ramakrishnan, Carl Vondrick, Raja Giryes, Philipp Krähenbühl, Vladlen Koltun

机构 * Apple(苹果公司)

专题命中 具身推理 :world model(abstract);分类 cs.AI、cs.CV、cs.LG

AI总结 研究发现多模态模型在解决空间谜题时会形成心理图像,通过微调Qwen3.5 VLM解决多种视觉推理任务,发现动作序列预测能提升解谜准确率,尤其在需要空间推理的任务中效果显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14902 2026-04-21 cs.AI cs.CL cs.CV cs.RO 67%

ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints

ADAPT:在未指定 affordance 约束下评估常识规划的基准测试

Pei-An Chen, Yong-Ching Liang, Jia-Fong Yeh, Hung-Ting Su, Yi-Ting Chen, Min Sun, Winston Hsu

机构 * National Taiwan University(国立台湾大学) National Yang Ming Chiao Tung University(国立阳明交通大学) National Tsing Hua University(国立清华大学)

专题命中 具身推理 :embodied agent(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 本文提出 ADAPT 模块,通过显式 affordance 推理提升智能具身代理的鲁棒性和任务成功率,展示了领域适应的视觉语言模型在 affordance 推理中的优越性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01765 2026-04-03 cs.CV cs.AI cs.RO 67%

DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning

DriveDreamer-Policy: 一种基于几何的世界-动作模型用于统一生成与规划

Yang Zhou, Xiaofeng Wang, Hao Shao, Letian Wang, Guosheng Zhao, Jiangnan Shao, Jiagang Zhu, Tingdong Yu, Zheng Zhu, Guan Huang, Steven L. Waslander

机构 * GigaAI University of Toronto(多伦多大学) CUHK MMLab(香港中文大学多媒体实验室)

专题命中 具身推理 :world model(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 本文提出DriveDreamer-Policy,结合深度生成、未来视频生成与运动规划,通过几何感知的世界表示提升生成与规划的连贯性与准确性。

Comments 11 pages, 4 figures; Project Website: https://drivedreamer-policy.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25741 2026-03-31 cs.CV cs.AI cs.RO 67%

Vega: Learning to Drive with Natural Language Instructions

Vega:通过自然语言指令学习驾驶

Sicheng Zuo, Yuxuan Li, Wenzhao Zheng, Zheng Zhu, Jie Zhou, Jiwen Lu

机构 * Tsinghua University(清华大学) GigaAI

专题命中 具身推理 :world model(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 本文提出Vega模型,通过自然语言指令生成和规划,提升自动驾驶的灵活性和个性化水平,基于大规模驾驶数据集InstructScene进行实验验证。

Comments Code is available at https://github.com/zuosc19/Vega

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24721 2026-03-27 cs.CV cs.AI cs.LG cs.MM 67%

Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models

可扩展的对象关系编码以提升大语言模型中的3D空间推理

Shengli Zhou, Minghang Zheng, Feng Zheng, Yang Liu

机构 * Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学计算机科学与工程系) Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)

专题命中 具身推理 :embodied agent(abstract);分类 cs.AI、cs.CV、cs.LG

AI总结 本文提出QuatRoPE,通过线性增长的对象数量输入长度和注意力层中的点积计算配对空间关系,提升大语言模型在3D空间推理中的性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17333 2026-03-19 cs.CL 67%

Grid Spatial Understanding: A Dataset for Textual Spatial Reasoning over Grids, Embodied Settings, and Coordinate Structures

网格空间理解:一个用于文本空间推理、具身场景和坐标结构的数据集

Risham Sidhu, Julia Hockenmaier

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 具身推理 :embodied agent(abstract);navigation(abstract)

AI总结 本文提出GSU数据集,用于评估LLM在导航、物体定位和结构组合任务中的空间推理能力,发现模型在具身代理参考框架和坐标列表识别上存在困难,且微调小模型可能达到前沿模型性能。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17438 2025-12-12 cs.AI cs.CV cs.LG cs.NE 67%

Object-centric proto-symbolic behavioural reasoning from pixels

基于像素的物体中心原型符号行为推理

Ruben van Bergen, Justus Hübotter, Alma Lago, Pablo Lanillos

机构 * Donders Institute, Radboud University(多纳尔斯研究所,拉布德大学) Cajal Neuroscience Center, Spanish National Research Council(卡哈尔神经科学中心,西班牙国家研究理事会)

专题命中 具身推理 :world model(abstract);分类 cs.AI、cs.CV、cs.LG

AI总结 本文提出了一种基于像素的物体中心深度学习架构,通过物体表示实现从感知到抽象推理的行为推理,展示了在合成环境中通过逻辑推理和连续控制任务的能力。

Comments Accepted for publication in Neural Networks journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00736 2025-12-02 cs.LG cs.AI cs.CV 67%

REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories

REM:通过多帧轨迹评估大语言模型的具身空间推理

Jacob Thompson, Emiliano Garcia-Lopez, Yonatan Bisk

机构 * Department of Computer Science(计算机科学系) Carnegie Mellon University(卡内基梅隆大学)

专题命中 具身推理 :navigation(abstract);分类 cs.AI、cs.CV、cs.LG

AI总结 REM通过多帧轨迹评估大语言模型的空间推理能力,揭示其在复杂空间任务中的不足,推动更 robust 的空间表示研究。

Journal ref Proceedings of the Conference on Language Modeling (COLM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00019 2025-12-02 cs.RO cs.AI cs.CV 67%

A Comprehensive Survey on Surgical Digital Twin

外科数字孪生的全面综述

Afsah Sharaf Khan, Falong Fan, Doohwan DH Kim, Abdurrahman Alshareef, Dong Chen, Justin Kim, Ernest Carter, Bo Liu, Jerzy W. Rozenblit, Bernard Zeigler

机构 * IEEE

专题命中 具身推理 :robotic(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 本文综述了外科数字孪生的技术现状与挑战,提出分类方法并识别了验证、安全性和数据治理等开放问题,旨在推动其在临床中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07463 2025-11-19 cs.RO cs.AI cs.CV 67%

DepthVision: Enabling Robust Vision-Language Models with GAN-Based LiDAR-to-RGB Synthesis for Autonomous Driving

Sven Kirchner, Nils Purschke, Ross Greer, Alois C. Knoll

机构 * Chair of Robotics, Artificial Intelligence and Real-time Systems, Technical University of Munich(机器人学、人工智能与实时系统教授会,慕尼黑技术大学) Computer Science and Engineering Department, University of California Merced(计算机科学与工程系,加州大学默塞德分校)

专题命中 具身推理 :robotics(abstract);分类 cs.RO、cs.AI、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03034 2025-09-04 cs.LG cs.AI cs.CR cs.CV cs.CY 67%

Rethinking Data Protection in the (Generative) Artificial Intelligence Era

Yiming Li, Shuo Shao, Yu He, Junfeng Guo, Tianwei Zhang, Zhan Qin, Pin-Yu Chen, Michael Backes, Philip Torr, Dacheng Tao, Kui Ren

机构 * The State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室) Nanyang Technological University(南洋理工大学) University of Maryland(马里兰大学) IBM Research(IBM研究院) CISPA Helmholtz Center for Information Security(CISPA 欧洲信息安全部分) University of Oxford(牛津大学)

专题命中 具身推理 :world model(abstract);分类 cs.AI、cs.CV、cs.LG

Comments Perspective paper for a broader scientific audience. The first two authors contributed equally to this paper. 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00465 2025-09-03 cs.RO cs.AI cs.CV 67%

Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning

Jiading Fang

机构 * Toyota Technological Institute at Chicago (TTIC)(丰田技术研究所(芝加哥))

专题命中 具身推理 :navigation(abstract);分类 cs.RO、cs.AI、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04380 2025-08-26 cs.CV cs.AI cs.LG 67%

EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding

Yuqi Wu, Wenzhao Zheng, Sicheng Zuo, Yuanhui Huang, Jie Zhou, Jiwen Lu

机构 * Department of Automation, Tsinghua University(自动化系,清华大学)

专题命中 具身推理 :embodied agent(abstract);分类 cs.AI、cs.CV、cs.LG

Comments Accepted by ICCV2025. Code: https://github.com/YkiWu/EmbodiedOcc

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07909 2025-08-12 cs.CV cs.AI cs.RO 67%

FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene Interaction

Dennis Rotondi, Fabio Scaparro, Hermann Blum, Kai O. Arras

机构 * Socially Intelligent Robotics Lab, Institute for Artificial Intelligence, University of Stuttgart(社会智能机器人实验室,人工智能研究所,斯图加特大学) Robot Perception and Learning Lab, LAMARR Institute for Machine Learning and Artificial Intelligence, University of Bonn(机器人感知与学习实验室,LAMARR机器学习与人工智能研究所,波恩大学)

专题命中 具身推理 :robotic(abstract);分类 cs.RO、cs.AI、cs.CV

Comments Paper accepted for IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01967 2025-07-04 q-bio.NC 67%

Ghost in the Machine: Examining the Philosophical Implications of Recursive Algorithms in Artificial Intelligence Systems

Llewellin RG Jegels

专题命中 具身推理 :robotics(abstract);world model(abstract)

Comments 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06235 2025-03-26 cs.CV cs.AI cs.RO 67%

NextStop: An Improved Tracker For Panoptic LIDAR Segmentation Data

Nirit Alkalay, Roy Orfaig, Ben-Zion Bobrovsky

专题命中 具身推理 :robotics(abstract);分类 cs.RO、cs.AI、cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00600 2024-11-04 cs.CV cs.AI cs.RO 67%

On Deep Learning for Geometric and Semantic Scene Understanding Using On-Vehicle 3D LiDAR

Li Li

专题命中 具身推理 :robotics(abstract);分类 cs.RO、cs.AI、cs.CV

Comments PhD thesis (Durham University, Computer Science), 149 pages (the 2024 BMVA Sullivan Doctoral Thesis Prize runner-up). Includes published content from arXiv:2407.10159 (ECCV 2024 ORAL), arXiv:2303.11203 (CVPR 2023), and arXiv:2406.10068 (3DV 2021), with minor revisions to the examined version: https://etheses.dur.ac.uk/15738/

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15373 2024-08-29 cs.CV cs.AI cs.LG 67%

Handling Geometric Domain Shifts in Semantic Segmentation of Surgical RGB and Hyperspectral Images

Silvia Seidlitz, Jan Sellner, Alexander Studier-Fischer, Alessandro Motta, Berkin Özdemir, Beat P. Müller-Stich, Felix Nickel, Lena Maier-Hein

专题命中 具身推理 :robotic(abstract);分类 cs.AI、cs.CV、cs.LG

Comments Silvia Seidlitz and Jan Sellner contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15843 2024-07-23 cs.CV cs.AI cs.RO 67%

CarFormer: Self-Driving with Learned Object-Centric Representations

Shadi Hamdan, Fatma Güney

专题命中 具身推理 :world model(abstract);分类 cs.RO、cs.AI、cs.CV

Comments Accepted to ECCV 2024, code and the pre-trained models can be found at https://kuis-ai.github.io/CarFormer/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01300 2024-07-19 cs.CV cs.AI cs.LG 67%

NeRF-MAE: Masked AutoEncoders for Self-Supervised 3D Representation Learning for Neural Radiance Fields

Muhammad Zubair Irshad, Sergey Zakharov, Vitor Guizilini, Adrien Gaidon, Zsolt Kira, Rares Ambrus

专题命中 具身推理 :robotics(abstract);分类 cs.AI、cs.CV、cs.LG

Comments Accepted to ECCV 2024. Project Page: https://nerf-mae.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07865 2024-05-31 cs.CV cs.AI cs.CL cs.LG 67%

Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Siddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang, Thomas Kollar, Dorsa Sadigh

专题命中 具身推理 :robotic(abstract);分类 cs.AI、cs.CV、cs.LG

Comments Published at ICML 2024. 22 pages, 11 figures. Training code and models: https://github.com/TRI-ML/prismatic-vlms. Evaluation code: https://github.com/TRI-ML/vlm-evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02288 2024-05-20 cs.CV cs.AI cs.RO 67%

Prospective Role of Foundation Models in Advancing Autonomous Vehicles

Jianhua Wu, Bingzhao Gao, Jincheng Gao, Jianhao Yu, Hongqing Chu, Qiankun Yu, Xun Gong, Yi Chang, H. Eric Tseng, Hong Chen, Jie Chen

专题命中 具身推理 :world model(abstract);分类 cs.RO、cs.AI、cs.CV

Comments 45 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15391 2024-02-26 cs.LG cs.AI cs.CV 67%

Genie: Generative Interactive Environments

Jake Bruce, Michael Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Bechtle, Feryal Behbahani, Stephanie Chan, Nicolas Heess, Lucy Gonzalez, Simon Osindero, Sherjil Ozair, Scott Reed, Jingwei Zhang, Konrad Zolna, Jeff Clune, Nando de Freitas, Satinder Singh, Tim Rocktäschel

专题命中 具身推理 :world model(abstract);分类 cs.AI、cs.CV、cs.LG

Comments https://sites.google.com/corp/view/genie-2024/

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12168 2024-01-23 cs.CV cs.CL cs.LG cs.RO 67%

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Boyuan Chen, Zhuo Xu, Sean Kirmani, Brian Ichter, Danny Driess, Pete Florence, Dorsa Sadigh, Leonidas Guibas, Fei Xia

专题命中 具身推理 :robotics(abstract);分类 cs.RO、cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏