arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2026-05-29 至 2026-05-29 共收录 12 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 机器人数据与评测 12 篇

2605.29564 2026-05-29 cs.RO 83%

VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation

VE2VF: 基于真实世界强化学习的视觉使能到无视觉蒸馏用于鲁棒接触丰富操作

Victor Kowalski, Chengxi Li, Dongheui Lee

机构 * Autonomous Systems, Technische Universitaet Wien (TU Wien)(自动系统,维也纳技术大学) Institute of Robotics and Mechatronics (DLR)(机器人与机电研究所)

专题命中 机器人数据与评测 :manipulation(title,abstract);robotic(abstract);分类 cs.RO

AI总结 提出一种人在环强化学习框架,通过教师-学生蒸馏将视觉使能策略的知识迁移到仅依赖本体感知的无视觉策略,在真实世界训练中实现鲁棒泛化,无需域随机化或数据增强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28883 2026-05-29 cs.AI cs.RO 81%

Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems

超低影响包裹式伐木(URIEL):提出一种利用空中机器人系统在热带森林中进行选择性可持续伐木和采后造林处理的新方法

Daniel Albiero, Gelton Fernando de Morais, Daniela Han, Flávio Roberto de Freitas Gonçalves, Artur Vitório Andrade Santos, Wesllen Lins de Araújo, Alessandra Maia Freire, Cláudio Kiyoshi Umezu, Mateus Peressin, Francesco Toscano, Admilson Írio Ribeiro, Alfeu J. Sguarezi Filho, Américo Ferraz Dias Neto, Angel Pontin Garcia

机构 * School of Agricultural Engineering, University of Campinas (UNICAMP)(坎皮纳斯大学农业工程学院) School of Mechanical Engineering, University of Campinas (UNICAMP)(坎皮纳斯大学机械工程学院) Depart. of Agricultural, Forestry, Food and Environmental Sciences, University of Basilicata(巴里奇塔大学农业、林业、食品与环境科学系) Sorocaba Environmental Engineering, São Paulo State University (UNESP)(圣保罗州立大学索罗卡巴环境工程) Center for Engineering, Modeling and Applied Social Sciences, Federal University of ABC (UFABC)(ABC联邦大学工程、建模和应用社会科学中心)

专题命中 机器人数据与评测 :robotics(title,abstract);分类 cs.RO、cs.AI

AI总结 提出URIEL方法,结合直升机伐木、机器人、AI和无人机采后造林处理,实现高经济可行性和几乎零附带损害,维持生态系统服务。

Comments 196 pages, 40 figures, A revolutionary technology to help protect tropical forests. It was developed, scaled, detailed, calculated, and simulated in an advanced computational environment, com viabilidade econômica e social. "E pur si muove"

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30346 2026-05-29 cs.CV 79%

YoCausal: How Far is Video Generation from World Model? A Causality Perspective

YoCausal: 视频生成距离世界模型还有多远?一个因果视角

You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee, Kaipeng Zhang, Yu-Lun Liu, Zhixiang Wang

机构 * National Yang Ming Chiao Tung University Shanda AI Research Tokyo

专题命中 机器人数据与评测 :world model(title,abstract);分类 cs.CV

AI总结 提出YoCausal基准,通过时间反转真实视频生成反事实样本,利用反向惊奇指数(RSI)和因果认知指数(CCI)评估视频扩散模型的因果理解能力,发现模型感知时间方向不等于理解因果关系,与人类水平存在显著差距。

Comments Project page: https://www.youzhexie.me/papers/YoCausal/index.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29410 2026-05-29 cs.RO 79%

A Progress-Aware Leader-Follower Midair Docking System for Dual-Drone Aerial Manipulation

面向双无人机空中操控的进度感知领航-跟随空中对接系统

Yifan Cai, Jan Ming Kevin Tan, Xiangqi Li, Chenzhe Jin, Narsimlu Kemsaram, Valerio Modugno

机构 * Department of Computer Science, University College London(计算机科学系,伦敦大学学院)

专题命中 机器人数据与评测 :manipulation(title,abstract);分类 cs.RO

AI总结 提出一种进度感知的领航-跟随双四旋翼空中对接平台,通过被动磁锁紧模块和阶段管理器实现可靠对接,并基于定量指标进行仿真与实验评估。

Comments This paper has been accepted for publication in the Proceedings of the 2026 IEEE 22nd International Conference on Automation Science and Engineering (CASE 2026), August 17-21, 2026, Shenyang, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23862 2026-05-29 cs.LG cs.AI cs.CL 62%

Graph Memory Transformer (GMT)

图记忆Transformer (GMT)

Nicola Zanarini, Niccolò Ferrari, Evelina Lamma

机构 * Bonfiglioli Engineering s.r.l.(博尼菲利工程公司) Department of Engineering, University of Ferrara(费拉拉大学工程学院) NAIS s.r.l.(NAIS公司)

专题命中 机器人数据与评测 :navigation(abstract);分类 cs.AI、cs.LG

AI总结 提出用显式学习的记忆图替换解码器-only Transformer中的前馈网络子层,保留自回归架构,实现可解释的记忆导航。

Comments 65 pages, 10 figures, 5 tables. Author list updated in arXiv metadata; no technical changes. Code available at https://github.com/Nemesis533/GMT-GraphMemoryTransformer

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29357 2026-05-29 cs.AI cs.LG cs.PL 62%

PassNet: Scaling Large Language Models for Graph Compiler Pass Generation

PassNet: 为图编译器通生成扩展大型语言模型

Yiqun Liu, Yingsheng Wu, Ruqi Yang, Enrong Zheng, Honglei Qiu, Sijun He, Tai Liang, Jingjing Wu, Yuhan Zhou, Yiwei Zhang, Dongyan Chen, Weihan Yi, Xinqi Li, Siqi Bao

机构 * Baidu, Inc.(百度公司)

专题命中 机器人数据与评测 :world model(abstract);分类 cs.AI、cs.LG

AI总结 针对编译器默认优化在长尾子图上性能不佳的问题,提出PassNet生态系统,包含大规模数据集和基准测试,通过微调小模型在少量轨迹上即可接近前沿模型性能。

Comments Code and data available at https://github.com/PaddlePaddle/PassNet

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29114 2026-05-29 cs.CR cs.LG cs.RO 62%

ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

ReasonBreak: 探测自动驾驶中具备推理能力的视觉-语言-行动模型的脆弱性

Mohammadreza Teymoorianfard, Jean-Philippe Monteuuis, Jonathan Petit, Amir Houmansadr

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Qualcomm(高通)

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.RO、cs.LG

AI总结 本文通过黑盒攻击方法,首次系统研究了具备推理能力的视觉-语言-行动模型在自动驾驶中面对真实输入扰动时的脆弱性,发现其推理和轨迹生成均易受攻击,导致碰撞率上升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22082 2026-05-29 cs.RO cs.LG 62%

CoRMA: Contrastive RMA for Contact-Rich Meta-Adaptation

CoRMA: 用于接触丰富元适应的对比RMA

Wentian Wang, Chutong Wen, Hongxu Ma, Wuhao Wang, Zhexiong Xue, Abdul Haseeb Nizamani, Dandi Zhou, Xinhai Sun, Jianqiao Zhu

机构 * Synthoid AI

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO、cs.LG

AI总结 提出CoRMA框架,通过语义接触上下文和对比学习实现力主导装配任务的元适应,无需演示或梯度更新,在仿真和真实机器人上优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00324 2026-05-29 math.OC cs.CV cs.RO eess.SP 62%

Dual Quaternion SE(3) Synchronization with Recovery Guarantees

对偶四元数 SE(3) 同步及其恢复保证

Jianing Zhao, Linglingzhi Zhu, Anthony Man-Cho So

机构 * Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong, Shatin, NT, Hong Kong(系统工程与工程管理系,香港中文大学(深圳)) H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology, Atlanta, GA, USA(H. Milton Stewart工业与系统工程学院,佐治亚理工学院)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.RO、cs.CV

AI总结 采用对偶四元数表示,通过谱初始化和对偶四元数广义幂法实现 SE(3) 同步,并给出误差界和线性收敛保证。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30347 2026-05-29 cs.CV cs.GR 57%

NeuROK: Generative 4D Neural Object Kinematics

NeuROK:生成式4D神经物体运动学

Chen Geng, Guangzhao He, Yue Gao, Yunzhi Zhang, Shangzhe Wu, Jiajun Wu

机构 * Stanford University(斯坦福大学) University of Cambridge(剑桥大学) Cornell University(康奈尔大学)

专题命中 机器人数据与评测 :world model(abstract);分类 cs.CV

AI总结 提出基于Transformer的编码器-解码器模型NeuROK,通过学习物体潜在运动学参数化,实现从静态3D物体生成逼真的4D动态变形,克服了传统方法对预定义物理模型和特定类别的依赖。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30100 2026-05-29 cs.LG 57%

Chess-World-Model: A 10M-Game Benchmark for Exact State Tracking from Chess Move Sequences

Chess-World-Model: 一个用于从国际象棋走棋序列精确状态跟踪的1000万对局基准

Benjamin Walker, Terry Lyons

机构 * Mathematical Institute, University of Oxford(牛津大学数学研究所) Department of Mathematics, Imperial College London(伦敦帝国理工学院数学系)

专题命中 机器人数据与评测 :world model(abstract);分类 cs.LG

AI总结 提出一个基于1000万真实国际象棋对局的大规模状态跟踪基准,通过预测合法走棋序列后的棋盘状态,测试模型学习转换规则的能力,并发现循环模型优于Transformer,且随机均匀分布子集能揭示规模掩盖的失败。

Comments 20 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.02288 2026-05-29 cs.CV 57%

LabBuilder: Protocol-Grounded 3D Layout Generation for Interactable and Safe Laboratory

LabBuilder: 基于协议的可交互且安全的3D实验室布局生成

Jianbao Cao, Zhangrui Zhao, Bohan Feng, Zixuan Hu, Rui Li, Haiyuan Wan, Chenxi Li, Jingyuan Li, Wenzhe Cai, Lei Bai, Wanli Ouyang, Lingyu Duan, Di Huang, Minting Pan, Sha Zhang, Xinzhu Ma, Shixiang Tang, Dongzhan Zhou

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Wuhan University(武汉大学) Beihang University(北航) Peking University(北京大学) Tsinghua University(清华大学) Shanghai Jiaotong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 机器人数据与评测 :navigation(abstract);分类 cs.CV

AI总结 提出LabBuilder系统,通过协议引导和约束感知优化,从文本描述生成安全且可执行的3D实验室布局,显著优于现有方法。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏