arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2026-05-29 至 2026-05-29 共收录 82 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 人机交互与遥操作 1 篇

2602.13436 2026-05-29 cs.RO 57%

Force Sensing for Wearable Human-Robot Interfaces via Fluidic Innervation

用于可穿戴人机界面的力传感:基于流体神经支配

Noah Rubin, Ava Schraeder, Hrishikesh Sahu, Thomas C. Bulea, Lillian Chin

机构 * Rehabilitation Medicine Department, National Institutes of Health (NIH) Clinical Center(国家卫生研究院(NIH)临床中心康复医学部门) Department of Electrical and Computer Engineering, University of Texas at Austin(德克萨斯大学奥斯汀分校电气与计算机工程系)

专题命中 人机交互与遥操作 :robotic(abstract);分类 cs.RO

AI总结 通过3D打印硅胶垫中的空气通道测量力,实现线性响应,并验证其在等长扭矩、动态运动和机器人外骨骼中的应用。

Comments 6 pages, 7 figures, accepted to BioRob 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 机器人数据与评测 12 篇

2605.29564 2026-05-29 cs.RO 83%

VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation

VE2VF: 基于真实世界强化学习的视觉使能到无视觉蒸馏用于鲁棒接触丰富操作

Victor Kowalski, Chengxi Li, Dongheui Lee

机构 * Autonomous Systems, Technische Universitaet Wien (TU Wien)(自动系统,维也纳技术大学) Institute of Robotics and Mechatronics (DLR)(机器人与机电研究所)

专题命中 机器人数据与评测 :manipulation(title,abstract);robotic(abstract);分类 cs.RO

AI总结 提出一种人在环强化学习框架,通过教师-学生蒸馏将视觉使能策略的知识迁移到仅依赖本体感知的无视觉策略,在真实世界训练中实现鲁棒泛化,无需域随机化或数据增强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28883 2026-05-29 cs.AI cs.RO 81%

Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems

超低影响包裹式伐木(URIEL):提出一种利用空中机器人系统在热带森林中进行选择性可持续伐木和采后造林处理的新方法

Daniel Albiero, Gelton Fernando de Morais, Daniela Han, Flávio Roberto de Freitas Gonçalves, Artur Vitório Andrade Santos, Wesllen Lins de Araújo, Alessandra Maia Freire, Cláudio Kiyoshi Umezu, Mateus Peressin, Francesco Toscano, Admilson Írio Ribeiro, Alfeu J. Sguarezi Filho, Américo Ferraz Dias Neto, Angel Pontin Garcia

机构 * School of Agricultural Engineering, University of Campinas (UNICAMP)(坎皮纳斯大学农业工程学院) School of Mechanical Engineering, University of Campinas (UNICAMP)(坎皮纳斯大学机械工程学院) Depart. of Agricultural, Forestry, Food and Environmental Sciences, University of Basilicata(巴里奇塔大学农业、林业、食品与环境科学系) Sorocaba Environmental Engineering, São Paulo State University (UNESP)(圣保罗州立大学索罗卡巴环境工程) Center for Engineering, Modeling and Applied Social Sciences, Federal University of ABC (UFABC)(ABC联邦大学工程、建模和应用社会科学中心)

专题命中 机器人数据与评测 :robotics(title,abstract);分类 cs.RO、cs.AI

AI总结 提出URIEL方法,结合直升机伐木、机器人、AI和无人机采后造林处理,实现高经济可行性和几乎零附带损害,维持生态系统服务。

Comments 196 pages, 40 figures, A revolutionary technology to help protect tropical forests. It was developed, scaled, detailed, calculated, and simulated in an advanced computational environment, com viabilidade econômica e social. "E pur si muove"

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30346 2026-05-29 cs.CV 79%

YoCausal: How Far is Video Generation from World Model? A Causality Perspective

YoCausal: 视频生成距离世界模型还有多远?一个因果视角

You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee, Kaipeng Zhang, Yu-Lun Liu, Zhixiang Wang

机构 * National Yang Ming Chiao Tung University Shanda AI Research Tokyo

专题命中 机器人数据与评测 :world model(title,abstract);分类 cs.CV

AI总结 提出YoCausal基准,通过时间反转真实视频生成反事实样本,利用反向惊奇指数(RSI)和因果认知指数(CCI)评估视频扩散模型的因果理解能力,发现模型感知时间方向不等于理解因果关系,与人类水平存在显著差距。

Comments Project page: https://www.youzhexie.me/papers/YoCausal/index.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29410 2026-05-29 cs.RO 79%

A Progress-Aware Leader-Follower Midair Docking System for Dual-Drone Aerial Manipulation

面向双无人机空中操控的进度感知领航-跟随空中对接系统

Yifan Cai, Jan Ming Kevin Tan, Xiangqi Li, Chenzhe Jin, Narsimlu Kemsaram, Valerio Modugno

机构 * Department of Computer Science, University College London(计算机科学系,伦敦大学学院)

专题命中 机器人数据与评测 :manipulation(title,abstract);分类 cs.RO

AI总结 提出一种进度感知的领航-跟随双四旋翼空中对接平台,通过被动磁锁紧模块和阶段管理器实现可靠对接,并基于定量指标进行仿真与实验评估。

Comments This paper has been accepted for publication in the Proceedings of the 2026 IEEE 22nd International Conference on Automation Science and Engineering (CASE 2026), August 17-21, 2026, Shenyang, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23862 2026-05-29 cs.LG cs.AI cs.CL 62%

Graph Memory Transformer (GMT)

图记忆Transformer (GMT)

Nicola Zanarini, Niccolò Ferrari, Evelina Lamma

机构 * Bonfiglioli Engineering s.r.l.(博尼菲利工程公司) Department of Engineering, University of Ferrara(费拉拉大学工程学院) NAIS s.r.l.(NAIS公司)

专题命中 机器人数据与评测 :navigation(abstract);分类 cs.AI、cs.LG

AI总结 提出用显式学习的记忆图替换解码器-only Transformer中的前馈网络子层,保留自回归架构,实现可解释的记忆导航。

Comments 65 pages, 10 figures, 5 tables. Author list updated in arXiv metadata; no technical changes. Code available at https://github.com/Nemesis533/GMT-GraphMemoryTransformer

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29357 2026-05-29 cs.AI cs.LG cs.PL 62%

PassNet: Scaling Large Language Models for Graph Compiler Pass Generation

PassNet: 为图编译器通生成扩展大型语言模型

Yiqun Liu, Yingsheng Wu, Ruqi Yang, Enrong Zheng, Honglei Qiu, Sijun He, Tai Liang, Jingjing Wu, Yuhan Zhou, Yiwei Zhang, Dongyan Chen, Weihan Yi, Xinqi Li, Siqi Bao

机构 * Baidu, Inc.(百度公司)

专题命中 机器人数据与评测 :world model(abstract);分类 cs.AI、cs.LG

AI总结 针对编译器默认优化在长尾子图上性能不佳的问题,提出PassNet生态系统,包含大规模数据集和基准测试,通过微调小模型在少量轨迹上即可接近前沿模型性能。

Comments Code and data available at https://github.com/PaddlePaddle/PassNet

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29114 2026-05-29 cs.CR cs.LG cs.RO 62%

ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

ReasonBreak: 探测自动驾驶中具备推理能力的视觉-语言-行动模型的脆弱性

Mohammadreza Teymoorianfard, Jean-Philippe Monteuuis, Jonathan Petit, Amir Houmansadr

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Qualcomm(高通)

专题命中 机器人数据与评测 :manipulation(abstract);分类 cs.RO、cs.LG

AI总结 本文通过黑盒攻击方法,首次系统研究了具备推理能力的视觉-语言-行动模型在自动驾驶中面对真实输入扰动时的脆弱性,发现其推理和轨迹生成均易受攻击,导致碰撞率上升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22082 2026-05-29 cs.RO cs.LG 62%

CoRMA: Contrastive RMA for Contact-Rich Meta-Adaptation

CoRMA: 用于接触丰富元适应的对比RMA

Wentian Wang, Chutong Wen, Hongxu Ma, Wuhao Wang, Zhexiong Xue, Abdul Haseeb Nizamani, Dandi Zhou, Xinhai Sun, Jianqiao Zhu

机构 * Synthoid AI

专题命中 机器人数据与评测 :robotic(abstract);分类 cs.RO、cs.LG

AI总结 提出CoRMA框架,通过语义接触上下文和对比学习实现力主导装配任务的元适应,无需演示或梯度更新,在仿真和真实机器人上优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00324 2026-05-29 math.OC cs.CV cs.RO eess.SP 62%

Dual Quaternion SE(3) Synchronization with Recovery Guarantees

对偶四元数 SE(3) 同步及其恢复保证

Jianing Zhao, Linglingzhi Zhu, Anthony Man-Cho So

机构 * Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong, Shatin, NT, Hong Kong(系统工程与工程管理系,香港中文大学(深圳)) H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology, Atlanta, GA, USA(H. Milton Stewart工业与系统工程学院,佐治亚理工学院)

专题命中 机器人数据与评测 :robotics(abstract);分类 cs.RO、cs.CV

AI总结 采用对偶四元数表示,通过谱初始化和对偶四元数广义幂法实现 SE(3) 同步,并给出误差界和线性收敛保证。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30347 2026-05-29 cs.CV cs.GR 57%

NeuROK: Generative 4D Neural Object Kinematics

NeuROK:生成式4D神经物体运动学

Chen Geng, Guangzhao He, Yue Gao, Yunzhi Zhang, Shangzhe Wu, Jiajun Wu

机构 * Stanford University(斯坦福大学) University of Cambridge(剑桥大学) Cornell University(康奈尔大学)

专题命中 机器人数据与评测 :world model(abstract);分类 cs.CV

AI总结 提出基于Transformer的编码器-解码器模型NeuROK,通过学习物体潜在运动学参数化,实现从静态3D物体生成逼真的4D动态变形,克服了传统方法对预定义物理模型和特定类别的依赖。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30100 2026-05-29 cs.LG 57%

Chess-World-Model: A 10M-Game Benchmark for Exact State Tracking from Chess Move Sequences

Chess-World-Model: 一个用于从国际象棋走棋序列精确状态跟踪的1000万对局基准

Benjamin Walker, Terry Lyons

机构 * Mathematical Institute, University of Oxford(牛津大学数学研究所) Department of Mathematics, Imperial College London(伦敦帝国理工学院数学系)

专题命中 机器人数据与评测 :world model(abstract);分类 cs.LG

AI总结 提出一个基于1000万真实国际象棋对局的大规模状态跟踪基准,通过预测合法走棋序列后的棋盘状态,测试模型学习转换规则的能力,并发现循环模型优于Transformer,且随机均匀分布子集能揭示规模掩盖的失败。

Comments 20 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.02288 2026-05-29 cs.CV 57%

LabBuilder: Protocol-Grounded 3D Layout Generation for Interactable and Safe Laboratory

LabBuilder: 基于协议的可交互且安全的3D实验室布局生成

Jianbao Cao, Zhangrui Zhao, Bohan Feng, Zixuan Hu, Rui Li, Haiyuan Wan, Chenxi Li, Jingyuan Li, Wenzhe Cai, Lei Bai, Wanli Ouyang, Lingyu Duan, Di Huang, Minting Pan, Sha Zhang, Xinzhu Ma, Shixiang Tang, Dongzhan Zhou

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Wuhan University(武汉大学) Beihang University(北航) Peking University(北京大学) Tsinghua University(清华大学) Shanghai Jiaotong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 机器人数据与评测 :navigation(abstract);分类 cs.CV

AI总结 提出LabBuilder系统,通过协议引导和约束感知优化,从文本描述生成安全且可执行的3D实验室布局,显著优于现有方法。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他机器人 9 篇

2509.19318 2026-05-29 eess.SP cs.RO 79%

Scensory: Real-Time Robotic Olfactory Perception for Joint Identification and Source Localization

Scensory:用于联合识别和源定位的实时机器人嗅觉感知

Yanbaihui Liu, Erica Babusci, Claudia K. Gunsch, Boyuan Chen

机构 * Duke University(杜克大学)

专题命中 其他机器人 :robotic(title,abstract);分类 cs.RO

AI总结 提出一种基于学习的机器人嗅觉框架Scensory,通过廉价交叉敏感VOC传感器阵列的短时序信号,利用神经网络解码时间动态特征,同时实现真菌种类识别(最高89.85%准确率)和源定位(最高87.31%准确率)。

Comments Our project website is at: http://generalroboticslab.com/Scensory

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29144 2026-05-29 cs.RO cs.SY eess.SY 70%

Learning and Adaptation in Wire Arc Additive Manufacturing Bead Geometry Control

线弧增材制造焊道几何控制中的学习与自适应

Chen-Lung Lu, John Wen

机构 * Rensselaer Polytechnic Institute(伦塞拉尔理工学院)

专题命中 其他机器人 :robotics(abstract);robotic(abstract);分类 cs.RO

AI总结 针对线弧增材制造中热场与几何耦合的非线性动态过程,提出基于循环神经网络和一步预测控制的数据驱动方法,并通过逐层预测误差更新模型实现自适应,实验验证了在焊道高度和宽度一致性上的显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30328 2026-05-29 cs.CV 57%

Supercharging Thermal Gaussian Splatting with Depth Estimation

利用深度估计增强热高斯泼溅

Manoj Biswanath, Chenxin Cai, Hannah Schieber, Daniel Roth, Benjamin Busam

机构 * Technical University of Munich(技术大学慕尼黑) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Munich Institute of Robotics and Machine Intelligence (MIRMI)(慕尼黑机器人与人工智能研究所) Human-Centered Computing and Extended Reality Lab(以人为本计算与扩展现实实验室) TUM University Hospital(技术大学慕尼黑医院)

专题命中 其他机器人 :robotics(abstract);分类 cs.CV

AI总结 提出一种仅使用热红外图像和深度估计的单模态方法TDg,通过热到深度高斯泼溅推导辐射场,在渲染质量和训练时间上优于多模态基线。

Comments 8 pages, 4 figures. Accepted and will be published in ISPRS proceedings (ISPRS Congress 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29687 2026-05-29 cs.AI cs.LO 57%

Reliable Reasoning with Large Language Models via Preference-Based Maximum Satisfiability

基于偏好最大可满足性的大语言模型可靠推理

Pedro Orvalho, Marta Kwiatkowska, Guillem Alenyà, Felip Manyà

机构 * Artificial Intelligence Research Institute (IIIA) Consejo Superior de Investigaciones Científicas (CSIC)(人工智能研究所(IIIA)西班牙国家科学研究委员会(CSIC)) Department of Computer Science University of Oxford(计算机科学系牛津大学) Institut de Robòtica i Informàtica Industrial (IRI-CSIC-UPC)(机器人与信息工业研究所(IRI-CSIC-UPC))

专题命中 其他机器人 :robotics(abstract);分类 cs.AI

AI总结 提出一种混合推理方法,通过LLM生成代码将自然语言问题编码为偏好最大可满足性问题,由精确求解器求解并独立验证,显著提高可行性。

Comments 17 pages, 1 figure, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01159 2026-05-29 cs.RO 57%

Remote telepresence over large distances via robot avatars: case studies

通过机器人化身进行远距离远程呈现:案例研究

Mohamed Elobaid, Stefano Dafarra, Ehsan Ranjbari, Giulio Romualdi, Tomohiro Chaki, Tomohiro Kawakami, Takahide Yoshiike, Daniele Pucci

机构 * Artificial and Mechanical Intelligence AMI (Italian Insititute of Technology)(人工与机械智能AMI(意大利理工学院)) Frontier Robotics, Innovative Research Excellence(前沿机器人,创新研究卓越;本田研发) Honda R&D(机器学习与优化,曼彻斯特大学) Machine Learning and Optimisation, The University of Manchester

专题命中 其他机器人 :robotic(abstract);分类 cs.RO

AI总结 本文探讨了如何调整一种新提出的化身系统架构,以适应不同形态的机器人(轮式、腿式及多种手部与运动学结构),在带宽受限条件下实现洲际远程呈现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29572 2026-05-29 cs.RO cs.HC 57%

Learning to Feel Materials from Multisensory Tactile Data via Interpretable Models

通过可解释模型从多感官触觉数据中学习感知材料

Li Zou, Yasemin Vardar

机构 * Delft University of Technology (TU Delft), Department of Cognitive Robotics(代尔夫特理工大学(TU Delft),认知机器人学系)

专题命中 其他机器人 :robotic(abstract);分类 cs.RO

AI总结 提出一个可解释的计算框架,利用多感官触觉数据(包括按压、静态接触和滑动交互)建模人类材料感知与识别,发现热觉和顺应性线索对感知建模和材料分类至关重要。

Comments 12 pages, 3 figures, journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29505 2026-05-29 cs.CV 57%

ESAM++: Efficient Online 3D Perception on the Edge

ESAM++:边缘上的高效在线3D感知

Qin Liu, Lavisha Aggarwal, Saptarashmi Bandyopadhyay, Vikas Bahirwani, Marc Niethammer, Ehsan Adeli, Andrea Colaco

机构 * Stanford University(斯坦福大学) Google(谷歌) UC San Diego(圣地亚哥大学)

专题命中 其他机器人 :robotics(abstract);分类 cs.CV

AI总结 提出ESAM++,一种轻量级可扩展的在线3D场景感知方法,通过3D稀疏特征金字塔网络(SFPN)在边缘设备上实现高效、准确的3D实例分割。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10727 2026-05-29 physics.bio-ph cond-mat.stat-mech physics.comp-ph 50%

Run-and-Tumble Escape in Pursuit-Evasion Dynamics of Intelligent Active Particles

智能活性粒子追逃动力学中的跑动-转向逃逸

Segun Goh, Dennis Haustein, Gerhard Gompper

专题命中 其他机器人 :robotic(abstract)

AI总结 通过确定性自转向追逐者和随机认知逃逸者的模型,研究了二维空间中智能活性粒子的追逃博弈,发现逃逸者采用高风险后向机动或持续前向微调策略可显著影响捕获时间。

Comments 6 figures

Journal ref Advanced Intelligent Systems 8, e202500852 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28992 2026-05-29 eess.IV 50%

FRAPPE: Full Input, Residual Output Autoencoding with Projection Pursuit Encoder

FRAPPE: 全输入、残差输出自编码与投影追踪编码器

Dan Jacobellis, Neeraja J. Yadwadkar

专题命中 其他机器人 :robotics(abstract)

AI总结 提出FRAPPE框架,通过全输入预测残差输出并使用投影追踪编码器,实现零开销变速率编码,在CPU上实现实时1080p编码,压缩效率优于AVIF。

详情

展开后加载摘要…

URL PDF HTML 收藏