arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

2026-05-01 至 2026-05-01 共收录 9 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身导航 9 篇

2604.27620 2026-05-01 cs.CV 83%

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation

SpaAct:基于课程适应的空间激活转换学习用于视觉语言导航

Pengna Li, Kangyi Wu, Shaoqing Xu, Fang Li, Hanbing Li, Lin Zhao, Kailin Lyu, Long Chen, Zhi-Xin Yang, Nanning Zheng

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) National Engineering Research Center for Visual Information and Applications(视觉信息与应用国家工程研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所) The State Key Laboratory of Internet of Things for Smart City(智能城市物联网国家重点实验室) Centre for Artificial Intelligence and Robotics(人工智能与机器人中心) University of Macau(澳门大学) Xiaomi EV(小i EV) School of Automation(自动化学院) Beijing Institute of Technology(北京理工大学) Institute of Automation(自动化研究所) Chinese Academy of Sciences(中国科学院)

专题命中 具身导航 :navigation(title,abstract);embodied agent(abstract);分类 cs.CV

AI总结 SpaAct通过引入空间激活任务和课程适应方法,提升视觉语言导航中动态空间感知能力,实现更高效的导航性能。

Comments Submmited to ACM MM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13559 2026-05-01 cs.AI 79%

OpAgent: Operator Agent for Web Navigation

OpAgent:网页导航的运算代理

Yuyu Guo, Wenjie Yang, Siyuan Yang, Ziyang Liu, Cheng Chen, Yuan Wei, Yun Hu, Yang Huang, Guoliang Hao, Dongsheng Yuan, Jianming Wang, Xin Chen, Hang Yu, Lei Lei, Peng Di

机构 * Ant Group(蚂蚁集团)

专题命中 具身导航 :navigation(title,abstract);分类 cs.AI

AI总结 本文提出OpAgent,通过在线强化学习优化网页导航策略,结合层级多任务微调和混合奖励机制,实现71.6%的高成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27010 2026-05-01 cs.HC 78%

Quantifying the Cost of Manual Navigation: A Comparison of Gesture-Based Magnification versus Direct Access Reading in Digital Layout-based Documents

量化手动导航的成本:手势放大与直接访问阅读在基于数字布局的文档中的比较

Sebastián Gallardo, Hui-Yin Wu, Dorian Mazauric, Pierre Kornprobst, Monica Di Meo, Stéphanie Baillif, Aurelie Calabrese

专题命中 具身导航 :navigation(title,abstract)

AI总结 研究比较手势放大与直接访问阅读在数字布局文档中的表现,发现大字号版块在阅读速度和目标定位效率上更优,且恢复了自然的阅读策略,同时降低用户负荷并提升偏好。

Journal ref IMX 2026 - International Conference on Interactive Media Experiences, Technological University of the Shannon, Jun 2026, Athlone, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20990 2026-05-01 cs.HC 71%

The Impact of Navigation on Proxemics in an Immersive Virtual Environment with Conversational Agents

导航对沉浸式虚拟环境中人际距离的影响:与对话代理的互动

Rose Connolly, Lauren Buck, Victor Zordan, Rachel McDonnell

专题命中 具身导航 :navigation(title)

AI总结 研究探讨了在沉浸式虚拟环境中,导航方式对人际距离的影响,发现 teleportation 使参与者与对话代理保持更近的距离,且女性与男性在距离上存在差异,自然行走则带来更高的自主感和身体所有权。

Comments for the associated supplementary video, see project page https://connolr3.github.io/TeleportationIPD Accepted for presentation at IEEE VR 2025 and for publication in a special issue of the IEEE Transactions on Visualization and Computer Graphics (IEEE TVCG) file was updated 30.04.2026 to include updated grant information

Journal ref IEEE Transactions on Visualization and Computer Graphics ( Volume: 31, Issue: 5, May 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27450 2026-05-01 cs.RO cs.AI 62%

RAY-TOLD: Ray-Based Latent Dynamics for Dense Dynamic Obstacle Avoidance with TDMPC

RAY-TOLD: 基于射线的任务导向潜在动力学用于密集动态障碍物避障与TDMPC

Seungho Han, Seokju Lee, Jeonguk Kang

机构 * School of Electrical Engineering, Hanyang University(翰阳大学电气工程学院) Mechatronics, Systems and Control Lab (MSC Lab), Department of Mechanical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(机械工程系,韩国科学技术院(KAIST)机电系统与控制实验室(MSC实验室)) Samsung Research, Samsung Electronics(三星研究所,三星电子)

专题命中 具身导航 :navigation(abstract);分类 cs.RO、cs.AI

AI总结 本文提出RAY-TOLD,结合物理基础MPPI的鲁棒性与强化学习的长视界,通过LiDAR中心的潜在动力学模型实现动态障碍物避障,提升导航可靠性与安全性。

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27383 2026-05-01 eess.IV cs.CV 57%

A Real-time Scale-robust Network for Glottis Segmentation in Nasal Transnasal Intubation

一种实时尺度鲁棒网络用于鼻内气管插管中的声带分割

Yang Zhou, Chaoyong Zhang, Ruoyi Hao, Huilin Pan, Yang Zhang, Hongliang Ren

机构 * School of Mechanical Engineering, Hubei University of Technology(湖北工业大学机械工程学院) National Key Laboratory for Novel Software Technology, Department of Computer Science and Technology, Nanjing University(南京大学新型软件技术国家重点实验室,计算机科学与技术系) Department of Electronic Engineering, The Chinese University of Hong Kong(香港中文大学电子工程系) Shun Hing Institute of Advanced Engineering, The Chinese University of Hong Kong(香港中文大学顺安先进工程研究院)

专题命中 具身导航 :navigation(abstract);分类 cs.CV

AI总结 本文提出了一种轻量级多感受野特征提取模块,用于提高鼻内气管插管中声带分割的鲁棒性和精度,实验表明其在三个数据集上表现优异,达到92.9%的mDice分数。

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27253 2026-05-01 cs.AI 57%

AutoSurfer -- Teaching Web Agents through Comprehensive Surfing, Learning, and Modeling

AutoSurfer -- 通过全面浏览、学习和建模教学网络代理

Fazle Elahi Faisal, Qianhui Wu, Baolin Peng, Jianfeng Gao

机构 * Microsoft Research(微软研究院)

专题命中 具身导航 :navigation(abstract);分类 cs.AI

AI总结 AutoSurfer通过系统性广度优先探索策略、任务合成引导和轨迹优化,实现全面覆盖网站操作空间,提升网络代理轨迹生成的准确性和多样性。

Comments 21 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17435 2026-05-01 cs.RO 57%

ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination

ImagineNav++: 通过场景想象促使视觉语言模型作为具身导航器

Teng Wang, Xinxin Zhao, Wenzhe Cai, Changyin Sun

机构 * School of Automation, Southeast University(东南大学自动化学院)

专题命中 具身导航 :navigation(abstract);分类 cs.RO

AI总结 本文提出ImagineNav++,通过场景想象将视觉语言模型用于无地图导航,利用想象模块生成高探索潜力的视点,并通过选择性聚焦记忆机制实现空间一致性,实验表明其在无地图导航中表现优异。

Comments 17 pages, 10 figures. arXiv admin note: text overlap with arXiv:2410.09874

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15350 2026-05-01 cs.RO cs.NE 57%

Nauplius Optimisation for Autonomous Hydrodynamics

幼体优化用于自主水动力学

Shyalan Ramesh, Scott Mann, Alex Stumpf

机构 * La Trobe University(拉特罗布大学)

专题命中 具身导航 :robotics(abstract);分类 cs.RO

AI总结 本文提出NOAH算法,结合水流感知漂移、不可逆锚定和群体通信,解决水下集群机器人在强流和持续感知中的优化问题。

Comments IEEE Access, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏