arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 304 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身导航 304 篇

2512.11736 2026-06-18 cs.RO 版本更新 89%

Bench-Push: Benchmarking Pushing-based Navigation and Manipulation Tasks for Mobile Robots

Bench-Push:基于推动的移动机器人导航与操作任务基准测试

Ninghan Zhong, Steven Caro, Megnath Ramesh, Rishi Bhatnagar, Avraiem Iskandar, Stephen L. Smith

机构 * Institute for Robotics and Intelligent Machines, Georgia Institute of Technology(机器人与智能机器研究所,佐治亚理工学院) Department of Electrical and Computer Engineering, University of Waterloo(电气与计算机工程系,滑铁卢大学) Department of Mechanical Engineering, University of Alberta(机械工程系,阿尔伯塔大学)

专题命中 具身导航 :manipulation(title,abstract);navigation(title,abstract);robotics(abstract);分类 cs.RO

AI总结 提出首个统一的推动式移动机器人导航与操作基准Bench-Push,包含多种模拟环境、新评估指标和基线实现,用于解决可移动障碍物环境中的机器人推动任务评估问题。

Comments Published in CRV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09497 2026-07-14 cs.RO cs.AI 版本更新 88%

Toward Autonomous Soft Robotic Endovascular Navigation via Imitation Learning

通过模仿学习实现自主软机器人血管内导航

Noah Barnes, Ji Woong Kim, Lingyun Di, Hannah Qu, Anuruddha Bhattacharjee, Miroslaw Janowski, Dheeraj Gandhi, Bailey Felix, Shaopeng Jiang, Olivia Young, Mark Fuge, Ryan D. Sochol, Jeremy D. Brown, Axel Krieger

机构 * Johns Hopkins University(约翰霍普金斯大学) Stanford University(斯坦福大学) McGill University(麦吉尔大学) University of Maryland(马里兰大学) Swiss Federal Institute of Technology in Lausanne (EPFL)(日内瓦联邦理工学院(EPFL)) ETH Zurich(苏黎世联邦理工学院)

专题命中 具身导航 :navigation(title,abstract);robotic(title,abstract);分类 cs.RO、cs.AI

AI总结 研究旨在实现自主软机器人血管内导航,开发基于变压器的模仿学习框架,通过目标条件等实现通用导航,在多种几何结构上训练并评估,策略成功率高,还进行了相关研究及扩展,提升了在未见几何结构上的成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17384 2026-07-07 cs.RO cs.CV 版本更新 88%

IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation

IndustryNav:探索动态工业导航中具身智能体的空间推理

Yifan Li, Lichi Li, Anh Dao, Xinyu Zhou, Wenjun Huang, Tianyi Ma, Yicheng Qiao, Zheda Mai, Daeun Lee, Zichen Chen, Pan Wang, Lehan Yang, Tianlong Wang, Zhen Tan, Sheng Li, Mohit Bansal, Yang Ni, Yu Kong

机构 * Michigan State University(密歇根州立大学) Independent Researcher(独立研究者) University of California, Irvine(加州大学尔湾分校) Ohio State University(俄亥俄州立大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of California, Santa Barbara(加州大学圣巴巴拉分校) University of Pittsburgh(匹兹堡大学) University of Virginia(弗吉尼亚大学) Arizona State University(亚利桑那州立大学) Purdue University Northwest(普渡大学西北分校)

专题命中 具身导航 :embodied agent(title,abstract);navigation(title,abstract);分类 cs.RO、cs.CV

AI总结 针对视觉大语言模型在空间推理的挑战,提出IndustryNav基准,利用动态工业场景评估,给出零样本导航管道及安全指标,研究发现模型在路径规划等方面有缺陷,凸显向主动探索等任务发展的需求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09241 2026-06-29 cs.CV cs.RO 版本更新 88%

RAE-NWM: Navigation World Model in Dense Visual Representation Space

RAE-NWM:密集视觉表示空间中的导航世界模型

Mingkun Zhang, Wangtian Shen, Fan Zhang, Haijian Qin, Zihao Pei, Ziyang Meng

机构 * Department of Precision Instrument, Tsinghua University(清华大学精密仪器系) University of Rochester(罗切斯特大学) Beijing Information Science and Technology University(北京信息科技大学)

专题命中 具身导航 :navigation(title,abstract);world model(title,abstract);分类 cs.RO、cs.CV

AI总结 提出RAE-NWM,在密集视觉表示空间中使用条件扩散Transformer建模导航动态,提升结构稳定性和动作精度,从而改善下游规划与导航。

Comments Code is available at: https://github.com/20robo/raenwm

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26148 2026-08-04 cs.RO 版本更新 88%

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation

具身智能体自主掌控:最小接口零样本智能体在视觉语言导航中可媲美工业级策略

Jian Zhou, Xunyi Zhao, Gengze Zhou, Zerui Li, Sihao Lin, Jiajun Liu, Qi Wu

机构 * Australian Institute for Machine Learning(澳大利亚机器学习研究所) Adelaide University(阿德莱德大学) Responsible AI Research Centre(负责任人工智能研究中心) CSIRO Data61(联邦科学与工业研究组织数据61中心) The University of Queensland(昆士兰大学)

专题命中 具身导航 :embodied agent(title,abstract);navigation(title,abstract);分类 cs.RO

AI总结 本研究提出智能具身控制,以零样本导航为测试平台,评估三种智能体框架,发现最小接口下部分框架可媲美工业级策略,同时揭示模型、框架与接口对具身智能体的互补作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23509 2026-07-21 cs.RO 版本更新 88%

Logic-Guided Socially-aware Robot Navigation World Model

逻辑引导的社会感知机器人导航世界模型

Weizheng Wang, Obi Ike, Soyun Choi, Sungeun Hong, Aniket Bera, Byung-Cheol Min

机构 * School of Applied and Creative Computing, Purdue University(应用与创意计算学院,普渡大学) Department of Computer Science, Purdue University(计算机科学系,普渡大学) Department of Applied Artificial Intelligence, Sungkyunkwan University(应用人工智能系,成均馆大学) Department of Computer Science and Department of Intelligent Systems Engineering, Indiana University Bloomington(计算机科学系和智能系统工程系,印第安纳大学布卢明顿分校)

专题命中 具身导航 :navigation(title,abstract);world model(title,abstract);分类 cs.RO

AI总结 提出NaviWM,通过结合结构化世界模型和逻辑驱动推理链,增强大语言模型在动态人类空间中生成社交合规且物理安全的导航决策的能力。

Comments The authors have decided to withdraw this manuscript due to concerns regarding its current scope, framing, and presentation. Please do not cite this version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18313 2026-06-23 cs.CV 版本更新 88%

OmniNWM: Omniscient Driving Navigation World Models

OmniNWM:全知驾驶导航世界模型

Bohan Li, Zhuang Ma, Dalong Du, Baorui Peng, Zhujin Liang, Zhenqiang Liu, Xianda Guo, Zheng Zhu, Chao Ma, Yueming Jin, Xin Jin, Hao Zhao, Wenjun Zeng

机构 * Shanghai Jiao Tong University(上海交通大学) Eastern Institute of Technology(东部技术研究所) PhiGent National University of Singapore(新加坡国立大学) Tsinghua University(清华大学) Wuhan University(武汉大学)

专题命中 具身导航 :navigation(title,abstract);world model(title,abstract);分类 cs.CV

AI总结 提出OmniNWM全知全景导航世界模型,在统一概率框架下处理状态、动作和奖励三个维度,实现多模态全景视频生成、零样本轨迹控制及内在奖励机制,在生成保真度和控制精度上达到SOTA。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14401 2026-07-07 cs.CV cs.AI 版本更新 87%

pFedNavi: Structure-Aware Personalized Federated Vision-Language Navigation for Embodied AI

pFedNavi:用于具身人工智能的结构感知个性化联邦视觉语言导航

Qingqian Yang, Hao Wang, Sai Qian Zhang, Jian Li, Yang Hua, Miao Pan, Tao Song, Zhengwei Qi, Haibing Guan

机构 * Openmind (Wuhu) Intelligent Robot Co., Ltd.(Openmind(芜湖)智能机器人有限公司) Shanghai Key Laboratory of Scalable Computing and Systems(上海可扩展计算与系统重点实验室) National Natural Science Foundation of China(国家自然科学基金委员会) US National Science Foundation(美国国家科学基金会)

专题命中 具身导航 :navigation(title,abstract);embodied AI(title);分类 cs.AI、cs.CV

AI总结 视觉语言导航需大量私有室内环境轨迹数据,存在隐私问题。本文提出针对视觉语言导航的结构感知个性化联邦学习框架pFedNavi,通过自适应确定特定客户端层并进行细粒度参数融合平衡全局知识与局部特性,性能优于基线。

Comments Accepted by the IEEE INFOCOM 2026 Workshop on Emerging Intelligent Networks (EIN)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25170 2026-07-17 cs.LG cs.AI cs.ET cs.RO 版本更新 87%

Grow-Prune-Freeze Networks: Adaptive & Continual Learning Technique for Olfactory Navigation

生长-剪枝-冻结网络:用于嗅觉导航的自适应与持续学习技术

Kordel K. France, Ovidiu Daescu

专题命中 具身导航 :navigation(title,abstract);robotics(abstract);world model(abstract);分类 cs.RO、cs.AI、cs.LG

AI总结 提出生长-剪枝-冻结(GPF)网络框架,通过动态调整策略网络层数实现持续学习,在湍流羽流导航任务中达到94%成功率,并推广到其他机器学习任务。

Comments Accepted as poster to the Reinforcement Learning in Big Worlds Workshop at the 2026 Reinforcement Learning Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16947 2026-06-29 cs.CV cs.RO 版本更新 87%

Image-based Geo-localization for Robotics: Are Black-box Vision-Language Models there yet?

基于图像的机器人地理定位:黑盒视觉语言模型是否已经足够?

Sania Waheed, Bruno Ferrarini, Michael Milford, Sarvapali D. Ramchurn, Shoaib Ehsan

机构 * University of Southampton(南安普顿大学) MyWay srl Queensland University of Technology(昆士兰科技大学) University of Essex(埃塞克斯大学)

专题命中 具身导航 :robotics(title,abstract);navigation(abstract,comments);robotic(abstract);分类 cs.RO、cs.CV

AI总结 本文首次系统研究黑盒生成式视觉语言模型作为独立零样本地理定位系统的潜力,发现其在粗粒度定位上表现良好,但在细粒度定位上因现实变化而显著退化。

Comments Accepted to the ICRA 2026 Workshop on Multi-Modal Spatial AI for Robust Navigation and Open-World Understanding (MM-SpatialAI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06833 2026-08-11 cs.RO 版本更新 85%

Unordered Landmark Visual Navigation

无序地标视觉导航

Hao Ren, Junzhe Zhu, Yihan Li, Zetong Bi, Le Zheng, Zhi Li, Yiqing Yuan, Zhaoliang Wan, Dizhe Zhang, Lu Qi, Hui Cheng

专题命中 具身导航 :navigation(title,abstract);embodied AI(abstract);分类 cs.RO

AI总结 针对无序图像导航的挑战,本文提出无需时间与里程计先验的ULVN框架,通过整合建图、定位与规划,在仿真和真实场景中性能优于现有最优方法。

Comments ECCV2026 Oral & Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00883 2026-08-05 cs.MM cs.CV cs.SD 版本更新 85%

Audio-Visual World Models: Learning Physically Grounded Multisensory Dynamics

视听世界模型:为具身智能体奠定多感官想象的基础

Jiahua Wang, Leqi Zheng, Jialong Wu, Yaoxin Mao, Shijie Cheng

机构 * Tsinghua University(清华大学) Beijing Institute of Technology(北京理工大学)

专题命中 具身导航 :world model(title,abstract);embodied agent(abstract);navigation(abstract);分类 cs.CV

AI总结 提出视听世界模型(AVWM)统一框架,通过条件扩散Transformer(AV-CDiT)联合预测双耳音频与视觉动态,在30小时基准AVW-4k上实现高保真多模态预测,并验证其在具身导航中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06998 2026-06-16 cs.RO 版本更新 85%

Raspi$^2$USBL: An open-source Raspberry Pi-Based Passive Inverted Ultra-Short Baseline Positioning System for Underwater Robotics

Raspi$^2$USBL:一种基于树莓派的开源被动倒置超短基线水下机器人定位系统

Jin Huang, Yingqiang Wang, Ying Chen

机构 * State Key Laboratory of Ocean Sensing, Ocean College of Zhejiang University(浙江省海洋传感重点实验室,浙江大学海洋学院) School of Oceanography, Shanghai Jiao Tong University(上海交通大学海洋学院)

专题命中 具身导航 :robotics(title,abstract);navigation(abstract);robotic(abstract);分类 cs.RO

AI总结 提出一种基于树莓派的低成本被动倒置超短基线定位系统,通过被动声学接收器和主动信标实现水下定位,在消声池、淡水湖和开放海域测试中达到0.1%斜距精度和0.1°方位角精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20055 2026-07-01 cs.RO cs.AI cs.CV 版本更新 85%

CoReLIN: Constraint-based Reasoning for Zero-shot Lifelong Interactive Navigation

CoReLIN: 基于约束推理的零样本终身交互式导航

Apoorva Vashisth, Manav Kulshrestha, Pranav Bakshi, Damon Conover, Guillaume Sartoretti, Aniket Bera

机构 * Purdue University(普渡大学) Indian Institute of Technology, Kharagpur(印度理工学院卡哈拉格普尔分校) DEVCOM Army Research Lab(DEVCOM陆军研究实验室) National University of Singapore(新加坡国立大学)

专题命中 具身导航 :navigation(title,abstract);manipulation(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 提出CoReLIN框架,利用LLM和场景图推理,在终身交互式导航中通过移动物体开辟路径并完成任务,显著提升长期效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21201 2026-06-15 cs.RO cs.AI cs.CV 版本更新 85%

Schrödinger's Navigator: Imagining an Ensemble of Futures for Zero-Shot Object Navigation

薛定谔的导航者:为零样本目标导航设想未来轨迹集合

Yu He, Da Huang, Zhenyang Liu, Zixiao Gu, Qiang Sun, Guangnan Ye, Yanwei Fu, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学) Shanghai University of International Business and Economics(上海对外经贸大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 具身导航 :navigation(title,abstract);world model(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 提出一种信念感知框架,在推理时通过轨迹条件化的3D世界模型设想多个未来场景,结合自适应遮挡物感知采样和未来感知价值图,提升零样本目标导航在遮挡严重环境中的隐蔽目标发现和风险感知路径选择。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10210 2026-07-07 cs.RO cs.CV 版本更新 84%

Nano-U: Efficient Terrain Segmentation for Tiny Robot Navigation

Nano-U:用于微型机器人导航的高效地形分割

Federico Pizzolato, Francesco Pasti, Nicola Bellotto

机构 * Dept of Information Engineering, University of Padua(信息工程系,帕多瓦大学)

专题命中 具身导航 :navigation(title);robotic(abstract,comments);robotics(abstract);分类 cs.RO、cs.CV

AI总结 本文提出Nano-U框架,通过量化感知蒸馏实现低功耗微型机器人高效地形分割,适用于Botanic Garden和TinyAgri数据集。

Comments Code repository: https://github.com/federico-pizz/Nano-U Accepted at Towards Autonomous Robotic Systems (TAROS) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26839 2026-08-10 cs.RO cs.CV 版本更新 84%

Ordinal Neural Collapse as a Representation Prior for Visual Navigation

序数神经坍缩作为视觉导航的表征先验

E-In Son, Jung-Taak Kim, Seung-Woo Seo

机构 * Seoul National University(首尔大学)

专题命中 具身导航 :navigation(title,abstract);robotic(abstract);分类 cs.RO、cs.CV

AI总结 针对视觉模仿学习中编码器学习到模糊、动作无关表征的问题,提出ORION方法,利用导航动作的序数结构显式组织表征空间,在仿真和真实环境中显著提升导航成功率。

Comments 27 pages, 14 figures. Supplementary material included

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27577 2026-08-04 cs.CV cs.RO 版本更新 84%

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation

结构化观察语言用于高效且通用的视觉-语言导航

Daojie Peng, Fulong Ma, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 具身导航 :navigation(title,abstract);embodied agent(abstract);分类 cs.RO、cs.CV

AI总结 本文提出SOL-Nav框架,通过将视觉观察转化为结构化语言描述,提升视觉-语言导航的效率和泛化能力,实验表明其在减少模型规模和训练数据依赖方面表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12630 2026-07-20 cs.RO cs.CV 版本更新 84%

Instance-Enriched Semantic Maps for Visual Language Navigation

用于视觉语言导航的实例增强语义地图

Jiho Hong, Eunae Kang, Sanghyun Kim, Young-Sik Shin

机构 * Department of Mechanical Engineering, Kyung Hee University(庆熙大学机械工程系) Advanced Institutes of Convergence Technology (AICT)(融合技术高级研究院) School of Mechanical Engineering, Kyungpook National University(庆北国立大学机械工程学院)

专题命中 具身导航 :navigation(title,abstract);embodied agent(abstract);分类 cs.RO、cs.CV

AI总结 研究视觉语言导航问题,提出实例增强语义地图框架,通过实例级二维半丰富信息映射、基于大语言模型的鲁棒查询处理和存储高效的语义表示,提升导航性能,在相关实验中取得优于基线的结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16806 2026-07-07 cs.AI cs.RO 版本更新 84%

Insect-inspired Visual Point-goal Navigation

一种高效的昆虫启发式方法用于视觉点目标导航

Yihe Lu, Barbara Webb

机构 * School of Informatics, University of Edinburgh(爱丁堡大学信息学院)

专题命中 具身导航 :navigation(title,abstract);embodied AI(abstract);分类 cs.RO、cs.AI

AI总结 本文提出了一种基于昆虫大脑结构的新型视觉点目标导航模型,通过模拟昆虫的联想学习和路径整合能力,实现高效导航。

Comments This work has been submitted for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17473 2026-06-24 cs.CV cs.AI 版本更新 84%

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation

双锚定:解决视觉语言导航中的状态漂移问题

Kangyi Wu, Pengna Li, Kailin Lyu, Xi Lin, Lin Zhao, Qingrong He, Jinjun Wang, Jianyi Liu

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) National Engineering Research Center for Visual Information and Applications(视觉信息与应用国家工程研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院) Xi’an Jiaotong University(西安交通大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Johns Hopkins University(约翰霍普金斯大学) Joy Future Academy, JD(京东探索研究院)

专题命中 具身导航 :navigation(title,abstract);world model(abstract);分类 cs.AI、cs.CV

AI总结 提出双锚定框架,通过指令进度锚定和记忆地标锚定分别解决进度漂移和记忆漂移,显著提升长场景导航成功率。

Comments Accepted by ECCV26

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28746 2026-07-10 cs.RO 版本更新 84%

He3-Seeker: Robotic Information Planning for Lunar Helium-3 Distribution Mapping

He3-Seeker:用于月球氦-3分布测绘的机器人信息规划

Dong Li, Yujie Zheng, Chengdeng Cao, Siyu Teng, Yuchen Li, Yang Gao, Long Chen

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Macau University of Science and Technology(澳门科学技术大学) China University of Petroleum-Beijing(中国石油大学(北京)) Wuhan University(武汉大学) Shenzhen University(深圳大学) Technical University of Munich(慕尼黑技术大学) Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 具身导航 :robotic(title,abstract);navigation(abstract);分类 cs.RO;robotics(comments)

AI总结 提出He3-Seeker框架,利用机器人信息规划引导自主导航与主动感知,实现基于多点钻探采样和原位分析的月球氦-3分布快速高保真测绘。

Comments Submitted to the International Conference on Space Robotics (iSpaRo) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03781 2026-08-14 cs.RO 版本更新 83%

OpenRC: An Open-Source Robotic Colonoscopy Framework for Multimodal Data Acquisition and Autonomy Research

OpenRC:一种用于多模态数据采集和自主性研究的开源机器人结肠镜框架

Siddhartha Kapuria, Mohammad Rafiee Javazm, Naruhiko Ikoma, Joga Ivatury, Mohammad Ali Nasseri, Nassir Navab, Farshid Alambeigi

机构 * Walker Department of Mechanical Engineering, The University of Texas at Austin(德克萨斯大学奥斯汀分校沃克机械工程系) Department of Surgical Oncology, Division of Surgery, The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心外科肿瘤学系) School of Medicine and Health, Technical University of Munich(慕尼黑工业大学医学与健康学院)

专题命中 具身导航 :robotic(title,abstract);navigation(abstract);分类 cs.RO

AI总结 OpenRC框架通过整合开源硬件和多模态数据集,为机器人结肠镜和手术自主性研究提供了可重复的基础,支持同时记录视频、操作指令、执行状态和末端位置,并验证了运动一致性和跨模态延迟。

Comments Abstract: Added repository and contribution statement

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19229 2026-08-07 cs.RO cs.AI cs.CV cs.LG cs.SY eess.SY 版本更新 83%

NavTrust: Benchmarking Trustworthiness for Embodied Navigation

NavTrust:体感导航的可信度基准测试

Huaide Jiang, Yash Chaudhary, Yuping Wang, Zehao Wang, Raghav Sharma, Manan Mehta, Yang Zhou, Lichao Sun, Zhiwen Fan, Zhengzhong Tu, Jiachen Li

机构 * Trustworthy Autonomous Systems Laboratory at the University of California, Riverside(加州大学河滨分校可信自主系统实验室) University of Michigan(密歇根大学) Workday University of Southern California(南加州大学) Texas A&M University(德克萨斯A&M大学) Lehigh University(莱斯大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 NavTrust通过系统性地对输入模态进行腐蚀,评估体感导航性能的鲁棒性,揭示了现有方法在真实环境中的性能缺陷,并提出改进策略。

Comments IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026); Project Website: https://navtrust.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10348 2026-07-30 cs.RO 版本更新 83%

Semantic Evidence Regulation via Relational Bias for Zero-Shot Object Navigation

通过关系归纳偏差重新思考具身导航

Weitao An, Chenghao Xu, Xu Yang, Cheng Deng

机构 * School of Electronic Engineering, Xidian University(西安电子科技大学电子工程学院) School of Information Science and Engineering, Hohai University(河海大学信息科学与工程学院)

专题命中 具身导航 :navigation(title,abstract);embodied agent(abstract);分类 cs.RO

AI总结 提出DB-Nav框架,利用激活偏置和抑制偏置双关系偏置重塑搜索空间,通过关系激活-抑制探索图调节前沿探索,显著提升目标导航成功率和路径效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11063 2026-07-22 cs.AI 版本更新 83%

AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation

AdvNav:视觉语言导航中基于行为引导的黑盒对抗攻击

Chenyang Li, Kaige Li, Zeyu Jiang, Changhao Chen

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 具身导航 :navigation(title,abstract);embodied AI(abstract);分类 cs.AI

AI总结 研究视觉语言导航系统对抗攻击,提出AdvNav框架,利用双粒度行为反馈构建替代目标,采用混合优化策略。在R2R数据集上评估,对两种模型取得较高攻击成功率,证明其有效性与通用性,揭示感知漏洞并为VLN模型设计提供见解。

Comments ACM International Conference on Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30293 2026-07-22 cs.RO 版本更新 83%

CSAR: Containerized System Architecture for Robotics

CSAR:面向机器人的容器化系统架构

Gregorio Ambrosio-Cestero, Cipriano Galindo, Javier Gonzalez-Jimenez, Jose-Raul Ruiz-Sarmiento

机构 * Machine Perception and Intelligent Robotics group (MAPIR), Dept. of System Engineering and Automation, Málaga Institute for Mechatronics Engineering and Cyber-Physical Systems (IMECH.UMA), University of Málaga(机器感知与智能机器人组(MAPIR),系统工程与自动化系,马拉加机电一体化工程与信息物理系统研究所(IMECH.UMA),马拉加大学)

专题命中 具身导航 :robotics(title,abstract);robotic(abstract);分类 cs.RO

AI总结 提出CSAR架构,结合LXC/LXD容器化、ROS 2/DDS通信和三层边缘基础设施,解决机器人团队在依赖隔离、兼容性、可复现性和异构部署中的挑战,并通过3D SLAM和GPU加速语义映射用例验证其有效性。

Comments 14 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06537 2026-07-14 cs.RO 版本更新 83%

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation

UniLM-Nav:零样本最后一英里导航的统一框架

Zhuofan Zhang, Tianxu Wang, Guoxi Zhang, Yixiong Lin, Xilin Wang, Hongming Xu, Qing Li, Song-Chun Zhu, Lifeng Fan

机构 * Tsinghua University(清华大学) State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,字节跳动公司人工智能研究院) Harbin Institute of Technology(哈尔滨工业大学) Peking University(北京大学)

专题命中 具身导航 :navigation(title,abstract);manipulation(abstract);分类 cs.RO

AI总结 研究移动操作中最后一英里导航问题,提出UniLM-Nav统一框架,通过多模态大语言模型后端分解任务为视图选择、功能接地和姿态推理,在OVMM基准上优于现有方法,还验证了在实际机器人上的适用性。

Comments Project page: https://unilm-nav.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11173 2026-07-02 cs.RO 版本更新 83%

Learning Category-level Last-meter Navigation from RGB Demonstrations of a Single-instance

从单实例RGB演示中学习类别级最后米导航

Tzu-Hsien Lee, Fidan Mahmudova, Karthik Desingh

机构 * University of Minnesota, Twin Cities(明尼苏达大学 Twin Cities 分校)

专题命中 具身导航 :navigation(title,abstract);manipulation(abstract);分类 cs.RO

AI总结 提出面向对象的模仿学习框架,利用RGB观测实现四足移动机械臂在最后米阶段的精确导航,无需深度或地图先验,在类别级泛化中达到高成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06480 2026-06-26 cs.RO 版本更新 83%

History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation

基于历史条件的时空视觉令牌剪枝用于高效视觉-语言导航

Qitong Wang, Yijun Liang, Ming Li, Tianyi Zhou, Christopher Rasmussen

机构 * Department of Computer and Information Sciences at the University of Delaware(德克萨斯大学达勒姆分校计算机与信息科学系) University of Maryland’s Department of Computer Science(马里兰大学计算机科学系) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 具身导航 :navigation(title,abstract);robotic(abstract);分类 cs.RO

AI总结 提出一种无需训练的时空视觉令牌剪枝框架,通过空间令牌选择和时空压缩减少冗余计算,在保持导航精度的同时显著提升推理效率,并在真实机器人上验证了低延迟指令跟随导航。

Comments International Conference on Intelligent Robots and Systems (IROS) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏