arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 14009 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身导航 14009 篇

2605.23257 2026-05-25 cs.RO cs.CV 81%

Turning Adaptation into Assets: Cross-Domain Bridging for Online Vision-Language Navigation

将适应转化为资产:面向在线视觉语言导航的跨域桥接

Zixuan Hu, Xuantuo Huang, Yancheng Li, Yichun Hu, Shengyong Xu, Ling-Yu Duan

机构 * School of Computer Science, Peking University, Beijing, China(北京大学计算机科学系) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室) School of Electronics, Peking University, Beijing, China(北京大学电子学院)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.CV

AI总结 提出 IDEA 框架,通过 Fisher 引导的软提示优化和动态资产库构建跨域桥接,解决视觉语言导航中测试时适应的灾难性遗忘和负迁移问题。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22816 2026-05-22 cs.RO cs.CV 81%

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation

AwareVLN: 基于自感知的视觉语言导航推理

Wenxuan Guo, Xiuwei Xu, Yichen Liu, Xiangyu Li, Hang Yin, Huangxing Chen, Wenzhao Zheng, Jianjiang Feng, Jie Zhou, Jiwen Lu

机构 * Tsinghua University(清华大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.CV

AI总结 本文提出AwareVLN框架,通过自感知推理机制实现端到端的视觉语言导航,解决了传统方法在理解代理、指令和场景关系上的不足,并在多个数据集上实现了优于现有方法的性能。

Comments Accepted to CVPR 2026. Project page: https://gwxuan.github.io/AwareVLN/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22036 2026-05-22 cs.CV cs.AI 81%

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation

GA-VLN: 用于高效视觉-语言导航的几何感知鸟瞰图表示

Jiahao Yang, Zihan Wang, Xiangyang Li, Xing Zhu, Yujun Shen, Yinghao Xu, Shuqiang Jiang

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Robbyant School of Computing, National University of Singapore(新加坡国立大学计算机学院) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.AI、cs.CV

AI总结 本文提出GA-VLN框架,通过引入几何感知的鸟瞰图表示(GA-BEV),整合显式和隐式几何信息,提升视觉-语言导航的效率和性能,实验表明其在仅使用导航数据的情况下取得了最先进的结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19501 2026-05-20 cs.RO cs.AI 81%

CANINE: Coaching Visually Impaired Users for Interactive Navigation with a Robot Guide Dog

CANINE: 为视觉障碍者提供交互导航的机器人导盲犬教学系统

Cunjun Yu, Zishuo Wang, Anxing Xiao, Linfeng Li, David Hsu

机构 * School of Computing(computing 学院) Smart Systems Institute(智能系统研究所)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.AI

AI总结 本文提出CANINE系统,通过个性化适应性语音反馈帮助视觉障碍者学习与机器人导盲犬的交互导航,通过分解复杂协调任务并分层训练提升学习效率和最终导航性能。

Comments Accepted to RSS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25646 2026-05-20 cs.CV cs.RO 81%

SAMe: A Semantic Anatomy Mapping Engine for Robotic Ultrasound

SAMe:一种用于机器人超声的语义解剖映射引擎

Jing Zhang, Duojie Chen, Wentao Jiang, Zihan Lou, Jianxin Liu, Xinwu Cui, Qinghong Zhao, Bo Du, Christoph F. Dietrich, Dacheng Tao

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Hubei Center for Applied Mathematics, Wuhan University(湖北应用数学中心,武汉大学) Department of Ultrasound, The Central Hospital of Wuhan(武汉市中心医院超声科) Department of Medical Ultrasound, Tongji Hospital, Tongji Medical College, Huazhong University of Science and Technology(同济医院,同济医学院,华中科技大学医学影像科) Department of Ultrasound in Medicine, Renmin Hospital of Wuhan University(武汉大学仁医医院医学超声科) University Hospital, Johann-Wolfgang-Goethe University Frankfurt am Main(法兰克福歌德大学医学院大学医院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 具身导航 :robotic(title,abstract);分类 cs.RO、cs.CV

AI总结 该研究提出SAMe,一种语义解剖映射引擎,通过提供显式的解剖先验层,解决机器人超声扫描初始化问题,实现了基于临床症状的解剖目标识别和控制指令生成,提高了自动扫描的准确性和效率。

Comments Supplementary information included. Code will be released at https://github.com/MiliLab/Echo-SAMe

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00634 2026-05-19 cs.RO cs.CV 81%

LiPS: Lightweight Panoptic Segmentation for Resource-Constrained Robotics

LiPS: 为资源受限机器人设计的轻量级全景分割

Calvin Galagain, Martyna Poreba, François Goulette, Cyrill Stachniss

机构 * Université Paris-Saclay, CEA LIST(巴黎-萨克雷大学,CEA LIST) U2IS, ENSTA, Institut Polytechnique de Paris(U2IS、ENSTA、巴黎理工学院) University of Bonn, Center for Robotics(波恩大学,机器人中心)

专题命中 具身导航 :robotics(title);robotic(abstract);分类 cs.RO、cs.CV

AI总结 本文提出LiPS,一种轻量级全景分割方法,通过简化特征提取和融合路径,在保持查询基于解码的同时,显著降低计算需求,实现与更重模型相当的精度和更高的吞吐量。

Comments Accepted to IEEE International Conference on Image Processing (ICIP) 2026, Paper #2070

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13934 2026-05-19 cs.RO cs.AI 81%

COLSON: Controllable Learning-Based Social Navigation via Diffusion-Based Reinforcement Learning

COLSON: 通过基于扩散的强化学习实现可控的社会导航

Kohei Matsumoto, Yuki Tomita, Yuki Hyodo, Ryo Kurazume

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.AI

AI总结 本文提出了一种基于扩散的强化学习方法,用于社会导航,通过灵活的动作分布提高了导航的适应性和可控性,同时能够适应未见过的场景。

Comments ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16947 2026-05-19 cs.CV cs.AI 81%

LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs

LightZeroNav: 基于轻量级VLMs的连续环境中零样本视觉语言导航

Kun Luo, Xiangyu Dong, Xiaoguang Ma, Haoran Zhao, Yaoming Zhou

机构 * Foshan Graduate School of Innovation, Northeastern University(创新研究生院,东北大学) Faculty of Robot Science and Engineering, Northeastern University(机器人科学与工程学院,东北大学) School of Aeronautic Science and Engineering, Beihang University(航空科学与工程学院,北航) QingniaoAI, China(清北AI,中国)

专题命中 具身导航 :navigation(title,abstract);分类 cs.AI、cs.CV

AI总结 本文提出LightZeroNav,通过轻量级VLMs解决连续环境中零样本视觉语言导航的三大瓶颈,无需特定训练或图搜索,在RGB观测和轻量级Qwen3-VL-8B模型下实现与GPT-4o相当的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09869 2026-05-18 cs.RO cs.CV 81%

ConsistNav: Closing the Action Consistency Gap in Zero-Shot Object Navigation with Semantic Executive Control

ConsistNav:通过语义执行控制关闭零样本物体导航中的动作一致性差距

Haosen Wang, Zhenyang Li, Yinqiang Zhang, Zongqi He, Lutao Jiang, Kai Li, Yizhou Zhao, Liaoyuan Fan, Wenjian Hou, Tingbang Liang, Yibin Wen, Defeng Gu

机构 * Sun Yat-sen University(中山大学) The University of Hong Kong(香港大学) Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) City University of Hong Kong(香港城市大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.CV

AI总结 本文提出ConsistNav框架,通过语义执行控制模块解决零样本物体导航中动作一致性问题,提升导航精度与鲁棒性。

Comments 13 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07642 2026-05-14 cs.AI cs.CL cs.CV 81%

Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents

拆解与构建:基于技能的视觉-语言导航代理混合

Tianyi Ma, Yue Zhang, Zehao Wang, Parisa Kordjamshidi

机构 * Michigan State University(密歇根州立大学) ESAT-PSI, KU Leuven(KU莱顿大学ESAT-PSI实验室)

专题命中 具身导航 :navigation(title,abstract);分类 cs.AI、cs.CV

AI总结 本文提出SkillNav框架,通过模块化技能推理提升视觉-语言导航性能,通过合成数据训练无监督的路由模型,实现对复杂场景的泛化能力。

Comments Accepted by ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07038 2026-05-11 cs.LG cs.MA cs.RO 81%

Learning Material-Aware Hamiltonian Risk Fields for Safe Navigation

学习材料感知的哈密顿风险场以实现安全导航

Aditya Sai Ellendula, Yi Wang, Chandrajit Bajaj

机构 * Department of Computer Science, University of Texas at Austin(德克萨斯大学奥斯汀分校计算机科学系) Oden Institute, University of Texas at Austin(德克萨斯大学奥斯汀分校奥登学院)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.LG

AI总结 本文提出一种基于材料感知的哈密顿风险场方法,通过引入上下文能量项,实现安全导航中的风险选择性控制,实验验证了其在不同场景下的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06487 2026-05-08 cs.CV cs.AI 81%

3D MRI Image Pretraining via Controllable 2D Slice Navigation Task

通过可控的2D切片导航任务进行3D MRI图像预训练

Yu Wang, Qingchao Chen

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Peking University(北京大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.AI、cs.CV

AI总结 本文提出通过可控2D切片导航任务预训练3D MRI图像,利用动作轨迹控制生成视频动作序列,提升解剖和空间表示学习能力。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11183 2026-05-08 cs.RO cs.CV cs.SY eess.SY 81%

Mitigating Error Accumulation in Continuous Navigation via Memory-Augmented Kalman Filtering

通过记忆增强的卡尔曼滤波缓解连续导航中的误差累积

Yin Tang, Jiawei Ma, Jinrui Zhang, Alex Jinpeng Wang, Deyu Zhang

机构 * Big Data Institute, Central South University, Changsha, China. Work done while working at CityUHK as a visiting scholar. Department of Computer Science \& Institute of Digital Medicine, City University of Hong Kong, Hong Kong, China School of Computer Science, Central South University, Changsha, China

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.CV

AI总结 本文提出NeuroKalman框架,通过先验预测和似然校正过程缓解连续导航中的状态漂移问题,实验表明其在TravelUAV基准上表现优异。

Comments ICML 2026 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02046 2026-05-05 cs.RO cs.LG 81%

NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks

NaviMaster: 学习GUI和具身导航任务的统一策略

Zhihao Luo, Wentao Yan, Jingyu Gong, Min Wang, Zhizhong Zhang, Xuhong Wang, Yuan Xie, Xin Tan

机构 * East China Normal University(东华大学) Shanghai AI Laboratory(上海人工智能实验室) SenseTime Research(商汤科技研究院)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.LG

AI总结 本文提出NaviMaster,通过统一框架整合GUI导航和具身导航,采用视觉目标轨迹收集管道、统一强化学习框架和距离感知奖励,提升泛化能力。

Comments ACL 2026 Main Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02854 2026-04-30 cs.RO cs.AI 81%

CoFL: Continuous Flow Fields for Language-Conditioned Navigation

CoFL:基于连续流场的语言条件导航

Haokun Liu, Zhaoqi Ma, Yicheng Chen, Masaki Kitagawa, Wentao Zhang, Zicen Xiong, Jinjie Li, Moju Zhao

机构 * Matterport3D ScanNet

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.AI

AI总结 CoFL通过端到端策略将鸟瞰图观测与语言指令映射为连续流场,实现工作空间条件下的场学习,提升导航精度与安全性,支持实时推理和闭环控制。

Comments 18 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09547 2026-04-30 cs.CV cs.AI 81%

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning

GoViG:通过多模态推理实现目标条件视觉导航指令生成

Fengyi Wu, Yifei Dong, Yilong Dai, Guangyu Chen, Qifeng Wu, Huiting Huang, Hang Wang, Qi Dai, Alexander G. Hauptmann, Zhi-Qi Cheng

机构 * University of Washington(华盛顿大学) The Hong Kong Polytechnic University(香港理工大学) Microsoft Research(微软研究院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.AI、cs.CV

AI总结 本文提出GoViG任务,通过多模态推理生成基于初始和目标状态的视觉导航指令,采用自回归多模态大语言模型提升空间准确性和语言清晰度,提出两种多模态推理策略并构建R2R-Goal数据集验证方法有效性。

Comments Accepted to ACL 2026 Findings. 22 pages, 12 figures, Code: https://github.com/F1y1113/GoViG

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13927 2026-04-30 cond-mat.mtrl-sci cs.AI cs.RO 81%

NIMS-OS: An automation software to implement a closed loop between artificial intelligence and robotic experiments in materials science

NIMS-OS:实现人工智能与材料科学机器人实验闭环的自动化软件

Ryo Tamura, Koji Tsuda, Shoichi Matsuda

机构 * Center for Basic Research on Materials(材料基础研究中心) National Institute for Materials Science(国家材料科学研究所) Graduate School of Frontier Sciences, The University of Tokyo(东京大学前沿科学研究生院) Research Center for Energy and Environmental Materials(GREEN)(能源与环境材料研究中心(GREEN)) Center for Advanced Battery Collaboration, Research Center for Energy and Environmental Materials (GREEN)(先进电池协作中心,能源与环境材料研究中心(GREEN))

专题命中 具身导航 :robotic(title,abstract);分类 cs.RO、cs.AI

AI总结 NIMS-OS通过整合AI技术与机器人实验,实现材料探索的闭环自动化,支持模块扩展和可视化工具,提供电解质自动探索的演示。

Comments 29 pages, 5 figures, 2 tables

Journal ref Science and Technology of Advanced Materials: Methods 3, 1, 2232297 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23970 2026-04-28 cs.AI cs.CV cs.HC cs.MA 81%

LLM-Guided Agentic Floor Plan Parsing for Accessible Indoor Navigation of Blind and Low-Vision People

基于大语言模型的代理楼层计划解析用于视障和低视力人士的无障碍室内导航

Aydin Ayanzadeh, Tim Oates

机构 * University of Maryland, Baltimore County(马里兰大学巴尔的摩分校)

专题命中 具身导航 :navigation(title,abstract);分类 cs.AI、cs.CV

AI总结 本文提出一种基于代理的框架,将单张楼层平面图转换为结构化的知识库,生成安全的无障碍导航指令。系统通过多代理模块生成空间知识图谱,并通过路径规划器生成导航指令,实验结果显示在真实场景和基准数据集上均优于单次调用基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09905 2026-04-22 cs.RO cs.CV 81%

Personalized Embodied Navigation for Portable Object Finding

面向便携物品寻找的个性化具身导航

Vishnu Sashank Dorbala, Bhrij Patel, Amrit Singh Bedi, Dinesh Manocha

机构 * University of Maryland, College Park(马里兰大学学院公园分校) University of Central Florida(中央佛罗里达大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.CV

AI总结 本文提出了一种面向动态环境的个性化习惯学习方法,通过引入Transit-Aware Planning算法提升便携物品寻找的性能,在模拟和现实环境中均取得显著效果。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10533 2026-04-21 cs.RO cs.CL cs.CV 81%

VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions

VLN-NF:具有假前提指令的视图与语言导航

Hung-Ting Su, Ting-Jun Wang, Jia-Fong Yeh, Min Sun, Winston H. Hsu

机构 * National Taiwan University(国立台湾大学) National Tsing Hua University(国立清华大学) DRIC

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.CV

AI总结 本文提出VLN-NF基准,通过引入假前提指令,要求智能体在目标不存在时进行探索并明确输出NOT-FOUND,采用ROAM方法在房间级导航与室内探索中取得最佳性能。

Comments ACL 2026 camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02972 2026-04-21 cs.CV cs.RO 81%

TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation

TagaVLM:面向视觉语言导航的拓扑感知全局动作推理

Jiaxing Liu, Zexi Zhang, Xiaoyan Li, Boyue Wang, Yongli Hu, Baocai Yin

机构 * School of Information Science and Technology, Beijing University of Technology, China(信息科学与技术学院,北京理工大学,中国) Imperial College London, U.K.(伦敦帝国理工学院,英国)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.CV

AI总结 TagaVLM通过引入拓扑结构增强视觉语言模型,提升空间推理能力,在R2R基准中取得最佳性能,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16967 2026-04-21 cs.RO cs.AI 81%

NaviFormer: A Deep Reinforcement Learning Transformer-like Model to Holistically Solve the Navigation Problem

NaviFormer:一种深度强化学习变换器模型,以整体解决导航问题

Daniel Fuertes, Andrea Cavallaro, Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García

机构 * Grupo de Tratamiento de Imágenes, Information Processing and Telecommunications Center, ETSI Telecomunicación, Universidad Politécnica de Madrid(图像处理与电信中心,信息处理与电信中心,电信工程学院,马德里理工大学) Idiap Research Institute(Idiap研究机构)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.AI

AI总结 NaviFormer通过预测高层路线和底层轨迹,整体解决导航问题,实验表明其在准确性和计算速度上表现优异,适用于实时任务。

Comments Published in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16298 2026-04-20 cs.CV cs.RO 81%

FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation

FineCog-Nav:整合细粒度认知模块以实现零样本多模态无人机导航

Dian Shao, Zhengzheng Xu, Peiyang Wang, Like Liu, Yule Wang, Jieqi Shi, Jing Huo

机构 * Northwestern Polytechnical University(西北工业大学) Nanjing University(南京大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.CV

AI总结 本文提出FineCog-Nav框架,通过细粒度认知模块提升无人机零样本多模态导航性能,构建了AerialVLN-Fine基准测试集,实验表明其在指令遵循、长视距规划和环境泛化方面优于基线方法。

Comments Accepted by CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19315 2026-04-16 cs.RO cs.AI 81%

Online Navigation Planning for Long-term Autonomous Operation of Underwater Gliders

长期自主运行的水下滑翔机在线导航规划

Victor-Alexandru Darvariu, Charlotte Z. Reed, Jan Stratmann, Bruno Lacerda, Benjamin Allsup, Stephen Woodward, Elizabeth Siddle, Trishna Saeharaseelan, Owain Jones, Dan Jones, Tobias Ferreira, Chloe Baker, Kevin Chaplin, James Kirk, Ashley Iceton-Morris, Ryan D. Patmore, Jeff Polton, Charlotte Williams, Christopher D. J. Auckland, Rob A. Hall, Alexandra Kokkinaki, Alvaro Lorenzo Lopez, Justin J. H. Buck, Nick Hawes

机构 * Oxford Robotics Institute(牛津机器人研究所) Department of Engineering Science, University of Oxford(牛津大学工程科学系) Optiver(Optiver公司) Stateful Robotics(Stateful Robotics公司) National Oceanography Centre(国家海洋研究中心) School of Environmental Sciences, University of East Anglia(东安格利亚大学环境科学学院) Scottish Association for Marine Science(苏格兰海洋科学协会)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.AI

AI总结 本文提出基于蒙特卡洛树搜索的在线导航规划方法,通过物理信息模拟器生成样本,实现水下滑翔机在不确定环境下的自主规划,提升任务持续时间和路径长度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12208 2026-04-15 cs.RO cs.AI 81%

Unveiling the Surprising Efficacy of Navigation Understanding in End-to-End Autonomous Driving

揭示端到端自动驾驶中导航理解的惊人效能

Zhihua Hua, Junli Wang, Pengfei LI, Qihao Jin, Bo Zhang, Kehua Sheng, Yilun Chen, Zhongxue Gan, Wenchao Ding

机构 * College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院) Didi Chuxing(滴滴出行) Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.AI

AI总结 本文提出SNG框架,通过真实导航模式高效表示全局导航信息,结合导航路径与分步信息,提升自动驾驶的全局与局部规划能力,实现无需辅助损失函数的高精度导航建模。

Comments 8 pages, 6 figures. ICRA 2026. Code available at https://fudan-magic-lab.github.io/SNG-VLA-web

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05467 2026-04-14 cs.CV cs.CL cs.RO 81%

MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation

MerNav:一种高度可泛化的记忆-执行-审查框架用于零样本物体目标导航

Dekang Qi, Shuang Zeng, Xinyuan Chang, Feng Xiong, Shichao Xie, Xiaolong Wu, Mu Xu

机构 * Amap, Alibaba Group(高德,阿里巴巴集团) Xi’an Jiaotong University(西安交通大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.CV

AI总结 本文提出MerNav框架,通过记忆、执行和审查模块提升视觉语言导航任务中的成功率和泛化能力,在多个数据集上均取得显著提升。

Comments 9 pages, 2 figures, 5 tables, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09604 2026-04-14 cs.AI cs.LG 81%

LLMs for Text-Based Exploration and Navigation Under Partial Observability

基于部分可观测性的文本基于探索与导航中的大型语言模型

Stephan Sandfuchs, Maximilian Melchert, Jörg Frochte

机构 * AKIS -- Interdisciplinary Institute for Applied AI and Data Science Ruhr(AKIS——鲁尔跨学科应用人工智能与数据科学研究所) Bochum University of Applied Sciences(波鸿应用科学大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.AI、cs.LG

AI总结 本文研究了大型语言模型在部分可观测环境下作为纯文本控制器的可行性,评估了九种不同LLM在探索和导航任务中的表现,发现推理调优模型在导航任务中表现更可靠,但效率仍低于理想路径。

Comments 15 pages, (to be published Springer Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering [LNICST] )

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08883 2026-04-13 cs.RO cs.AI 81%

HTNav: A Hybrid Navigation Framework with Tiered Structure for Urban Aerial Vision-and-Language Navigation

HTNav:一种具有分层结构的混合导航框架用于城市空域视觉-语言导航

Chengjie Fan, Cong Pan, Zijian Liu, Ningzhong Liu, Jie Qin

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(教育部脑机智能技术重点实验室)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.AI

AI总结 本文提出HTNav框架,结合模仿学习与强化学习,通过分层决策机制提升复杂城市环境中的导航精度与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02826 2026-04-13 cs.LG cs.AI 81%

From Navigation to Refinement: Revealing the Two-Stage Nature of Flow-based Diffusion Models through Oracle Velocity

从导航到细化:通过Oracle速度揭示基于流的扩散模型的两阶段本质

Haoming Liu, Jinnuo Liu, Yanhao Li, Liuyang Bai, Yunkai Ji, Yuanhe Guo, Shenji Wan, Hongyi Wen

机构 * Center for Data Science, New York University Shanghai(上海纽约大学数据科学中心) New York University(纽约大学)

专题命中 具身导航 :navigation(title,abstract);分类 cs.AI、cs.LG

AI总结 本文通过分析基于流的扩散模型的边际速度场,揭示其两阶段训练目标,解释了模型在早期阶段的全局布局生成和后期阶段的细节记忆行为,为模型设计和算法改进提供了指导。

Comments Accepted to CVPR 2026 (Findings track); 16 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20549 2026-04-10 cs.CV cs.RO 81%

Deep Learning-Powered Visual SLAM Aimed at Assisting Visually Impaired Navigation

基于深度学习的视觉SLAM:辅助视障导航

Marziyeh Bamdad, Hans-Peter Hutter, Alireza Darvishy

机构 * Institute of Computer Science, ZHAW School of Engineering(苏黎世应用科技大学工程学院计算机科学研究所) Department of Informatics, University of Zurich(苏黎世大学信息学系)

专题命中 具身导航 :navigation(title,abstract);分类 cs.RO、cs.CV

AI总结 本文提出SELM-SLAM3框架,结合SuperPoint和LightGlue提升特征提取与匹配鲁棒性,在低纹理、运动模糊等挑战性条件下实现更精确的定位与跟踪,提升视障导航可靠性。

Comments 8 pages, 7 figures, 4 tables. Published in the Proceedings of the 20th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications (VISIGRAPP 2025), VISAPP

详情

展开后加载摘要…

URL PDF HTML 收藏