arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 4106 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 模仿学习与强化学习 4106 篇

2007.08616 2020-07-20 cs.LG cs.AI cs.RO stat.ML 78%

Collision Avoidance Robotics Via Meta-Learning (CARML)

Abhiram Iyer, Aravind Mahadevan

专题命中 模仿学习与强化学习 :robotics(title);分类 cs.RO、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.06223 2020-05-14 cs.AI cs.LG cs.NE cs.RO 78%

DREAM Architecture: a Developmental Approach to Open-Ended Learning in Robotics

Stephane Doncieux, Nicolas Bredeche, Léni Le Goff, Benoît Girard, Alexandre Coninx, Olivier Sigaud, Mehdi Khamassi, Natalia Díaz-Rodríguez, David Filliat, Timothy Hospedales, A. Eiben, Richard Duro

专题命中 模仿学习与强化学习 :robotics(title);分类 cs.RO、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.10923 2020-03-25 cs.RO cs.AI cs.LG eess.SP 78%

Autonomous UAV Navigation: A DDPG-based Deep Reinforcement Learning Approach

Omar Bouhamed, Hakim Ghazzai, Hichem Besbes, Yehia Massoud

专题命中 模仿学习与强化学习 :navigation(title);分类 cs.RO、cs.AI、cs.LG

Comments This paper is accepted for publication in IEEE International Symposium on Circuits and Systems (ISCAS'20), Seville, Spain, Oct. 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.09351 2020-03-05 cs.LG cs.AI cs.NE cs.RO stat.ML 78%

Multi-objective Model-based Policy Search for Data-efficient Learning with Sparse Rewards

Rituraj Kaushik, Konstantinos Chatzilygeroudis, Jean-Baptiste Mouret

专题命中 模仿学习与强化学习 :robotics(abstract);robotic(abstract);robot learning(comments,journal_ref);分类 cs.RO、cs.AI、cs.LG

Comments Conference on Robot Learning (CoRL)- 2018; Code at https://github.com/resibots/kaushik_2018_multi-dex ; Video at https://youtu.be/9ZLwUxAAq6M

Journal ref Proceedings of the Conference on Robot Learning, PMLR 87:839-855, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.02543 2018-02-27 cs.RO cs.AI cs.LG 78%

Socially Compliant Navigation through Raw Depth Inputs with Generative Adversarial Imitation Learning

Lei Tai, Jingwei Zhang, Ming Liu, Wolfram Burgard

专题命中 模仿学习与强化学习 :navigation(title);分类 cs.RO、cs.AI、cs.LG

Comments ICRA 2018 camera-ready version. 7 pages, video link: https://www.youtube.com/watch?v=0hw0GD3lkA8

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.05255 2017-12-14 nlin.AO cond-mat.dis-nn cond-mat.stat-mech cs.NE q-bio.NC 78%

Criticality as It Could Be: organizational invariance as self-organized criticality in embodied agents

Miguel Aguilera, Manuel G. Bedia

专题命中 模仿学习与强化学习 :embodied agent(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1612.01086 2017-03-28 cs.AI cs.LG cs.RO 78%

Deep Learning of Robotic Tasks without a Simulator using Strong and Weak Human Supervision

Bar Hilleli, Ran El-Yaniv

专题命中 模仿学习与强化学习 :robotic(title);分类 cs.RO、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.08967 2015-12-01 cs.AI cs.LG cs.RO 78%

Robotic Search & Rescue via Online Multi-task Reinforcement Learning

Lisa Lee

专题命中 模仿学习与强化学习 :robotic(title);分类 cs.RO、cs.AI、cs.LG

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.09627 2026-06-04 cs.LG cs.RO cs.SY eess.SY 77%

Barrier-Certified Adaptive Reinforcement Learning with Applications to Brushbot Navigation

具有应用的障碍证书自适应强化学习:Brushbot导航

Motoya Ohnishi, Li Wang, Gennaro Notomista, Magnus Egerstedt

机构 * School of Electrical Engineering, Royal Institute of Technology(皇家理工学院电气工程学院) Georgia Institute of Technology(佐治亚理工学院) RIKEN Center for Advanced Intelligence Project(日本理化学研究所高级智能研究中心) School of Mechanical Engineering(机械工程学院)

专题命中 模仿学习与强化学习 :navigation(title);分类 cs.RO、cs.LG;robotics(journal_ref)

AI总结 本文提出了一种安全学习框架,结合自适应模型学习算法和障碍证书,用于具有可能非平稳智能体动态的系统。通过稀疏优化技术提取模型的动态结构,并利用学习的模型结合控制障碍证书来约束策略(反馈控制器),以保持安全性,即避免特定的不利状态空间区域。在某些条件下,保证了在安全被非平稳性破坏后,以李雅普诺夫稳定性的方式恢复安全。此外,将动作-价值函数近似重新公式化,使任何基于内核的非线性函数估计方法都能应用于我们的自适应学习框架。最后,保证了障碍证书策略优化的解是全局最优的,确保在温和条件下进行贪心策略改进。所得到的框架通过四旋翼无人机的模拟进行验证,该无人机此前在安全学习文献中被假设为平稳性,然后在动态未知、高度复杂且非平稳的Brushbot机器人上进行测试。

Comments ©2019 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

Journal ref Published in IEEE Transactions on Robotics, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19578 2024-10-21 cs.RO cs.LG cs.NE 77%

Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics

Norman Di Palo, Edward Johns

专题命中 模仿学习与强化学习 :robotics(title,comments);分类 cs.RO、cs.LG

Comments Published at Robotics: Science and Systems (RSS) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.05719 2022-09-16 cs.LG cs.RO stat.ML 77%

Smooth Exploration for Robotic Reinforcement Learning

Antonin Raffin, Jens Kober, Freek Stulp

专题命中 模仿学习与强化学习 :robotic(title);分类 cs.RO、cs.LG;robot learning(journal_ref)

Comments Code: https://github.com/DLR-RM/stable-baselines3/ Training scripts: https://github.com/DLR-RM/rl-baselines3-zoo/

Journal ref Proceedings of the 5th Conference on Robot Learning, PMLR 164:1634-1644, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.00127 2021-09-02 cs.AI cs.RO cs.SY eess.SY 77%

Cognitive science as a source of forward and inverse models of human decisions for robotics and control

Mark K. Ho, Thomas L. Griffiths

专题命中 模仿学习与强化学习 :robotics(title,comments);分类 cs.RO、cs.AI

Comments Invited submission for Annual Review of Control, Robotics, and Autonomous Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.06250 2021-04-19 cs.LG cs.RO 77%

Structured learning of rigid-body dynamics: A survey and unified view from a robotics perspective

A. René Geist, Sebastian Trimpe

专题命中 模仿学习与强化学习 :robotics(title,comments);分类 cs.RO、cs.LG

Comments - Added to title "... from a robotics perspective" - Added to Section 1.4 summary of used notation - Significantly extended the intro to mechanics in Section 2 - Adjusted Figure 4 to show the connection between model errors - Corrected typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08491 2026-08-11 cs.AI 新提交 77%

TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models

TrustRoboReward:面向多范式机器人奖励模型的偏好有序保序分数编辑方法

Yidong Wang, Yan Zhan, Ziteng Feng, Zhenyu Cui, Ziyi Zhou, Renzhao Liang, Jiaxuan Zhu, Zilei Yang, Yiran Zhao, Zhongkuan Mao, Bo Jia, Hanchu Ni, Chenggang Xie, Biao Liu, Yi Zhang, Yong Dai, Xiaozhu Ju, Wei Ye, Shikun Zhang

机构 * Peking University(北京大学) Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) University of Science and Technology of China(中国科学技术大学) Southeast University(东南大学) Southern University of Science and Technology(南方科技大学) Beijing University of Aeronautics and Astronautics(北京航空航天大学) Beijing Language and Culture University(北京语言大学) Sichuan University(四川大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 模仿学习与强化学习 :embodied AI(abstract);manipulation(abstract);robotic(abstract);分类 cs.AI

AI总结 针对现有机器人奖励模型的跨范式偏好与分数不一致问题,本文提出 TrustRoboReward 框架,通过 POISE 方法解决反转冲突,训练的 Qwen3-VL-4B 性能接近 GPT-5-mini,优于 RoboReward 基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22169 2026-07-21 cs.RO 版本更新 77%

VersualRL: Closed-Loop Verbal Reinforcement Learning with Visual Execution Feedback for Task-Level Robot Planning

闭环言语强化学习用于任务级机器人规划

Dmitrii Plotnikov, Iaroslav Kolomiets, Dmitrii Maliukov, Dmitrij Kosenkov, Daniia Zinniatullina, Artem Trandofilov, Georgii Gazaryan, Kirill Bogatikov, Timofei Kozlov, Mikhail Konenkov, Miguel Altamirano Cabrera, Dzmitry Tsetserukou

机构 * Intelligent Space Robotics Laboratory, Skolkovo Institute of Science and Technology(智能空间机器人实验室,斯克尔科夫科学与技术研究所)

专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);robotic(abstract);分类 cs.RO

AI总结 提出闭环言语强化学习框架,通过大语言模型与视觉语言模型交互,在符号规划层迭代优化行为树,实现可解释的任务级规划与执行不确定性下的自适应。

Comments Accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23371 2026-06-23 cs.RO 新提交 77%

TSD: A Physics-Inspired Trajectory Saliency Detector for Efficient Imitation Learning

TSD:一种受物理启发的轨迹显著性检测器用于高效模仿学习

Yiming Zhao, Gongrui Ma, Qingkai Li, Mingguo Zhao

机构 * Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 模仿学习与强化学习 :robot learning(abstract);manipulation(abstract);robotic(abstract);分类 cs.RO

AI总结 提出一种无需训练、即插即用的轨迹显著性检测器TSD,通过空间熵和向心加速度两个物理指标识别关键轨迹,实现数据集压缩和扩展,平均减少25%数据量仍保持或提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22509 2026-06-23 cs.AI 新提交 77%

Imagine to Ensure Safety in Hierarchical Reinforcement Learning

想象以确保分层强化学习的安全性

Gregory Gorbov, Artem Latyshev, Aleksandr I. Panov

机构 * Cognitive AI Systems Lab(认知人工智能系统实验室) Moscow Independent Research Institute of Artificial Intelligence(莫斯科独立人工智能研究所) FRC Computer Science RAS(俄罗斯科学院联邦研究中心计算机科学研究所)

专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);world model(abstract);分类 cs.AI

AI总结 提出结合可学习世界模型与高低层策略的分层方法,通过高层生成安全子目标、低层利用想象滚动减少不安全行为,在长时域任务中显著优于现有安全强化学习基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09381 2026-06-09 cs.RO 新提交 77%

ReGIL: Retrieval-Guided Imitation Learning from a Single Demonstration

ReGIL: 基于检索引导的单一示范模仿学习

Yuying Zhang, Francesco Verdoja, Wenyan Yang, Ville Kyrki

机构 * Aalto University(阿尔托大学)

专题命中 模仿学习与强化学习 :robot learning(abstract);manipulation(abstract);robotic(abstract);分类 cs.RO

AI总结 提出ReGIL框架,将单一示范作为外部记忆,通过检索引导探索、生成正则化缓冲和构建奖励,在LIBERO和Meta-World基准及真实机器人任务中显著提升成功率和训练效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08039 2026-06-09 cs.RO 新提交 77%

MuJoCo-Drones-Gym: A GPU-Accelerated Multi-Drone Simulator for Control and Reinforcement Learning

MuJoCo-Drones-Gym: 用于控制和强化学习的GPU加速多无人机模拟器

Manan Tayal

机构 * TAU-Intelligence

专题命中 模仿学习与强化学习 :robotics(abstract);navigation(abstract);robotic(abstract);分类 cs.RO

AI总结 提出基于MuJoCo物理引擎的GPU加速多无人机模拟器MuJoCo-Drones-Gym,支持任意数量Crazyflie 2.x纳米四旋翼,提供模块化物理模型、动作接口和观测空间,集成PettingZoo多智能体强化学习,涵盖悬停、速度跟踪等七种任务环境。

Comments 18 pages, 8 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14380 2026-06-08 cs.RO 版本更新 77%

CRAFT: Coaching Reinforcement Learning Autonomously using Foundation Models for Multi-Robot Coordination Tasks

CRAFT:利用基础模型自主教练强化学习以完成多机器人协调任务

Seoyeon Choi, Kanghyun Ryu, Jonghoon Ock, Negar Mehr

机构 * Department of Mechanical Engineering, University of California Berkeley(机械工程系,加州大学伯克利分校)

专题命中 模仿学习与强化学习 :robotics(abstract);manipulation(abstract);navigation(abstract);分类 cs.RO

AI总结 提出CRAFT框架,利用大语言模型分解任务、生成奖励函数,并通过视觉语言模型优化,实现多机器人协调学习,在四足导航和双臂操作任务中验证有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26478 2026-05-27 cs.RO cs.AI cs.CV cs.LG cs.SY eess.SY 77%

Efficient On-policy Visual-RL via Stochastic Decoupled Policy Gradient

基于随机解耦策略梯度的高效在策略视觉强化学习

Haoxiang You, Yilang Liu, Davis Zong, Qian Wang, Teeratham Vitchutripop, Qi Wang, Daniel Rakita, Ian Abraham

机构 * Yale University(耶鲁大学) Shanghai Jiao Tong University(上海交通大学) University of Sydney(悉尼大学)

专题命中 模仿学习与强化学习 :robotics(abstract);manipulation(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 提出随机解耦策略梯度(SDPG)方法,通过轨迹滚动的随机扰动估计策略梯度,在单GPU上数小时内端到端训练多样化的视觉运动控制策略,显著降低计算和内存开销,并在视觉MuJoCo基准测试中优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00361 2026-04-17 cs.LG 77%

Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach

基于原始能力的分层强化学习的直接偏好优化:双层方法

Utsav Singh, Souradip Chakraborty, Wesley A. Suttle, Brian M. Sadler, Derrik E. Asher, Anit Kumar Sahu, Mubarak Shah, Vinay P. Namboodiri, Amrit Singh Bedi

机构 * IIT Kanpur(印度理工学院坎浦尔分校) University of Maryland, College Park(马里兰大学学院公园分校) U.S. Army Research Laboratory(美国陆军研究实验室) University of Texas, Austin(德克萨斯大学奥斯汀分校) Oracle(Oracle公司) University of Central Florida(中央佛罗里达大学) University of Bath(巴斯大学)

专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);robotic(abstract);分类 cs.LG

AI总结 本文提出DIPPER框架,通过双层优化解决分层强化学习中的非平稳性和不可行子目标问题,采用直接偏好优化提升高层策略性能,实验证明在机器人导航和操作任务中取得显著改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10962 2026-04-14 cs.RO 77%

ScoRe-Flow: Complete Distributional Control via Score-Based Reinforcement Learning for Flow Matching

ScoRe-Flow:通过基于分数的强化学习实现完整分布控制的流匹配

Xiaotian Qiu, Lukai Chen, Jinhao Li, Qi Sun, Cheng Zhuo, Guohao Dai

机构 * Zhejiang University(浙江大学) Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 模仿学习与强化学习 :embodied AI(abstract);manipulation(abstract);robotic(abstract);分类 cs.RO

AI总结 本文提出ScoRe-Flow方法,通过基于分数的强化学习微调,实现对流匹配中均值和方差的解耦控制,提升了训练效率和任务成功率。

Comments 20 pages, 19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24241 2026-03-26 eess.SY cs.LG cs.SY 77%

C-STEP: Continuous Space-Time Empowerment for Physics-informed Safe Reinforcement Learning of Mobile Agents

C-STEP:连续时空赋能用于物理引导的移动智能体安全强化学习

Guihlerme Daubt, Adrian Redder

机构 * Faculty of Technology, Bielefeld University, Germany(比勒菲尔德大学技术学院,德国) Department of Electrical Engineering and Information Technology, Paderborn University, Germany(帕德博恩大学电气工程与信息科技系,德国)

专题命中 模仿学习与强化学习 :robotics(abstract);navigation(abstract);robotic(abstract);分类 cs.LG

AI总结 本文提出C-STEP方法,通过结合导航奖励设计物理引导的内在奖励,提升移动智能体在复杂环境中的安全导航能力,减少碰撞并优化任务完成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23804 2026-03-02 cs.LG 77%

Actor-Critic Pretraining for Proximal Policy Optimization

行为克隆预训练用于近端策略优化

Andreas Kernbach, Amr Elsheikh, Nicolas Grupp, René Nagel, Marco F. Huber

专题命中 模仿学习与强化学习 :robotics(abstract);manipulation(abstract);robotic(abstract);分类 cs.LG

AI总结 本文提出了一种基于行为克隆的actor-critic预训练方法,用于近端策略优化,通过初始化actor和critic网络提升样本效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22824 2025-12-30 cs.LG 77%

TEACH: Temporal Variance-Driven Curriculum for Reinforcement Learning

TEACH: 基于时间方差的强化学习课程学习

Gaurav Chaudhary, Laxmidhar Behera

专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);robotic(abstract);分类 cs.LG

AI总结 TEACH提出基于时间方差的课程学习方法,通过动态优先选择高不确定性目标,提升多目标强化学习的样本效率和策略学习效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03706 2025-10-07 cs.RO cs.AI cs.CV cs.LG 77%

EmbodiSwap for Zero-Shot Robot Imitation Learning

Eadom Dessalene, Pavan Mantripragada, Michael Maynord, Yiannis Aloimonos

机构 * department of Computer Science, University of Maryland, College Park, MD, 20742(计算机科学系,马里兰大学, College Park, MD)

专题命中 模仿学习与强化学习 :robotics(abstract);manipulation(abstract);分类 cs.RO、cs.AI、cs.CV

Comments Video link: https://drive.google.com/file/d/1UccngwgPqUwPMhBja7JrXfZoTquCx_Qe/view?usp=sharing

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00739 2025-10-02 cs.LG 77%

TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement Learning

Marco Bagatella, Matteo Pirotta, Ahmed Touati, Alessandro Lazaric, Andrea Tirinzoni

机构 * FAIR at Meta(Meta 的 FAIR) ETH Zurich(苏黎世联邦理工学院)

专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);world model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25402 2025-10-01 cs.RO 77%

Parallel Heuristic Search as Inference for Actor-Critic Reinforcement Learning Models

Hanlan Yang, Itamar Mishani, Luca Pivetti, Zachary Kingston, Maxim Likhachev

专题命中 模仿学习与强化学习 :robot learning(abstract);manipulation(abstract);robotic(abstract);分类 cs.RO

Comments Submitted for Publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05674 2025-09-15 cs.RO 77%

Integrating Diffusion-based Multi-task Learning with Online Reinforcement Learning for Robust Quadruped Robot Control

Xinyao Qin, Xiaoteng Ma, Yang Qi, Qihan Liu, Chuanyi Xue, Ning Gui, Qinyu Dong, Jun Yang, Bin Liang

机构 * Department of Automation, Tsinghua University(自动化系,清华大学)

专题命中 模仿学习与强化学习 :robotics(abstract);manipulation(abstract);robotic(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏