Collision Avoidance Robotics Via Meta-Learning (CARML)
专题命中 模仿学习与强化学习 :robotics(title);分类 cs.RO、cs.AI、cs.LG
视觉与机器人
机器人、具身智能、机器人学习、操作、导航和具身世界模型。
专题命中 模仿学习与强化学习 :robotics(title);分类 cs.RO、cs.AI、cs.LG
专题命中 模仿学习与强化学习 :robotics(title);分类 cs.RO、cs.AI、cs.LG
专题命中 模仿学习与强化学习 :navigation(title);分类 cs.RO、cs.AI、cs.LG
Comments This paper is accepted for publication in IEEE International Symposium on Circuits and Systems (ISCAS'20), Seville, Spain, Oct. 2020
专题命中 模仿学习与强化学习 :robotics(abstract);robotic(abstract);robot learning(comments,journal_ref);分类 cs.RO、cs.AI、cs.LG
Comments Conference on Robot Learning (CoRL)- 2018; Code at https://github.com/resibots/kaushik_2018_multi-dex ; Video at https://youtu.be/9ZLwUxAAq6M
Journal ref Proceedings of the Conference on Robot Learning, PMLR 87:839-855, 2018
专题命中 模仿学习与强化学习 :navigation(title);分类 cs.RO、cs.AI、cs.LG
Comments ICRA 2018 camera-ready version. 7 pages, video link: https://www.youtube.com/watch?v=0hw0GD3lkA8
专题命中 模仿学习与强化学习 :embodied agent(title,abstract)
专题命中 模仿学习与强化学习 :robotic(title);分类 cs.RO、cs.AI、cs.LG
专题命中 模仿学习与强化学习 :robotic(title);分类 cs.RO、cs.AI、cs.LG
Comments 7 pages
具有应用的障碍证书自适应强化学习:Brushbot导航
机构 * School of Electrical Engineering, Royal Institute of Technology(皇家理工学院电气工程学院) ; Georgia Institute of Technology(佐治亚理工学院) ; RIKEN Center for Advanced Intelligence Project(日本理化学研究所高级智能研究中心) ; School of Mechanical Engineering(机械工程学院)
专题命中 模仿学习与强化学习 :navigation(title);分类 cs.RO、cs.LG;robotics(journal_ref)
AI总结 本文提出了一种安全学习框架,结合自适应模型学习算法和障碍证书,用于具有可能非平稳智能体动态的系统。通过稀疏优化技术提取模型的动态结构,并利用学习的模型结合控制障碍证书来约束策略(反馈控制器),以保持安全性,即避免特定的不利状态空间区域。在某些条件下,保证了在安全被非平稳性破坏后,以李雅普诺夫稳定性的方式恢复安全。此外,将动作-价值函数近似重新公式化,使任何基于内核的非线性函数估计方法都能应用于我们的自适应学习框架。最后,保证了障碍证书策略优化的解是全局最优的,确保在温和条件下进行贪心策略改进。所得到的框架通过四旋翼无人机的模拟进行验证,该无人机此前在安全学习文献中被假设为平稳性,然后在动态未知、高度复杂且非平稳的Brushbot机器人上进行测试。
Comments ©2019 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
Journal ref Published in IEEE Transactions on Robotics, 2019
专题命中 模仿学习与强化学习 :robotics(title,comments);分类 cs.RO、cs.LG
Comments Published at Robotics: Science and Systems (RSS) 2024
专题命中 模仿学习与强化学习 :robotic(title);分类 cs.RO、cs.LG;robot learning(journal_ref)
Comments Code: https://github.com/DLR-RM/stable-baselines3/ Training scripts: https://github.com/DLR-RM/rl-baselines3-zoo/
Journal ref Proceedings of the 5th Conference on Robot Learning, PMLR 164:1634-1644, 2022
专题命中 模仿学习与强化学习 :robotics(title,comments);分类 cs.RO、cs.AI
Comments Invited submission for Annual Review of Control, Robotics, and Autonomous Systems
专题命中 模仿学习与强化学习 :robotics(title,comments);分类 cs.RO、cs.LG
Comments - Added to title "... from a robotics perspective" - Added to Section 1.4 summary of used notation - Significantly extended the intro to mechanics in Section 2 - Adjusted Figure 4 to show the connection between model errors - Corrected typos
TrustRoboReward:面向多范式机器人奖励模型的偏好有序保序分数编辑方法
机构 * Peking University(北京大学) ; Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) ; University of Science and Technology of China(中国科学技术大学) ; Southeast University(东南大学) ; Southern University of Science and Technology(南方科技大学) ; Beijing University of Aeronautics and Astronautics(北京航空航天大学) ; Beijing Language and Culture University(北京语言大学) ; Sichuan University(四川大学) ; Beijing University of Posts and Telecommunications(北京邮电大学)
专题命中 模仿学习与强化学习 :embodied AI(abstract);manipulation(abstract);robotic(abstract);分类 cs.AI
AI总结 针对现有机器人奖励模型的跨范式偏好与分数不一致问题,本文提出 TrustRoboReward 框架,通过 POISE 方法解决反转冲突,训练的 Qwen3-VL-4B 性能接近 GPT-5-mini,优于 RoboReward 基线。
闭环言语强化学习用于任务级机器人规划
机构 * Intelligent Space Robotics Laboratory, Skolkovo Institute of Science and Technology(智能空间机器人实验室,斯克尔科夫科学与技术研究所)
专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);robotic(abstract);分类 cs.RO
AI总结 提出闭环言语强化学习框架,通过大语言模型与视觉语言模型交互,在符号规划层迭代优化行为树,实现可解释的任务级规划与执行不确定性下的自适应。
Comments Accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026
TSD:一种受物理启发的轨迹显著性检测器用于高效模仿学习
机构 * Department of Automation, Tsinghua University(清华大学自动化系)
专题命中 模仿学习与强化学习 :robot learning(abstract);manipulation(abstract);robotic(abstract);分类 cs.RO
AI总结 提出一种无需训练、即插即用的轨迹显著性检测器TSD,通过空间熵和向心加速度两个物理指标识别关键轨迹,实现数据集压缩和扩展,平均减少25%数据量仍保持或提升性能。
想象以确保分层强化学习的安全性
机构 * Cognitive AI Systems Lab(认知人工智能系统实验室) ; Moscow Independent Research Institute of Artificial Intelligence(莫斯科独立人工智能研究所) ; FRC Computer Science RAS(俄罗斯科学院联邦研究中心计算机科学研究所)
专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);world model(abstract);分类 cs.AI
AI总结 提出结合可学习世界模型与高低层策略的分层方法,通过高层生成安全子目标、低层利用想象滚动减少不安全行为,在长时域任务中显著优于现有安全强化学习基线。
ReGIL: 基于检索引导的单一示范模仿学习
机构 * Aalto University(阿尔托大学)
专题命中 模仿学习与强化学习 :robot learning(abstract);manipulation(abstract);robotic(abstract);分类 cs.RO
AI总结 提出ReGIL框架,将单一示范作为外部记忆,通过检索引导探索、生成正则化缓冲和构建奖励,在LIBERO和Meta-World基准及真实机器人任务中显著提升成功率和训练效率。
MuJoCo-Drones-Gym: 用于控制和强化学习的GPU加速多无人机模拟器
机构 * TAU-Intelligence
专题命中 模仿学习与强化学习 :robotics(abstract);navigation(abstract);robotic(abstract);分类 cs.RO
AI总结 提出基于MuJoCo物理引擎的GPU加速多无人机模拟器MuJoCo-Drones-Gym,支持任意数量Crazyflie 2.x纳米四旋翼,提供模块化物理模型、动作接口和观测空间,集成PettingZoo多智能体强化学习,涵盖悬停、速度跟踪等七种任务环境。
Comments 18 pages, 8 figures, 7 tables
CRAFT:利用基础模型自主教练强化学习以完成多机器人协调任务
机构 * Department of Mechanical Engineering, University of California Berkeley(机械工程系,加州大学伯克利分校)
专题命中 模仿学习与强化学习 :robotics(abstract);manipulation(abstract);navigation(abstract);分类 cs.RO
AI总结 提出CRAFT框架,利用大语言模型分解任务、生成奖励函数,并通过视觉语言模型优化,实现多机器人协调学习,在四足导航和双臂操作任务中验证有效性。
基于随机解耦策略梯度的高效在策略视觉强化学习
机构 * Yale University(耶鲁大学) ; Shanghai Jiao Tong University(上海交通大学) ; University of Sydney(悉尼大学)
专题命中 模仿学习与强化学习 :robotics(abstract);manipulation(abstract);分类 cs.RO、cs.AI、cs.CV
AI总结 提出随机解耦策略梯度(SDPG)方法,通过轨迹滚动的随机扰动估计策略梯度,在单GPU上数小时内端到端训练多样化的视觉运动控制策略,显著降低计算和内存开销,并在视觉MuJoCo基准测试中优于基线方法。
基于原始能力的分层强化学习的直接偏好优化:双层方法
机构 * IIT Kanpur(印度理工学院坎浦尔分校) ; University of Maryland, College Park(马里兰大学学院公园分校) ; U.S. Army Research Laboratory(美国陆军研究实验室) ; University of Texas, Austin(德克萨斯大学奥斯汀分校) ; Oracle(Oracle公司) ; University of Central Florida(中央佛罗里达大学) ; University of Bath(巴斯大学)
专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);robotic(abstract);分类 cs.LG
AI总结 本文提出DIPPER框架,通过双层优化解决分层强化学习中的非平稳性和不可行子目标问题,采用直接偏好优化提升高层策略性能,实验证明在机器人导航和操作任务中取得显著改进。
ScoRe-Flow:通过基于分数的强化学习实现完整分布控制的流匹配
机构 * Zhejiang University(浙江大学) ; Shanghai Jiao Tong University(上海交通大学) ; Shanghai Innovation Institute(上海创新研究院)
专题命中 模仿学习与强化学习 :embodied AI(abstract);manipulation(abstract);robotic(abstract);分类 cs.RO
AI总结 本文提出ScoRe-Flow方法,通过基于分数的强化学习微调,实现对流匹配中均值和方差的解耦控制,提升了训练效率和任务成功率。
Comments 20 pages, 19 figures
C-STEP:连续时空赋能用于物理引导的移动智能体安全强化学习
机构 * Faculty of Technology, Bielefeld University, Germany(比勒菲尔德大学技术学院,德国) ; Department of Electrical Engineering and Information Technology, Paderborn University, Germany(帕德博恩大学电气工程与信息科技系,德国)
专题命中 模仿学习与强化学习 :robotics(abstract);navigation(abstract);robotic(abstract);分类 cs.LG
AI总结 本文提出C-STEP方法,通过结合导航奖励设计物理引导的内在奖励,提升移动智能体在复杂环境中的安全导航能力,减少碰撞并优化任务完成。
行为克隆预训练用于近端策略优化
专题命中 模仿学习与强化学习 :robotics(abstract);manipulation(abstract);robotic(abstract);分类 cs.LG
AI总结 本文提出了一种基于行为克隆的actor-critic预训练方法,用于近端策略优化,通过初始化actor和critic网络提升样本效率。
TEACH: 基于时间方差的强化学习课程学习
专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);robotic(abstract);分类 cs.LG
AI总结 TEACH提出基于时间方差的课程学习方法,通过动态优先选择高不确定性目标,提升多目标强化学习的样本效率和策略学习效率。
机构 * department of Computer Science, University of Maryland, College Park, MD, 20742(计算机科学系,马里兰大学, College Park, MD)
专题命中 模仿学习与强化学习 :robotics(abstract);manipulation(abstract);分类 cs.RO、cs.AI、cs.CV
Comments Video link: https://drive.google.com/file/d/1UccngwgPqUwPMhBja7JrXfZoTquCx_Qe/view?usp=sharing
机构 * FAIR at Meta(Meta 的 FAIR) ; ETH Zurich(苏黎世联邦理工学院)
专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);world model(abstract);分类 cs.LG
专题命中 模仿学习与强化学习 :robot learning(abstract);manipulation(abstract);robotic(abstract);分类 cs.RO
Comments Submitted for Publication
机构 * Department of Automation, Tsinghua University(自动化系,清华大学)
专题命中 模仿学习与强化学习 :robotics(abstract);manipulation(abstract);robotic(abstract);分类 cs.RO