arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 4111 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 模仿学习与强化学习 4111 篇

2402.15160 2024-03-04 cs.LG cs.AI 76%

Spatially-Aware Transformer for Embodied Agents

Junmo Cho, Jaesik Yoon, Sungjin Ahn

专题命中 模仿学习与强化学习 :embodied agent(title);分类 cs.AI、cs.LG

Comments ICLR 2024 Spotlight. First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.07457 2024-03-01 cs.LG cs.AI 76%

When Demonstrations Meet Generative World Models: A Maximum Likelihood Framework for Offline Inverse Reinforcement Learning

Siliang Zeng, Chenliang Li, Alfredo Garcia, Mingyi Hong

专题命中 模仿学习与强化学习 :world model(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10735 2023-11-21 cs.AI cs.RO 76%

Safe Navigation: Training Autonomous Vehicles using Deep Reinforcement Learning in CARLA

Ghadi Nehme, Tejas Y. Deo

专题命中 模仿学习与强化学习 :navigation(title);分类 cs.RO、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17330 2023-10-27 cs.LG cs.AI 76%

CQM: Curriculum Reinforcement Learning with a Quantized World Model

Seungjae Lee, Daesol Cho, Jonghae Park, H. Jin Kim

专题命中 模仿学习与强化学习 :world model(title);分类 cs.AI、cs.LG

Comments Accepted to NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16663 2023-09-29 cs.RO cs.LG 76%

HyperPPO: A scalable method for finding small policies for robotic control

Shashank Hegde, Zhehui Huang, Gaurav S. Sukhatme

专题命中 模仿学习与强化学习 :robotic(title);分类 cs.RO、cs.LG

Comments Website: https://sites.google.com/usc.edu/hyperppo

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06754 2023-06-13 cs.RO cs.AI 76%

Reinforcement Learning in Robotic Motion Planning by Combined Experience-based Planning and Self-Imitation Learning

Sha Luo, Lambert Schomaker

专题命中 模仿学习与强化学习 :robotic(title);分类 cs.RO、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12350 2023-05-30 cs.RO cs.LG cs.SY eess.SY 76%

Unsupervised Reward Shaping for a Robotic Sequential Picking Task from Visual Observations in a Logistics Scenario

Vittorio Giammarino, Andrew J Meyer, Kai Biegun

专题命中 模仿学习与强化学习 :robotic(title);分类 cs.RO、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.09034 2023-03-08 cs.RO cs.LG 76%

Model-Based Inverse Reinforcement Learning from Visual Demonstrations

Neha Das, Sarah Bechtle, Todor Davchev, Dinesh Jayaraman, Akshara Rai, Franziska Meier

专题命中 模仿学习与强化学习 :robotic(abstract,comments);manipulation(abstract);分类 cs.RO、cs.LG;robot learning(journal_ref)

Comments Accepted at the 4th Conference on Robotic Learning (CoRL 2020), Cambridge MA, USA

Journal ref Proceedings of the 2020 Conference on Robot Learning, PMLR 155:1930-1942, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.00421 2022-11-15 cs.RO cs.LG 76%

Air Learning: A Deep Reinforcement Learning Gym for Autonomous Aerial Robot Visual Navigation

Srivatsan Krishnan, Behzad Boroujerdian, William Fu, Aleksandra Faust, Vijay Janapa Reddi

专题命中 模仿学习与强化学习 :navigation(title);分类 cs.RO、cs.LG

Comments To Appear in Springer Machine Learning Journal (Special Issue on Reinforcement Learning for Real Life). Updating the title to match the Springer Machine Learning Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.14749 2022-09-27 cs.RO cs.AI 76%

Adaptive Risk-Tendency: Nano Drone Navigation in Cluttered Environments with Distributional Reinforcement Learning

Cheng Liu, Erik-Jan van Kampen, Guido C. H. E. de Croon

专题命中 模仿学习与强化学习 :navigation(title);分类 cs.RO、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.13842 2021-11-01 cs.RO cs.LG 76%

Model Predictive Actor-Critic: Accelerating Robot Skill Acquisition with Deep Reinforcement Learning

Andrew S. Morgan, Daljeet Nandha, Georgia Chalvatzaki, Carlo D'Eramo, Aaron M. Dollar, Jan Peters

专题命中 模仿学习与强化学习 :manipulation(abstract);robotic(abstract);robotics(comments,journal_ref);分类 cs.RO、cs.LG

Comments IEEE International Conference on Robotics and Automation (ICRA), Xi'an, China, 2021

Journal ref 2021 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.03138 2021-02-08 cs.RO cs.AI 76%

An advantage actor-critic algorithm for robotic motion planning in dense and dynamic scenarios

Chengmin Zhou, Bingding Huang, Pasi Fränti

专题命中 模仿学习与强化学习 :robotic(title);分类 cs.RO、cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.09134 2020-12-17 cs.MA cs.LG cs.RO 76%

Multi-agent navigation based on deep reinforcement learning and traditional pathfinding algorithm

Hongda Qiu

专题命中 模仿学习与强化学习 :navigation(title);分类 cs.RO、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.02646 2020-08-07 cs.RO cs.AI 76%

Deep Reinforcement Learning for Tactile Robotics: Learning to Type on a Braille Keyboard

Alex Church, John Lloyd, Raia Hadsell, Nathan F. Lepora

专题命中 模仿学习与强化学习 :robotics(title);分类 cs.RO、cs.AI

Comments Accepted in RAL and IROS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.14398 2020-05-29 cs.LG cs.RO stat.ML 76%

Robotic Table Tennis with Model-Free Reinforcement Learning

Wenbo Gao, Laura Graesser, Krzysztof Choromanski, Xingyou Song, Nevena Lazic, Pannag Sanketi, Vikas Sindhwani, Navdeep Jaitly

专题命中 模仿学习与强化学习 :robotic(title);分类 cs.RO、cs.LG

Comments V2: new URL of supplementary video. 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.07802 2019-08-29 cs.LG cs.AI cs.HC 76%

Learning 3D Navigation Protocols on Touch Interfaces with Cooperative Multi-Agent Reinforcement Learning

Quentin Debard, Jilles Steeve Dibangoye, Stéphane Canu, Christian Wolf

专题命中 模仿学习与强化学习 :navigation(title);分类 cs.AI、cs.LG

Comments 17 pages, 8 figures. Accepted at The European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases 2019 (ECMLPKDD 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.02805 2019-01-04 cs.NE cs.AI cs.RO 76%

DeepTraffic: Crowdsourced Hyperparameter Tuning of Deep Reinforcement Learning Systems for Multi-Agent Dense Traffic Navigation

Lex Fridman, Jack Terwilliger, Benedikt Jenik

专题命中 模仿学习与强化学习 :navigation(title);分类 cs.RO、cs.AI

Comments Neural Information Processing Systems (NIPS 2018) Deep Reinforcement Learning Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.05256 2018-12-14 cs.RO cs.AI 76%

Learning to Communicate: A Machine Learning Framework for Heterogeneous Multi-Agent Robotic Systems

Hyung-Jin Yoon, Huaiyu Chen, Kehan Long, Heling Zhang, Aditya Gahlawat, Donghwan Lee, Naira Hovakimyan

专题命中 模仿学习与强化学习 :robotic(title);分类 cs.RO、cs.AI

Comments AIAA SciTech 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.07543 2017-01-27 cs.LG astro-ph.IM cs.RO 76%

FPGA Architecture for Deep Learning and its application to Planetary Robotics

Pranay Gankidi, Jekan Thangavelautham

专题命中 模仿学习与强化学习 :robotics(title);分类 cs.RO、cs.LG

Comments 8 pages, 10 figures in Proceedings of the IEEE Aerospace Conference 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06451 2025-05-13 cs.RO 76%

Adaptive Wiping: Adaptive contact-rich manipulation through few-shot imitation learning with Force-Torque feedback and pre-trained object representations

Chikaha Tsuji, Enrique Coronado, Pablo Osorio, Gentiane Venture

机构 * Department of Mechano-Informatics, The University of Tokyo(东京大学机械信息学系) National Institute of Advanced Industrial Science and Technology(国家先进工业科学与技术研究院) Department of Mechanical Systems Engineering, Tokyo University of Agriculture and Technology(东京农业大学机械系统工程系) Department of Mechanical Engineering, The University of Tokyo(东京大学机械工程系)

专题命中 模仿学习与强化学习 :manipulation(title);分类 cs.RO;robotics(journal_ref)

Journal ref IEEE Robotics and Automation Letters, vol.10, no.1, pp.240-247, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06139 2024-12-10 cs.LG cs.SY eess.SY 76%

Bounded Exploration with World Model Uncertainty in Soft Actor-Critic Reinforcement Learning Algorithm

Ting Qiao, Henry Williams, David Valencia, Bruce MacDonald

专题命中 模仿学习与强化学习 :world model(title);分类 cs.LG;robotics(comments)

Comments 8 pages, 7 figures. Accepted as a poster presentation in the Australian Robotics and Automation Association (2023)

Journal ref ISBN: 978-0-6455655-2-2 ISSN: 1448-2053

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.03195 2022-04-08 cs.RO 76%

3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic Surgery

Bin Li, Ruofeng Wei, Jiaqi Xu, Bo Lu, Chi-Hang Yee, Chi-Fai Ng, Pheng-Ann Heng, Qi Dou, Yun-Hui Liu

专题命中 模仿学习与强化学习 :robotic(title);分类 cs.RO;robotics(comments)

Comments 7 pages, 7 figures, 2022 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.13687 2021-12-21 cs.LG 76%

panda-gym: Open-source goal-conditioned environments for robotic learning

Quentin Gallouédec, Nicolas Cazin, Emmanuel Dellandréa, Liming Chen

专题命中 模仿学习与强化学习 :robotic(title);分类 cs.LG;robot learning(comments)

Comments NeurIPS 2021 Workshop on Robot Learning: Self-Supervised and Lifelong Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.05762 2018-10-25 cs.RO 76%

GPU-Accelerated Robotic Simulation for Distributed Reinforcement Learning

Jacky Liang, Viktor Makoviychuk, Ankur Handa, Nuttapong Chentanez, Miles Macklin, Dieter Fox

专题命中 模仿学习与强化学习 :robotic(title);分类 cs.RO;robot learning(comments)

Comments Accepted and to appear at the Conference on Robot Learning (CoRL) 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17760 2026-07-21 cs.LG cs.AI cs.RO 新提交 75%

Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning

泛化与引导:用于少样本逆强化学习的奖励分解

Ziyi Liu, Grace Zhang

专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);分类 cs.RO、cs.AI、cs.LG

AI总结 研究少样本逆强化学习(FM - IRL)问题,提出多任务判别器近邻引导的IRL(MPG)方法,通过学习两个互补奖励组件,在多种有显著变化的任务上验证有效性,平均成功率达81.2%,优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09336 2026-07-13 cs.LG cs.AI cs.RO 新提交 75%

Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning

用于高效离线强化学习的捷径轨迹规划

Guanquan Wang, Yoshimasa Tsuruoka

机构 * The University of Tokyo(东京大学)

专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);分类 cs.RO、cs.AI、cs.LG

AI总结 研究针对离线强化学习中轨迹规划器的问题,提出捷径轨迹规划(STP)框架,将捷径模型作为轨迹生成器,单阶段训练条件捷径轨迹模型,支持可调推理,用增强可行性感知校正的评论家选候选计划,在多任务基准测试中性能强且简化训练管道。

Comments 16 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16861 2026-07-13 cs.RO cs.AI cs.LG 版本更新 75%

ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning

ReinforceGen:具有自动数据生成和强化学习的混合技能策略

Zihan Zhou, Animesh Garg, Ajay Mandlekar, Caelan Garrett

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所) Georgia Institute of Technology(佐治亚理工学院) NVIDIA Research(NVIDIA研究)

专题命中 模仿学习与强化学习 :robotics(abstract);manipulation(abstract);分类 cs.RO、cs.AI、cs.LG

AI总结 针对机器人长期操纵挑战,ReinforceGen系统结合任务分解、数据生成、模仿学习与运动规划形成初始方案,经强化学习微调,在Robosuite数据集基准测试中成功率达80%,消融研究显示微调使性能平均提升89%,实际评估也有显著改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23640 2026-06-23 cs.LG cs.AI cs.RO stat.ML 新提交 75%

Learning Process Rewards via Success Visitation Matching for Efficient RL

通过成功访问匹配学习过程奖励以实现高效强化学习

Raymond Tsao, Andrew Wagenmaker, Sergey Levine

机构 * UC Berkeley(加州大学伯克利分校)

专题命中 模仿学习与强化学习 :manipulation(abstract);robotic(abstract);分类 cs.RO、cs.AI、cs.LG

AI总结 针对稀疏奖励导致的信用分配难题,提出通过判别器区分成功/失败轨迹,生成密集过程奖励以匹配成功访问分布,在不改变最优策略下加速机器人控制策略微调。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09758 2026-06-09 cs.RO cs.AI cs.LG 新提交 75%

Difference-Aware Retrieval Policies for Imitation Learning

差异感知的模仿学习检索策略

Quinn Pfeifer, Ethan Pronovost, Paarth Shah, Khimya Khetarpal, Siddhartha Srinivasa, Abhishek Gupta

机构 * Paul G. Allen School of Computer Science & Engineering, University of Washington(华盛顿大学保罗·G·艾伦计算机科学与工程学院) Toyota Research Institute(丰田研究所) Google DeepMind(谷歌DeepMind) Mila

专题命中 模仿学习与强化学习 :manipulation(abstract);robotic(abstract);分类 cs.RO、cs.AI、cs.LG

AI总结 提出DARP,一种半参数检索式模仿学习方法,通过基于k近邻的局部邻域结构重参数化,解决行为克隆的分布外泛化问题,在连续控制和机器人操作任务中性能提升15-46%。

Comments 12 pages, 7 figures, 3 tables. Accepted to ICLR 2026. Code and demos available at https://weirdlabuw.github.io/darp-site/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00083 2026-06-02 cs.LG cs.AI cs.RO 75%

From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models

从演示到奖励:VLM奖励模型的测试时提示优化

Christian Gumbsch, Leonardo Barcellona, Lennard Schünemann, Platon Karageorgis, Andrii Zadaianchuk, Zehao Wang, Sergey Zakharov, Fabien Despinoy, Rahaf Aljundi, Efstratios Gavves

机构 * University of Amsterdam(阿姆斯特丹大学) Catholic University of Leuven(鲁汶天主大学) Toyota Research Institute(丰田研究院) Toyota Motor Europe(丰田欧洲公司)

专题命中 模仿学习与强化学习 :robotics(abstract);robotic(abstract);分类 cs.RO、cs.AI、cs.LG

AI总结 提出Demo2Reward方法,利用少量专家演示在测试时优化VLM奖励模型的提示指令,减少假阳性并保持真阳性,无需额外训练即可提升下游策略学习。

详情

展开后加载摘要…

URL PDF HTML 收藏