arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 4116 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 模仿学习与强化学习 4116 篇

1710.06034 2018-03-30 cs.LG stat.ML 57%

Stochastic Variance Reduction for Policy Gradient Estimation

Tianbing Xu, Qiang Liu, Jian Peng

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.LG

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.10467 2018-02-01 cs.AI cs.PL cs.SE 57%

Deep Reinforcement Learning for Programming Language Correction

Rahul Gupta, Aditya Kanade, Shirish Shevade

专题命中 模仿学习与强化学习 :navigation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.09344 2017-12-29 cs.AI 57%

Whatever Does Not Kill Deep Reinforcement Learning, Makes It Stronger

Vahid Behzadan, Arslan Munir

专题命中 模仿学习与强化学习 :manipulation(abstract);分类 cs.AI

Comments arXiv admin note: text overlap with arXiv:1701.04143

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.07887 2017-12-22 cs.MA cs.AI 57%

Multiagent-based Participatory Urban Simulation through Inverse Reinforcement Learning

Soma Suzuki

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.10173 2017-12-01 cs.LG stat.ML 57%

Hierarchical Policy Search via Return-Weighted Density Estimation

Takayuki Osa, Masashi Sugiyama

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.LG

Comments The 32nd AAAI Conference on Artificial Intelligence (AAAI 2018), 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.07597 2017-09-25 cs.AI 57%

Inverse Reinforcement Learning with Conditional Choice Probabilities

Mohit Sharma, Kris M. Kitani, Joachim Groeger

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.06347 2017-08-29 cs.LG 57%

Proximal Policy Optimization Algorithms

John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, Oleg Klimov

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.04818 2017-07-18 cs.CV 57%

RED: Reinforced Encoder-Decoder Networks for Action Anticipation

Jiyang Gao, Zhenheng Yang, Ram Nevatia

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.01077 2017-06-06 cs.AI 57%

Actor-Critic for Linearly-Solvable Continuous MDP with Partially Known Dynamics

Tomoki Nishi, Prashant Doshi, Michael R. James, Danil Prokhorov

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.AI

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.02239 2017-05-17 cs.AI 57%

Functions that Emerge through End-to-End Reinforcement Learning - The Direction for Artificial General Intelligence -

Katsunari Shibata

专题命中 模仿学习与强化学习 :robot learning(abstract);分类 cs.AI

Comments The Multi-disciplinary Conference on Reinforcement Learning and Decision Making (RLDM) 2017, 5 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1502.05477 2017-04-24 cs.LG 57%

Trust Region Policy Optimization

John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, Pieter Abbeel

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.LG

Comments 16 pages, ICML 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.01030 2017-03-06 cs.LG 57%

Deeply AggreVaTeD: Differentiable Imitation Learning for Sequential Prediction

Wen Sun, Arun Venkatraman, Geoffrey J. Gordon, Byron Boots, J. Andrew Bagnell

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.LG

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.01986 2016-10-07 cs.LG 57%

Active exploration in parameterized reinforcement learning

Mehdi Khamassi, Costas Tzafestas

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.LG

Comments Submitted to EWRL2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.04156 2016-09-28 stat.ML cs.LG q-bio.NC 57%

Neuroprosthetic decoder training as imitation learning

Josh Merel, David Carlson, Liam Paninski, John P. Cunningham

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.09152 2016-08-23 cs.LG 57%

Actor-critic versus direct policy search: a comparison based on sample complexity

Arnaud de Froissard de Broissia, Olivier Sigaud

专题命中 模仿学习与强化学习 :robot learning(abstract);分类 cs.LG

Comments Proceedings JFPDA (Journees Francaises Planification Decision Apprentissage)

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.04996 2016-08-18 cs.AI 57%

Open Problem: Approximate Planning of POMDPs in the class of Memoryless Policies

Kamyar Azizzadenesheli, Alessandro Lazaric, Animashree Anandkumar

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:1602.07764

Journal ref 29th Annual Conference on Learning Theory (2016) 1639--1642

详情

展开后加载摘要…

URL PDF HTML 收藏
1509.08455 2015-09-29 stat.ML cs.LG 57%

Efficient Empowerment

Maximilian Karl, Justin Bayer, Patrick van der Smagt

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1409.7552 2015-09-17 stat.ML cs.LG 57%

The Advantage of Cross Entropy over Entropy in Iterative Information Gathering

Johannes Kulick, Robert Lieck, Marc Toussaint

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.LG

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1506.02080 2015-06-09 cs.LG stat.ML 57%

Local Nonstationarity for Efficient Bayesian Optimization

Ruben Martinez-Cantin

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1410.8233 2014-10-31 cs.AI 57%

Do Artificial Reinforcement-Learning Agents Matter Morally?

Brian Tomasik

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.AI

Comments 37 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1303.2651 2014-04-01 cs.LG cs.IR 57%

Hybrid Q-Learning Applied to Ubiquitous recommender system

Djallel Bouneffouf

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.LG

Comments arXiv admin note: substantial text overlap with arXiv:1301.4351, arXiv:1303.2308

详情

展开后加载摘要…

URL PDF HTML 收藏
1303.7201 2013-03-29 cs.AI 57%

Design for a Darwinian Brain: Part 2. Cognitive Architecture

Chrisantha Fernando, Vera Vasas

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.AI

Comments Submitted as Part 2 to Living Machines 2013, Natural History Museum, London. Code available on github as it is being developed to implement the cognitive architecture above, here... https://github.com/ctf20/DarwinianNeurodynamics

详情

展开后加载摘要…

URL PDF HTML 收藏
1301.3630 2013-03-22 cs.LG 57%

Behavior Pattern Recognition using A New Representation Model

Qifeng Qiao, Peter A. Beling

专题命中 模仿学习与强化学习 :navigation(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1303.3679 2013-03-18 cs.RO 57%

Minimum-violation LTL Planning with Conflicting Specifications

Jana Tumova, Luis I. Reyes Castro, Sertac Karaman, Emilio Frazzoli, Daniela Rus

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.RO

Comments extended version of the ACC 2013 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
1208.0984 2012-08-07 cs.LG 57%

APRIL: Active Preference-learning based Reinforcement Learning

Riad Akrour, Marc Schoenauer, Michèle Sebag

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.LG

Journal ref ECML PKDD 2012 7524 (2012) 116-131

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13579 2026-07-16 cs.RO cs.AI cs.LG 新提交 56%

Agile perceptive multi-skill locomotion for quadrupedal robots in the wild

野外四足机器人的敏捷感知多技能运动

Jun-Gill Kang, Jaehyun Park, Tae-Gyu Song, Joon-Ha Kim, Seungwoo Hong, Hae-Won Park

机构 * Agency for Defense Development(国防发展局) Korea Advanced Institute of Science and Technology(韩国科学技术院) DIDEN Robotics(迪登机器人公司) Korea University(韩国大学)

专题命中 模仿学习与强化学习 :分类 cs.RO、cs.AI、cs.LG;robotics(comments)

AI总结 研究使四足机器人在复杂地形实现多技能运动的问题,提出 APT-RL 框架,利用机载感知和计算自主转换技能,通过轨迹优化生成数据集训练技能,经实验验证该框架能让机器人在复杂环境敏捷机动,稳健穿越多样障碍物。

Comments Project page: https://skillquadsr.github.io/ ,This is the author's version of the work. It is posted here by permission of the AAAS for personal use, not for redistribution. The definitive version was published in Science Robotics on 7.15.2026; doi: 10.1126/scirobotics.adz7397. Jun-Gill Kang and Jaehyun Park are co-first authors. Seungwoo Hong and Hae-Won Park are co-corresponding authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09436 2026-04-13 cs.RO cs.AI cs.LG cs.SY eess.SY 56%

Temporal Transfer Learning for Traffic Optimization with Coarse-grained Advisory Autonomy

基于粗粒度建议自主性的交通优化时序迁移学习

Jung-Hoon Cho, Sirui Li, Jeongyun Kim, Cathy Wu

机构 * Department of Civil and Environmental Engineering, Massachusetts Institute of Technology(麻省理工学院土木与环境工程系) Laboratory for Information & Decision Systems, Massachusetts Institute of Technology(麻省理工学院信息与决策系统实验室) Department of Mechanical and Automotive Engineering, Seoul National University of Science and Technology(首尔科学技术大学机械与汽车工程系)

专题命中 模仿学习与强化学习 :分类 cs.RO、cs.AI、cs.LG;robotics(journal_ref)

AI总结 本文提出时序迁移学习算法,通过零样本迁移解决自动驾驶中的粗粒度建议任务,验证了其在混合交通场景中的有效性。

Comments 18 pages, 12 figures

Journal ref IEEE Transactions on Robotics, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.07299 2026-04-07 cs.RO cs.AI cs.LG 56%

Control Synthesis from Linear Temporal Logic Specifications using Model-Free Reinforcement Learning

基于线性时序逻辑规范的控制合成使用无模型强化学习

Alper Kamil Bozkurt, Yu Wang, Michael M. Zavlanos, Miroslav Pajic

机构 * Duke University(杜克大学)

专题命中 模仿学习与强化学习 :分类 cs.RO、cs.AI、cs.LG;robotics(journal_ref)

AI总结 本文提出一种强化学习框架,用于在未知随机环境中基于线性时序逻辑规范合成控制策略,通过无模型强化学习最大化满足LTL公式概率,确保算法收敛至最优策略。

Journal ref 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 2020, pp. 10349-10355

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20230 2026-03-24 cs.RO cs.AI cs.LG 56%

Beyond Scalar Rewards: Distributional Reinforcement Learning with Preordered Objectives for Safe and Reliable Autonomous Driving

超越标量奖励:用于安全可靠自动驾驶的预序目标分布强化学习

Ahmed Abouelazm, Jonas Michel, Daniel Bogdoll, Philip Schörner, J. Marius Zöllner

机构 * FZI Research Center for Information Technology(弗劳恩霍夫研究所信息技术研究中心) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

专题命中 模仿学习与强化学习 :分类 cs.RO、cs.AI、cs.LG;robotics(comments)

AI总结 本文提出预序多目标MDP框架,通过引入量化主导指标提升自动驾驶安全性和可靠性,实验表明其在Carla中表现出更优的性能和更稳健的策略。

Comments First and Second authors contributed equally; Accepted to the 2026 IEEE International Conference on Robotics and Automation (ICRA 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05091 2026-02-06 cs.AI cs.LG cs.RO physics.space-ph 56%

Evaluating Robustness and Adaptability in Learning-Based Mission Planning for Active Debris Removal

评估基于学习的任务规划在主动清除任务中的鲁棒性和适应性

Agni Bandyopadhyay, Günther Waxenegger-Wilfing

机构 * Faculty of Mathematics and Computer Science, Julius-Maximilians-Universität Würzburg(数学与计算机科学学院,维尔堡大学)

专题命中 模仿学习与强化学习 :分类 cs.RO、cs.AI、cs.LG;robotics(comments)

AI总结 本文评估了基于学习的任务规划在主动清除任务中的鲁棒性和适应性,比较了三种规划器在不同约束条件下的性能,发现领域随机化PPO在适应性上表现更优,而MCTS在处理约束变化时更稳定但计算成本高。

Comments Presented at Conference: International Conference on Space Robotics (ISPARO,2025) At: Sendai,Japan

详情

展开后加载摘要…

URL PDF HTML 收藏