arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 4113 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 模仿学习与强化学习 4113 篇

2601.21548 2026-01-30 cs.RO cs.AI cs.ET 62%

Training slow silicon neurons to control extremely fast robots with spiking reinforcement learning

训练慢硅神经元以通过脉冲强化学习控制极快的机器人

Irene Ambrosini, Ingo Blakowski, Dmitrii Zendrikov, Cristiano Capone, Luna Gava, Giacomo Indiveri, Chiara De Luca, Chiara Bartolozzi

机构 * Institute of Neuroinformatics, UZH and ETH Zurich(神经信息研究所,苏黎世联邦理工学院和苏黎世联邦理工学院) Istituto Italiano di Tecnologia(意大利技术研究所) Technical University of Munich(慕尼黑技术大学) Natl. Center for Radiation Protection and Computational Physics, Istituto Superiore di Sanità(辐射防护与计算物理国家中心,意大利卫生超级研究所) Digital Society Initiative, University of Zurich(数字社会倡议,苏黎世大学)

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.RO、cs.AI

AI总结 通过脉冲强化学习训练慢硅神经元,实现高速机器人控制,展示脑启发方法在快速互动任务中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00564 2026-01-30 cs.LG cs.AI 62%

Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining

通过联合优化的世界-动作模型扩展离线模型基于的强化学习

Jie Cheng, Ruixi Qiao, Yingwei Ma, Binhua Li, Gang Xiong, Qinghai Miao, Yongbin Li, Yisheng Lv

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Alibaba Group(阿里巴巴集团)

专题命中 模仿学习与强化学习 :world model(abstract);分类 cs.AI、cs.LG

AI总结 JOWA通过联合优化的世界-动作模型扩展离线RL,实现高效泛化和高性能

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19969 2026-01-29 cs.RO cs.LG 62%

E2HiL: Entropy-Guided Sample Selection for Efficient Real-World Human-in-the-Loop Reinforcement Learning

E2HiL: 基于熵引导的高效真实世界人机协同强化学习样本选择

Haoyuan Deng, Yuanjiang Xue, Haoyang Du, Boyang Zhou, Zhenyu Wu, Ziwei Wang

机构 * Nanyang Technological University, Singapore(南洋理工大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 模仿学习与强化学习 :manipulation(abstract);分类 cs.RO、cs.LG

AI总结 E2HiL通过熵引导的样本选择方法,提高了真实世界人机协同强化学习的样本效率和成功率,减少了人工干预需求。

Comments Project page: https://e2hil.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11234 2026-01-28 cs.LG cs.AI 62%

Bayes Adaptive Monte Carlo Tree Search for Offline Model-based Reinforcement Learning

贝叶斯自适应蒙特卡洛树搜索用于离线模型驱动强化学习

Jiayu Chen, Le Xu, Wentse Chen, Jeff Schneider

机构 * The University of Hong Kong(香港大学) Tsinghua University(清华大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 模仿学习与强化学习 :world model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出贝叶斯自适应蒙特卡洛树搜索算法,用于提升离线模型驱动强化学习的性能,显著优于现有方法。

Comments This paper is accepted in ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18751 2026-01-27 cs.LG cs.AI 62%

Trust, Don't Trust, or Flip: Robust Preference-Based Reinforcement Learning with Multi-Expert Feedback

信任、不信任或翻转:基于多专家反馈的鲁棒偏好基于强化学习

Seyed Amir Hosseini, Maryam Abdolali, Amirhosein Tavakkoli, Fardin Ayar, Ehsan Javanmardi, Manabu Tsukada, Mahdi Javanmardi

机构 * K. N. Toosi University of Technology(K. N. Toosi大学) Amirkabir University of Technology(阿美尔卡比尔技术大学) The University of Tokyo(东京大学)

专题命中 模仿学习与强化学习 :manipulation(abstract);分类 cs.AI、cs.LG

AI总结 TriTrust-PBRL通过多专家反馈联合学习奖励模型和信任参数,实现对抗性偏好下的鲁棒偏好强化学习。

Comments Equal contribution: Seyed Amir Hosseini and Maryam Abdolali. Corresponding author: Maryam Abdolali (maryam.abdolali@kntu.ac.ir)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16405 2026-01-26 cs.RO cs.LG 62%

Reinforcement Learning-Based Energy-Aware Coverage Path Planning for Precision Agriculture

基于强化学习的能源感知覆盖路径规划用于精准农业

Beining Wu, Zihao Ding, Leo Ostigaard, Jun Huang

机构 * EECS Department, South Dakota State University(南达科他州立大学电子工程与计算机科学系)

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.RO、cs.LG

AI总结 本文提出基于SAC强化学习的能源感知覆盖路径规划方法,通过CNN和LSTM优化覆盖效率与能耗,实验显示其在精准农业中覆盖效率和能耗控制优于传统算法。

Comments Accepted by RACS '25: International Conference on Research in Adaptive and Convergent Systems, November 16-19, 2025, Ho Chi Minh, Vietnam. 10 pages, 5 figures

Journal ref Proceedings of the 2025 International Conference on Research in Adaptive and Convergent Systems.(2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10492 2026-01-26 cs.LG cs.AI 62%

UACER: An Uncertainty-Adaptive Critic Ensemble Framework for Robust Adversarial Reinforcement Learning

UACER:一种用于鲁棒对抗强化学习的不确定性自适应批评者集合框架

Jiaxi Wu, Tiantian Zhang, Yuxing Wang, Yongzhe Chang, Xueqian Wang

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.AI、cs.LG

AI总结 UACER通过多样化批评者集合和时间变化衰减不确定性机制,提升鲁棒对抗强化学习的稳定性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09680 2026-01-21 cs.CV cs.AI 62%

Object-Centric Latent Action Learning

基于对象的潜在动作学习

Albina Klepach, Alexander Nikulin, Ilya Zisman, Denis Tarasov, Alexander Derevyagin, Andrei Polubarov, Nikita Lyubaykin, Igor Kiselev, Vladislav Kurenkov

专题命中 模仿学习与强化学习 :embodied AI(abstract);分类 cs.AI、cs.CV

AI总结 本文提出了一种基于对象的潜在动作学习框架,通过自监督对象中心预训练解耦代理与干扰背景动态,提升动作标签的鲁棒性,从而改善模仿学习和代理适应效率。

Comments Accepted by AAAI 2026 (Oral). Source code: https://github.com/dunnolab/object-centric-lapo

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13794 2026-01-21 cs.GR cs.LG cs.RO 62%

MimicKit: A Reinforcement Learning Framework for Motion Imitation and Control

MimicKit:一种用于运动模仿与控制的强化学习框架

Xue Bin Peng

机构 * Simon Fraser University(西蒙弗雷泽大学) NVIDIA(英伟达)

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.RO、cs.LG

AI总结 MimicKit 通过运动模仿和强化学习结合,提供统一的运动控制训练框架,支持计算机图形学和机器人学的研究与应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00549 2026-01-14 cs.LG cs.AI 62%

Robust Single-Agent Reinforcement Learning for Regional Traffic Signal Control Under Demand Fluctuations

鲁棒的单智能体强化学习用于应对需求波动的区域交通信号控制

Qiang Li, Jin Niu, Lina Yu

专题命中 模仿学习与强化学习 :world model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种鲁棒的单智能体强化学习框架,用于应对交通需求波动的区域交通信号控制,通过集中决策和高效学习模型有效减少交通队列长度。

Comments A critical error in the methodology. The reported congestion control effects were not caused by the proposed signal timing optimization, but by an incorrect traffic volume scaling factor during evaluation. The traffic demand was not properly amplified, resulting in misleading performance gains. Due to the substantial nature of the error, completion of revisions is not feasible in the short term

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19717 2025-12-24 cs.LG cs.AI 62%

Thermodynamic Focusing for Inference-Time Search: Practical Methods for Target-Conditioned Sampling and Prompted Inference

推理时间搜索的热力学聚焦:目标条件采样和提示推理的实用方法

Zhan Zhang

机构 * Zhan Zhang(张湛)

专题命中 模仿学习与强化学习 :navigation(abstract);分类 cs.AI、cs.LG

AI总结 本文提出ICFA算法,通过目标条件重新加权提升推理时间搜索效率,结合结构化提示与混合架构实现更高效的样本利用。

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17846 2025-12-22 cs.RO cs.AI 62%

Planning as Descent: Goal-Conditioned Latent Trajectory Synthesis in Learned Energy Landscapes

规划作为下降:在学习的能量景观中进行目标条件化潜在轨迹合成

Carlos Vélez García, Miguel Cazorla, Jorge Pomares

机构 * University of Alicante(阿利坎特大学)

专题命中 模仿学习与强化学习 :manipulation(abstract);分类 cs.RO、cs.AI

AI总结 PaD通过学习目标条件化能量函数实现离线无奖励规划,以验证为基础进行轨迹合成,取得95%的成功率并优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16813 2025-12-19 cs.NI cs.AI cs.DC cs.LG eess.SP 62%

Coordinated Anti-Jamming Resilience in Swarm Networks via Multi-Agent Reinforcement Learning

通过多智能体强化学习实现蜂群网络中的协同抗干扰韧性

Bahman Abolhassani, Tugba Erpek, Kemal Davaslioglu, Yalin E. Sagduyu, Sastry Kompella

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于QMIX算法的多智能体强化学习框架,用于提升蜂群网络在反应式干扰下的通信韧性,通过仿真验证其在对抗环境中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15120 2025-12-18 cs.LG cs.AI 62%

Automatic Reward Shaping from Multi-Objective Human Heuristics

多目标人类启发式自动奖励塑造

Yuqing Xie, Jiayu Chen, Wenhao Tang, Ya Zhang, Chao Yu, Yu Wang

机构 * Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.AI、cs.LG

AI总结 MORSE通过双层优化框架自动整合多目标人类启发式奖励,提升机器人任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21127 2025-12-16 cs.LG cs.AI 62%

Meta Policy Switching for Secure UAV Deconfliction in Adversarial Airspace

元策略切换用于对抗性空域中的安全无人机脱困

Deepak Kumar Panda, Weisi Guo

机构 * Faculty of Engineering and Applied Sciences, Cranfield University(工程与应用科学学院,克兰菲尔德大学)

专题命中 模仿学习与强化学习 :navigation(abstract);分类 cs.AI、cs.LG

AI总结 本研究提出元策略切换框架,通过动态选择稳健策略提升无人机在对抗性空域中的安全导航能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09410 2025-12-11 cs.RO cs.LG cs.MA 62%

Generalizable Collaborative Search-and-Capture in Cluttered Environments via Path-Guided MAPPO and Directional Frontier Allocation

在 cluttered 环境中通过路径引导的 MAPPO 和方向性前沿分配实现可泛化的协作搜索与捕获

Jialin Ying, Zhihao Li, Zicheng Dong, Guohua Wu, Yihuan Liao

机构 * Department of Automation, Central South University, Changsha, China(自动化系,中南大学,长沙,中国)

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.RO、cs.LG

AI总结 本文提出 PGF-MAPPO 方法,通过路径引导的 MAPPO 和方向性前沿分配,在 cluttered 环境中实现高效的协作搜索与捕获,展示了出色的泛化能力和性能优势。

Comments 7 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.10585 2025-12-11 cs.RO cs.AI 62%

Prediction-aware and Reinforcement Learning based Altruistic Cooperative Driving

具有预测意识和强化学习的利他性协作驾驶

Rodolfo Valiente, Mahdi Razzaghpour, Behrad Toghi, Ghayoor Shah, Yaser P. Fallah

专题命中 模仿学习与强化学习 :navigation(abstract);分类 cs.RO、cs.AI

AI总结 本文提出一种基于预测意识和强化学习的自动驾驶方法,通过混合预测网络和社交优化框架提升安全性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14259 2025-12-09 cs.LG cs.RO 62%

Quantization-Free Autoregressive Action Transformer

无量化自回归动作变压器

Ziyad Sheebaelhamd, Michael Tschannen, Michael Muehlebach, Claire Vernade

机构 * University of Tübingen(图宾根大学) Google DeepMind(谷歌DeepMind) Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.RO、cs.LG

AI总结 无量化自回归动作变压器通过生成无限词汇变换器直接参数化连续策略,简化流程并提升在模拟机器人任务中的性能。

Journal ref 39th Conference on Neural Information Processing Systems, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15582 2025-12-08 cs.RO cs.AI 62%

Momentum-constrained Hybrid Heuristic Trajectory Optimization Framework with Residual-enhanced DRL for Visually Impaired Scenarios

具有残差增强深度强化学习的动量约束混合启发式轨迹优化框架用于视障场景

Yuting Zeng, Zhiwen Zheng, You Zhou, JiaLing Xiao, Yongbin Yu, Manping Fan, Bo Gong, Liyong Ren

专题命中 模仿学习与强化学习 :navigation(abstract);分类 cs.RO、cs.AI

AI总结 本文提出一种结合残差增强深度强化学习的动量约束混合启发式轨迹优化框架,用于提升视障场景下的导航鲁棒性与安全性。

Comments Upon further internal evaluation, we found that the current version does not adequately represent the clarity and completeness that we intend for this work. To avoid possible misunderstanding caused by this preliminary form, we request withdrawal. A refined version will be prepared privately before any further dissemination

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19785 2025-12-03 cs.LG cs.AI 62%

medDreamer: Model-Based Reinforcement Learning with Latent Imagination on Complex EHRs for Clinical Decision Support

medDreamer:基于复杂电子病历的潜在想象的模型驱动强化学习用于临床决策支持

Qianyi Xu, Gousia Habib, Feng Wu, Dilruk Perera, Mengling Feng

机构 * National University of Singapore(新加坡国立大学)

专题命中 模仿学习与强化学习 :world model(abstract);分类 cs.AI、cs.LG

AI总结 medDreamer通过结合潜在想象和自适应特征集成模块,实现基于复杂电子病历的模型驱动强化学习,以提升个性化治疗推荐的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06981 2025-12-02 cs.AI cs.LG 62%

Deep RL Needs Deep Behavior Analysis: Exploring Implicit Planning by Model-Free Agents in Open-Ended Environments

深度强化学习需要深度行为分析:通过无模型智能体在开放性环境中探索隐式规划

Riley Simmons-Edler, Ryan P. Badman, Felix Baastad Berg, Raymond Chua, John J. Vastola, Joshua Lunger, William Qian, Kanaka Rajan

机构 * Department of Neurobiology, Harvard Medical School(哈佛医学院神经生物学系) Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(哈佛大学自然与人工智能研究学院) Department of Mathematics, NTNU(NTNU数学系) School of Computer Science, McGill University & Mila(麦吉尔大学计算机科学学院及Mila) Department of Computer Science, University of Toronto(多伦多大学计算机科学系) Biophysics Graduate Program, Harvard University(哈佛大学生物物理学研究生项目)

专题命中 模仿学习与强化学习 :world model(abstract);分类 cs.AI、cs.LG

AI总结 本文通过ForageWorld环境研究DRL智能体的行为,发现无模型智能体可通过涌现动态展现规划行为,提出通用分析框架用于研究复杂智能体的学习动态。

Comments Published at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03187 2025-12-01 cs.LG cs.RO 62%

Periodic Skill Discovery

周期性技能发现

Jonghae Park, Daesol Cho, Jusuk Lee, Dongseok Shim, Inkyu Jang, H. Jin Kim

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.RO、cs.LG

AI总结 周期性技能发现(PSD)通过无监督学习发现具有不同周期的技能,提升复杂机器人任务中的表现和多样性。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18617 2025-11-26 cs.RO cs.CV 62%

AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations

AutoFocus-IL:基于视觉语言模型的数据高效视觉模仿学习中的显著性图

Litian Gong, Fatemeh Bahrani, Yutai Zhou, Amin Banayeeanzade, Jiachen Li, Erdem Bıyık

机构 * Department of Electrical and Computer Engineering, University of California, Riverside, USA(电气与计算机工程系,加州大学河滨分校) Thomas Lord Department of Computer Science, University of Southern California, USA(汤姆斯·劳德计算机科学系,南加州大学)

专题命中 模仿学习与强化学习 :manipulation(abstract);分类 cs.RO、cs.CV

AI总结 AutoFocus-IL通过视觉语言模型自动生成显著性图,提升视觉模仿学习的数据效率和泛化能力,无需额外人类标注。

Comments 8 pages, 6 figures. Code and datasets available at http://autofocus-il.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19528 2025-11-26 cs.RO cs.AI 62%

Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories

发现、学习与强化:通过多样化强化学习生成轨迹扩展视觉-语言-动作预训练

Rushuai Yang, Zhiyuan Feng, Tianxiang Zhang, Kaixin Wang, Chuheng Zhang, Li Zhao, Xiu Su, Yi Chen, Jiang Bian

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Tsinghua University(清华大学) Wuhan University(武汉大学) Central South University(中南大学) Microsoft Research(微软研究院)

专题命中 模仿学习与强化学习 :manipulation(abstract);分类 cs.RO、cs.AI

AI总结 本文提出DLR框架,通过多样化强化学习生成轨迹,提升VLA预训练的多样性和扩展性,实现更广泛的状态-动作空间覆盖。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18606 2025-11-25 cs.RO cs.LG 62%

How to Train Your Latent Control Barrier Function: Smooth Safety Filtering Under Hard-to-Model Constraints

如何训练您的潜在控制障碍函数:在难以建模的约束下实现平滑的安全过滤

Kensuke Nakamura, Arun L. Bishop, Steven Man, Aaron M. Johnson, Zachary Manchester, Andrea Bajcsy

专题命中 模仿学习与强化学习 :manipulation(abstract);分类 cs.RO、cs.LG

AI总结 本文提出LatentCBF方法,通过梯度惩罚和价值训练解决潜在空间中不兼容价值函数的问题,实现平滑安全过滤并提升任务完成率。

Comments 3 figures, 10 tables, 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15284 2025-11-20 cs.RO cs.AI 62%

Path Planning through Multi-Agent Reinforcement Learning in Dynamic Environments

Jonas De Maeyer, Hossein Yarahmadi, Moharram Challenger

机构 * Department of Computer Science University of Antwerp (UA)(安特卫普大学计算机科学系) Department of Computer Engineering, Faculty of Engineering, Ayatollah Boroujerdi University(阿亚图拉·博鲁杰尔迪大学工程学院计算机工程系) Department of Computer Science University of Antwerp (UA) and Flanders Make(安特卫普大学计算机科学系和弗拉芒制造)

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.RO、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14153 2025-11-19 cs.CV cs.AI 62%

LENS: Learning to Segment Anything with Unified Reinforced Reasoning

Lianghui Zhu, Bin Ouyang, Yuxuan Zhang, Tianheng Cheng, Rui Hu, Haocheng Shen, Longjin Ran, Xiaoxin Chen, Li Yu, Wenyu Liu, Xinggang Wang

机构 * vivo Mobile Communication Co., Ltd.(vivo移动通信有限公司)

专题命中 模仿学习与强化学习 :robotics(abstract);分类 cs.AI、cs.CV

Comments Code is released at https://github.com/hustvl/LENS

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11992 2025-11-18 cs.MA cs.AI cs.LG 62%

Goal-Oriented Multi-Agent Reinforcement Learning for Decentralized Agent Teams

Hung Du, Hy Nguyen, Srikanth Thudumu, Rajesh Vasa, Kon Mouzakis

专题命中 模仿学习与强化学习 :navigation(abstract);分类 cs.AI、cs.LG

Comments Accepted poster at the IEEE Consumer Communications & Networking Conference (CCNC) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.09010 2025-11-17 cs.RO cs.LG cs.SY eess.SY 62%

DiAReL: Reinforcement Learning with Disturbance Awareness for Robust Sim2Real Policy Transfer in Robot Control

Mohammadhossein Malmir, Josip Josifovski, Noah Klarmann, Alois Knoll

机构 * Department of Computer Engineering, School of Computation, Information and Technology, Technical University of Munich(计算机工程系,计算、信息与技术学院,慕尼黑技术大学) Rosenheim University of Applied Sciences(罗森海姆应用技术大学)

专题命中 模仿学习与强化学习 :robotic(abstract);分类 cs.RO、cs.LG

Comments Accepted for publication in IEEE Transactions on Control Systems Technology (TCST)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09852 2025-11-13 cs.RO cs.LG cs.MA 62%

Strategic Coordination of Drones via Short-term Distributed Optimization and Long-term Reinforcement Learning

Chuhao Qin, Evangelos Pournaras

机构 * School of Computer Science, University of Leeds(利兹大学计算机科学学院)

专题命中 模仿学习与强化学习 :navigation(abstract);分类 cs.RO、cs.LG

Comments 23 pages, 16 figures, accepted by Applied Soft Computing

详情

展开后加载摘要…

URL PDF HTML 收藏