arXivDaily arXiv每日学术速递 周一至周五更新

作者

Richard S. Sutton

Reinforcement Learning

共收录 77
1705.04185 2017-05-15 cs.AI cs.LG

A First Empirical Study of Emphatic Temporal Difference Learning

Sina Ghiassian, Banafsheh Rafiee, Richard S. Sutton

Comments 5 pages, Accepted to NIPS Continual Learning and Deep Networks workshop, 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.03967 2017-05-12 cs.LG

GQ($λ$) Quick Reference and Implementation Guide

Adam White, Richard S. Sutton

详情

展开后加载摘要…

URL PDF HTML 收藏
1612.02879 2017-04-28 cs.LG cs.AI stat.ML

Learning Representations by Stochastic Meta-Gradient Descent in Neural Networks

Vivek Veeriah, Shangtong Zhang, Richard S. Sutton

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.03006 2017-02-13 cs.LG

Multi-step Off-policy Learning Without Importance Sampling Ratios

Ashique Rupam Mahmood, Huizhen Yu, Richard S. Sutton

Comments 24 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1512.04087 2016-09-09 cs.AI cs.LG

True Online Temporal-Difference Learning

Harm van Seijen, A. Rupam Mahmood, Patrick M. Pilarski, Marlos C. Machado, Richard S. Sutton

Comments This is the published JMLR version. It is a much improved version. The main changes are: 1) re-structuring of the article; 2) additional analysis on the forward view; 3) empirical comparison of traditional and new forward view; 4) added discussion of other true online papers; 5) updated discussion for non-linear function approximation

Journal ref Journal of Machine Learning Research (JMLR), 17(145):1-40, 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1503.04269 2016-07-21 cs.LG

An Emphatic Approach to the Problem of Off-policy Temporal-Difference Learning

Richard S. Sutton, A. Rupam Mahmood, Martha White

Comments 29 pages This is a significant revision based on the first set of reviews. The most important change was to signal early that the main result is about stability, not convergence

Journal ref Journal of Machine Learning Research 17(73): 1-29, 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.02807 2016-06-10 cs.HC cs.AI

Face valuing: Training user interfaces with facial expressions and reinforcement learning

Vivek Veeriah, Patrick M. Pilarski, Richard S. Sutton

Comments 7 pages, 4 figures, IJCAI 2016 - Interactive Machine Learning Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
1508.04582 2015-08-20 cs.LG

Learning to Predict Independent of Span

Hado van Hasselt, Richard S. Sutton

Comments 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1507.07147 2015-07-28 cs.LG

True Online Emphatic TD($λ$): Quick Reference and Implementation Guide

Richard S. Sutton

详情

展开后加载摘要…

URL PDF HTML 收藏
1507.01569 2015-07-07 cs.LG cs.AI

Emphatic Temporal-Difference Learning

A. Rupam Mahmood, Huizhen Yu, Martha White, Richard S. Sutton

Comments 9 pages, accepted for presentation at European Workshop on Reinforcement Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
1507.00353 2015-07-03 cs.AI cs.LG stat.ML

An Empirical Evaluation of True Online TD(λ)

Harm van Seijen, A. Rupam Mahmood, Patrick M. Pilarski, Richard S. Sutton

Comments European Workshop on Reinforcement Learning (EWRL) 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1504.05539 2015-04-22 cs.LG

Temporal-Difference Networks

Richard S. Sutton, Brian Tanner

Comments 8 pages, 3 figures, presented at the 2004 conference on Neural Information Processing Systems. in Advances in Neural Information Processing Systems 17 (proceedings of the 2004 conference), Saul, L. K., Weiss, Y., and Bottou, L. (Eds)

详情

展开后加载摘要…

URL PDF HTML 收藏
1205.4839 2015-03-19 cs.LG

Off-Policy Actor-Critic

Thomas Degris, Martha White, Richard S. Sutton

Comments Full version of the paper, appendix and errata included; Proceedings of the 2012 International Conference on Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
1112.1133 2015-03-19 cs.LG cs.RO

Multi-timescale Nexting in a Reinforcement Learning Robot

Joseph Modayil, Adam White, Richard S. Sutton

Comments (11 pages, 5 figures, This version to appear in the Proceedings of the Conference on the Simulation of Adaptive Behavior, 2012)

详情

展开后加载摘要…

URL PDF HTML 收藏
1309.4714 2013-09-19 cs.AI cs.LG cs.RO

Temporal-Difference Learning to Assist Human Decision Making during the Control of an Artificial Limb

Ann L. Edwards, Alexandra Kearney, Michael Rory Dawson, Richard S. Sutton, Patrick M. Pilarski

Comments 5 pages, 4 figures, This version to appear at The 1st Multidisciplinary Conference on Reinforcement Learning and Decision Making, Princeton, NJ, USA, Oct. 25-27, 2013

详情

展开后加载摘要…

URL PDF HTML 收藏
1301.2343 2013-01-14 cs.AI cs.LG

Planning by Prioritized Sweeping with Small Backups

Harm van Seijen, Richard S. Sutton

详情

展开后加载摘要…

URL PDF HTML 收藏
1206.6262 2012-06-28 cs.AI cs.LG

Scaling Life-long Off-policy Learning

Adam White, Joseph Modayil, Richard S. Sutton

详情

展开后加载摘要…

URL PDF HTML 收藏