作者
Richard S. Sutton
Reinforcement Learning
An Empirical Comparison of Off-policy Prediction Learning Algorithms in the Four Rooms Environment
Comments 13 pages
Learning and Planning in Average-Reward Markov Decision Processes
Comments In Proceedings of ICML 2021
An Empirical Comparison of Off-policy Prediction Learning Algorithms on the Collision Task
Does the Adam Optimizer Exacerbate Catastrophic Forgetting?
Comments 9 pages in main text + 3 pages of references + 16 pages of appendices, 6 figures in main text + 21 figures in appendices, 6 tables in appendices; source code available at https://github.com/dylanashley/catastrophic-forgetting/tree/arxiv
Planning with Expectation Models for Control
Policy Iterations for Reinforcement Learning Problems in Continuous Time and Space -- Fundamental Theory and Methods
Comments To appear in Automatica. All the Appendices are provided
Journal ref Automatica vol. 126, 109421 (2021)
Understanding the Pathologies of Approximate Policy Evaluation when Combined with Greedification in Reinforcement Learning
Document-editing Assistants and Model-based Reinforcement Learning as a Path to Conversational AI
Comments Currently under review
Inverse Policy Evaluation for Value-based Sequential Decision-making
Comments Submitted to NeurIPS 2020
Planning with Expectation Models
Behaviour Suite for Reinforcement Learning
Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning
Comments AAAI 2020
Learning Sparse Representations Incrementally in Deep Reinforcement Learning
Discounted Reinforcement Learning Is Not an Optimization Problem
Comments Accepted for presentation at the Optimization Foundations of Reinforcement Learning Workshop at NeurIPS 2019
Learning Feature Relevance Through Step Size Adaptation in Temporal-Difference Learning
Should All Temporal Difference Learning Use Emphasis?
Understanding Multi-Step Deep Reinforcement Learning: A Systematic Study of the DQN Target
On Generalized Bellman Equations and Temporal-Difference Learning
Comments Minor revision; 41 pages; to appear in Journal on Machine Learning Research, 2018
Journal ref Journal of Machine Learning Research 19(48):1-49, 2018
Online Off-policy Prediction
Comments 68 pages
Predicting Periodicity with Temporal Difference Learning
Per-decision Multi-step Temporal Difference Learning with Control Variates
Journal ref (2018). In Conference on Uncertainty in Artificial Intelligence. http://auai.org/uai2018/proceedings/papers/282.pdf
Two geometric input transformation methods for fast online reinforcement learning with neural nets
Comments 16 pages
Reactive Reinforcement Learning in Asynchronous Environments
Comments 11 pages, 7 figures, currently under journal peer review
Multi-step Reinforcement Learning: A Unifying Algorithm
Comments Appeared at the Thirty-Second AAAI Conference on Artificial Intelligence (AAAI-18)
Journal ref (2018). In AAAI Conference on Artificial Intelligence. https://www.aaai.org/ocs/index.php/AAAI/AAAI18/paper/view/16294
Integrating Episodic Memory into a Reinforcement Learning Agent using Reservoir Sampling
A Deeper Look at Experience Replay
Comments NIPS 2017 Deep Reinforcement Learning Symposium
TIDBD: Adapting Temporal-difference Step-sizes Through Stochastic Meta-descent
Comments Version as submitted to the 31st Conference on Neural Information Processing Systems (NIPS 2017) on May 19, 2017. 9 pages, 5 figures. Extended version in preparation for journal submission
Directly Estimating the Variance of the λ-Return Using Temporal-Difference Methods
Communicative Capital for Prosthetic Agents
Comments 33 pages, 10 figures; unpublished technical report undergoing peer review