arXivDaily arXiv每日学术速递 周一至周五更新

作者

Richard S. Sutton

Reinforcement Learning

共收录 77
2608.01475 2026-08-04 cs.LG 新提交

Plasticity of Growing and Elastic Neural Networks in Online Continual Learning

在线持续学习中生长型与弹性型神经网络的可塑性

Jeong Min Kong, Richard S. Sutton

机构 * University of California, Los Angeles (UCLA)(加利福尼亚大学洛杉矶分校) University of Alberta(阿尔伯塔大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所)

AI总结 本文研究在线持续学习中生长型与弹性型神经网络的可塑性,实验显示自适应生长型和弹性型神经网络可维持高准确率、不丧失可塑性,或为在线持续学习提供有前景的算法类别。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19357 2026-06-19 cs.RO cs.AI 新提交

Physical Atari: A Robust and Accessible Platform for Real-time Reinforcement Learning on Robots

Physical Atari: 一个用于机器人实时强化学习的鲁棒且可访问的平台

Khurram Javed, Joseph Modayil, Gloria Kennickell, Richard S. Sutton, John Carmack

机构 * Keen Technologies University of Alberta, Canada(阿尔伯塔大学,加拿大) Openmind Research Institute(Openmind研究机构)

AI总结 提出Physical Atari平台,通过机器人操作Atari控制器和实时渲染游戏帧,实现物理世界中的强化学习研究,验证了算法可直接在机器人上学习,并指出分布偏移会显著降低策略性能。

Comments To appear at RLC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
1206.3285 2026-06-03 cs.AI cs.LG cs.SY eess.SY

Dyna-Style Planning with Linear Function Approximation and Prioritized Sweeping

具有线性函数逼近和优先级扫描的Dyna风格规划

Richard S. Sutton, Csaba Szepesvari, Alborz Geramifard, Michael P. Bowling

AI总结 本文提出一种基于模型的Dyna风格规划方法,扩展至线性函数逼近,证明其收敛性,并引入线性Dyna的优先级扫描算法。

Comments Appears in Proceedings of the Twenty-Fourth Conference on Uncertainty in Artificial Intelligence (UAI2008)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03915 2026-06-02 cs.LG math.OC

Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning

异步随机逼近及其在平均奖励强化学习中的应用

Huizhen Yu, Yi Wan, Richard S. Sutton

机构 * Department of Computing Science, University of Alberta(计算科学系,阿尔伯塔大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔人工智能研究所(Amii))

AI总结 研究异步随机逼近算法的稳定性与收敛性,通过扩展Borkar-Meyn稳定性证明方法和Hirsch-Benaïm动力学系统方法,为平均奖励强化学习中的相对值迭代算法提供理论基础。

Comments 34 pages. This version contains only the asynchronous stochastic approximation material from version 2 of the original report; the reinforcement-learning material has been moved to a separate, stand-alone paper (arXiv:2512.06218). Minor corrections and additional remarks have been incorporated. A shorter version of this paper is to appear in the SIAM Journal on Control and Optimization

Journal ref SIAM Journal on Control and Optimization, 64(3):1456-1481, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24238 2026-05-26 cs.AI

Toward Enactive Artificial Intelligence

走向生成式人工智能

Banafsheh Rafiee, Richard Sutton

机构 * Independent Researcher(独立研究者) Department of Computing Science, University of Alberta, Canada(阿尔伯塔大学计算机科学系) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所)

AI总结 本文主张将生成式认知方法融入人工智能,强调感知与行动不可分割、具身性和自主性,并指出强化学习在结构上与生成式原则存在共鸣但仍有差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19033 2026-04-22 cs.LG cs.AI

Intentional Updates for Streaming Reinforcement Learning

为流式强化学习设计的意图更新

Arsalan Sharifnassab, Mohamed Elsayed, Kris De Asis, A. Rupam Mahmood, Richard S. Sutton

机构 * Openmind Research Institute(Openmind研究 institutes) Department of Computing Science, University of Alberta(阿尔伯塔大学计算机科学系)

AI总结 本文提出意图更新方法,通过定义更新目标来优化流式强化学习的稳定性与性能,实验表明其在流式环境下表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06218 2025-12-09 cs.LG math.OC

Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration

在半马尔可夫决策过程中的平均奖励强化学习中通过相对价值迭代

Huizhen Yu, Yi Wan, Richard S. Sutton

机构 * Department of Computing Science, University of Alberta(计算科学系,阿尔伯塔大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所)

AI总结 本文提出了一种在半马尔可夫决策过程中利用相对价值迭代算法进行平均奖励强化学习的方法,并引入新的单调性条件以提高算法收敛性。

Comments 24 pages. This paper presents the reinforcement-learning material previously contained in version 2 of arXiv:2409.03915, which is now being split into two stand-alone papers. Minor corrections and improvements to the main results have also been made in the course of this reformatting

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19539 2025-07-29 cs.LG cs.AI stat.ML

Swift-Sarsa: Fast and Robust Linear Control

Khurram Javed, Richard S. Sutton

机构 * University of Alberta(阿尔伯塔大学)

Comments Presented at RLDM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02342 2025-07-10 cs.LG cs.AI math.OC

MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters

Arsalan Sharifnassab, Saber Salehkaleybar, Richard Sutton

机构 * Openmind Research Institute, Canada(开放心智研究机构,加拿大) Leiden Institute of Advanced Computer Science, Leiden University, Leiden, Netherlands(莱顿先进计算机科学研究所,莱顿大学,莱顿,荷兰) Department of Computing Science, University of Alberta, Edmonton, Canada(计算科学系,阿尔伯塔大学,爱德蒙顿,加拿大)

Journal ref ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09999 2024-10-31 cs.LG cs.AI

Reward Centering

Abhishek Naik, Yi Wan, Manan Tomar, Richard S. Sutton

Comments In Proceedings of RLC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14951 2024-09-04 cs.LG cs.AI

An Idiosyncrasy of Time-discretization in Reinforcement Learning

Kris De Asis, Richard S. Sutton

Comments RLC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16262 2024-08-30 cs.LG math.OC

On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes

Yi Wan, Huizhen Yu, Richard S. Sutton

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15091 2024-08-15 cs.LG math.OC

A Note on Stability in Asynchronous Stochastic Approximation without Communication Delays

Huizhen Yu, Yi Wan, Richard S. Sutton

Comments Corrected typos and a minor error; parts of this material will be included in a separate future arXiv preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14361 2024-07-23 cs.LG cs.AI

Auxiliary task discovery through generate-and-test

Banafsheh Rafiee, Sina Ghiassian, Jun Jin, Richard Sutton, Jun Luo, Adam White

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13812 2024-04-11 cs.LG

Maintaining Plasticity in Deep Continual Learning

Shibhansh Dohare, J. Fernando Hernandez-Garcia, Parash Rahman, A. Rupam Mahmood, Richard S. Sutton

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17401 2024-02-01 cs.LG cs.AI

Step-size Optimization for Continual Learning

Thomas Degris, Khurram Javed, Arsalan Sharifnassab, Yuxin Liu, Richard Sutton

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01569 2023-12-27 cs.AI cs.LG

Iterative Option Discovery for Planning, by Planning

Kenny Young, Richard S. Sutton

Comments Fixed incorrect arrows on some figures in the appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.03466 2023-09-19 cs.LG cs.AI

Reward-Respecting Subtasks for Model-Based Reinforcement Learning

Richard S. Sutton, Marlos C. Machado, G. Zacharias Holland, David Szepesvari, Finbarr Timbers, Brian Tanner, Adam White

Journal ref Artificial Intelligence, first published online September 6, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.13757 2023-07-25 cs.LG cs.AI

Toward Efficient Gradient-Based Value Estimation

Arsalan Sharifnassab, Richard Sutton

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15625 2023-06-28 cs.LG cs.AI

Value-aware Importance Weighting for Off-policy Reinforcement Learning

Kristopher De Asis, Eric Graves, Richard S. Sutton

Comments CoLLAs 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11173 2023-03-23 cs.AI cs.LG

The Alberta Plan for AI Research

Richard S. Sutton, Michael Bowling, Patrick M. Pilarski

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01613 2022-11-29 cs.LG

Doubly-Asynchronous Value Iteration: Making Value Iteration Asynchronous in Actions

Tian Tian, Kenny Young, Richard S. Sutton

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.15141 2022-11-08 cs.LG

On Convergence of Average-Reward Off-Policy Control Algorithms in Weakly Communicating MDPs

Yi Wan, Richard S. Sutton

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.02808 2022-10-19 cs.LG cs.AI

Average-Reward Off-Policy Policy Evaluation with Function Approximation

Shangtong Zhang, Yi Wan, Richard S. Sutton, Shimon Whiteson

Comments ICML 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.04590 2022-10-12 cs.AI

From Eye-blinks to State Construction: Diagnostic Benchmarks for Online Representation Learning

Banafsheh Rafiee, Zaheer Abbas, Sina Ghiassian, Raksha Kumaraswamy, Richard Sutton, Elliot Ludvig, Adam White

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12515 2022-10-03 cs.LG cs.AI

Toward Discovering Options that Achieve Faster Planning

Yi Wan, Richard S. Sutton

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13252 2022-06-07 cs.AI

The Quest for a Common Model of the Intelligent Decision Maker

Richard S. Sutton

Comments Will appear as an extended abstract at the fifth Multi-disciplinary Conference on Reinforcement Learning and Decision Making, held in Providence, Rhode Island, June 8-11, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06325 2022-05-06 cs.LG

Continual Backprop: Stochastic Gradient Descent with Persistent Randomness

Shibhansh Dohare, Richard S. Sutton, A. Rupam Mahmood

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.09701 2022-02-22 cs.LG

A History of Meta-gradient: Gradient Methods for Meta-learning

Richard S. Sutton

Comments 3 pages of text, 54 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.15236 2022-01-03 cs.LG cs.AI

Learning Agent State Online with Recurrent Generate-and-Test

Amir Samani, Richard S. Sutton

详情

展开后加载摘要…

URL PDF HTML 收藏