Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems
长期决策问题中基于成对偏好的强化学习
机构 * School of Computer Science, McGill University, Montreal, Quebec, Canada(麦吉尔大学计算机科学学院) ; Mila - Quebec AI Institute, Montreal, Quebec, Canada(魁北克人工智能研究所) ; Department of Electrical Engineering, Stanford University, Stanford, California, USA(斯坦福大学电气工程系)
AI总结 针对长期决策问题中基于成对偏好的强化学习效率低且缺乏马尔可夫策略最优性保证的问题,提出马尔可夫决策竞赛模型,证明平稳马尔可夫策略最优性、求解复杂度为P,并给出亚线性收敛算法,在高维长期问题中显著提升学习效率。
Comments Accepted for ICML 2026. v2 has an updated abstract and introduction. Results and conclusions are unchanged