The Harder Path: Last Iterate Convergence for Uncoupled Learning in Zero-Sum Games with Bandit Feedback
更艰难的道路:零和博弈中无耦合学习的最后迭代收敛
机构 * Inria - FairPlay(Inria-公平游戏实验室) ; Stealth AI Startup / Inria / ENS(隐形AI初创公司 / Inria / 索邦大学) ; Criteo AI Lab, Paris, France(Criteo AI实验室,巴黎,法国)
AI总结 研究零和矩阵博弈在重复玩和带隙反馈下的学习问题,提出无耦合算法保证在无通信情况下最后迭代收敛到纳什均衡,发现收敛至纳什均衡对性能不利,最佳速率为Ω(T^{-1/4}),优于常规的Ω(T^{-1/2})。
Comments Accepted at the 42nd International Conference on Machine Learning (ICML 2025)