Decentralized Best-Response-Based Learning in Two-Player Zero-Sum Stochastic Games: A Finite-Sample Analysis
两人零和随机博弈中基于最优响应的去中心化学习:有限样本分析
机构 * Purdue University(普渡大学) ; University of Maryland, College Park(马里兰大学学院公园分校) ; Caltech(加州理工学院) ; MIT(麻省理工学院)
AI总结 本文对两人零和矩阵博弈和随机博弈中的去中心化学习进行有限样本分析,提出基于最优响应的学习算法,并证明其样本复杂度。
Comments A preliminary version [arXiv:2303.03100] of this paper, with a subset of the results that are presented here, was presented at NeurIPS 2023