Variance-Reduced Q-Learning over Static and Time-Varying Networks
静态和时变网络上的方差减少Q学习
机构 * Department of Electrical and Computer Engineering, North Carolina State University(北卡罗来纳州立大学电气与计算机工程系) ; Department of Electrical and Computer Engineering, University of California San Diego(加利福尼亚大学圣地亚哥分校电气与计算机工程系)
AI总结 研究多个智能体在同一MDP下的分散强化学习问题,提出基于轮次的分布式Q学习算法VRDQ,该算法在静态和时变网络中能实现高概率有限时间收敛,样本复杂度加速且通信量仅需\(\tilde{O}(1)\),改善了通信成本。
Comments Accepted at the 2026 American Control Conference (ACC 2026)