High entropy leads to symmetry-equivariant policies in Dec-POMDPs
高熵导致 Dec-POMDP 中的对称等变策略
机构 * FLAIR, Department of Engineering Science, University of Oxford(奥德赛实验室,工程科学系,牛津大学) ; Collaborative Artificial Intelligence, University of Stuttgart(协同人工智能,斯图加特大学)
AI总结 证明在 Dec-POMDP 中,足够高的熵正则化可确保策略梯度收敛到对称等变联合策略,并通过实验发现高熵系数能提升跨种子交叉对战的回报。