Global Optimality for Constrained Exploration via Penalty Regularization
通过惩罚正则化实现约束探索的全局最优性
机构 * Florian Wolf: , Ilyas Fatkhullin: , Niao He: 1The Computing \& Mathematical Sciences Department, California Institute of Technology, Pasadena, CA. 2Department of Computer Science, ETH Zurich, Switzerland. 3ETH AI Center, ETH Zurich, Switzerland.
AI总结 本文提出Policy Gradient Penalty方法,通过二次惩罚正则化解决约束下的探索问题,实现全局收敛性和近优策略。