CSPO: Constraint-Sensitive Policy Optimization for Safe Reinforcement Learning
CSPO: 面向安全强化学习的约束敏感策略优化
机构 * University of Luxembourg(卢森堡大学)
专题命中 安全训练 :safety(abstract);分类 cs.AI
AI总结 提出约束敏感策略优化(CSPO),通过引入局部约束敏感性修正原目标,加速安全恢复并减少振荡,在导航与运动基准上取得更高约束回报。
Comments Accepted as a Spotlight paper at the 43rd International Conference on Machine Learning (ICML 2026)