Safe Reinforcement Learning with Preference-based Constraint Inference
基于偏好的约束推断的安全强化学习
机构 * Department of Automation, Tsinghua University, Beijing, China ; Laboratory for Information \& Decision Systems, Massachusetts Institute of Technology, Cambridge, MA, USA
专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG
AI总结 提出偏好约束强化学习(PbCRL),通过引入死区机制和信噪比损失,从人类偏好中推断安全约束,实现更好的约束对齐和策略学习。
Comments Accepted by the 43rd International Conference on Machine Learning (ICML 2026)