Generalization Limits of Reinforcement Learning Alignment
强化学习对齐的泛化极限
机构 * Aladdin Security Inc(阿拉丁安全公司)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);safety(abstract);分类 cs.AI、cs.LG
AI总结 本文研究了强化学习对齐技术的泛化能力限制,提出复合越狱方法,通过结合多种攻击技术提高攻击成功率,验证了安全训练无法广泛泛化 hypothesis。
Comments 7 pages, 2 figures, 2 tables, accepted at JSAI 2026