Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion
通过偏好维度扩展解释并打破安全与有益性天花板
机构 * Huazhong University of Science and Technology(华中科技大学) ; Nanyang Technological University(南洋理工大学) ; Tsinghua University(清华大学) ; Chongqing University(重庆大学)
专题命中 偏好对齐 :safety(title);alignment(abstract);harmlessness(abstract);分类 cs.AI
AI总结 本文提出MORA方法,通过多维奖励整合解决多目标对齐中的安全与有益性矛盾,实验显示在序列对齐中提升5%-12.4%,同时对齐中整体奖励提升4.6%。