Understanding helpfulness and harmless tension in reward models
理解奖励模型中的有用性与无害性张力
机构 * University of Copenhagen(哥本哈根大学)
专题命中 后训练与偏好优化 :RLHF(abstract,abstract_cn);language model(abstract);分类 cs.CL、cs.LG
AI总结 通过激活分析和消融实验,发现奖励模型中有用性和无害性目标存在干扰,共享神经元对模型行为影响不成比例,导致对齐张力。
Comments The source code used in this study is publicly available at: https://github.com/EshaanT/RM-alignment\_tension