PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration
PEBS: 用于RLHF奖励模型校准的每评分者经验贝叶斯收缩
专题命中 偏好对齐 :RLHF(title,title_cn);分类 cs.AI、cs.LG;alignment(comments)
AI总结 针对RLHF中奖励模型忽略评分者个体差异的问题,提出PEBS方法,对每个评分者拟合仿射校准器并应用经验贝叶斯收缩,在PRISM和PluriHarms数据集上分别降低RMSE 8.58%和9.66%。
Comments Accepted at the ICML 2026 Workshop on Pluralistic Alignment. Code: https://github.com/deadsmash07/pebs-pluralistic