Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
通过信息论指导消除奖励模型中的归纳偏置
机构 * Qwen Large Model Application Team, Alibaba(阿里巴巴大模型应用团队) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Shenzhen Research Institute of Big Data(深圳大数据研究院)
AI总结 本文提出了一种基于信息论的奖励模型去偏方法DIR,通过最大化奖励模型评分与人类偏好对之间的互信息,同时最小化奖励模型输出与偏好输入偏置属性之间的互信息,从而有效缓解归纳偏置问题并提升RLHF性能。
Comments Published as a conference paper at The International Conference on Learning Representations (ICLR) 2026