Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner
朝着解耦偏好优化动态:压制失败者,保留胜利者
机构 * South China University of Technology(南方科技大学) ; Columbia University(哥伦比亚大学) ; University of Waterloo(滑铁卢大学)
AI总结 本文提出奖励校准方法,通过解耦带宽条件实现压制失败者并保留胜利者的优化动态,提升了下游性能。
Journal ref International Conference on Machine Learning(ICML) 2026