Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization
Margin Adaptive DPO: 利用奖励模型实现偏好优化中的细粒度控制
机构 * Independent Researcher(独立研究者)
专题命中 后训练与偏好优化 :preference optimization(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
AI总结 提出Margin-Adaptive Direct Preference Optimization (MADPO)方法,通过奖励模型估计偏好边界并自适应调整DPO损失权重,实现实例级别的细粒度控制,在摘要任务上优于现有方法。