DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization
DGPO: 超越成对偏好与方向一致的组级优化
机构 * Information Hub, The Hong Kong University of Science and Technology (Guangzhou), China(香港科学与技术大学(广州)信息中心,中国) ; The Hong Kong University of Science and Technology, Hong Kong SAR(香港科学与技术大学,香港特别行政区)
专题命中 其他推理 :reasoning(abstract);分类 cs.CL
AI总结 本文提出DGPO框架,通过组级监督信号和多候选比较建模方向一致性,提升偏好优化的准确性和多样性,实验显示在多个数据集上平均提升3.6%。