Maximizing the efficiency of human feedback in AI alignment: a comparative analysis
最大化人类反馈在AI对齐中的效率:比较分析
机构 * University College Dublin, Ireland(都柏林大学学院)
专题命中 偏好对齐 :RLHF(summary_cn,abstract);alignment(title,abstract);分类 cs.AI
AI总结 本文探讨了RLHF中偏好推断的替代采样与评估策略,提出Swiss InfoGain方法在受限标注预算下表现更优,且更节省样本,提升了偏好学习的效率与鲁棒性。
Comments 17 pages, 6 figures, 6 algorithms. AICS2025
Journal ref 33rd International Conference on Artificial Intelligence and Cognitive Science (AICS 2025)