Not All Preferences Are Created Equal: Stability-Aware and Gradient-Efficient Alignment for Reasoning Models
并非所有偏好都同等重要:面向推理模型的稳定性感知与梯度高效对齐
机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空航天信息研究所) ; Baidu Inc.(百度公司) ; Department of Computer Science, University of Toronto(多伦多大学计算机科学系) ; School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.AI
AI总结 SAGE通过动态框架提升推理模型对齐的稳定性与梯度效率,通过信噪比优化和稳定性感知评分函数实现更高效的训练。