REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
REOPD:面向在线策略蒸馏的可靠性自适应奖励外推
机构 * Peking University(北京大学) ; Southwest Jiaotong University(西南交通大学) ; University of Science and Technology of China(中国科学技术大学) ; School of Computer Science and Technology,University of Science and Technology of China(中国科学技术大学计算机科学与技术学院) ; University of Electronic Science and Technology of China(电子科技大学) ; School of Automation Engineering,University of Electronic Science and Technology of China(电子科技大学自动化工程学院) ; Renmin University of China(中国人民大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Autonomous Driving, Shanghai Artificial Intelligence Laboratory(上海人工智能实验室自动驾驶研究中心)
AI总结 本文提出面向在线策略蒸馏的可靠性自适应奖励外推框架REOPD,通过token级系数实现细粒度可靠性自适应,在多场景下优于基线方法,无需额外模型即可提升训练稳定性与性能。