Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models
在学习中保持好奇心:通过自适应自蒸馏实现熵保持的监督微调以提升大推理模型
机构 * City University of Hong Kong(香港城市大学) ; The Hong Kong University of Science and Technology(香港科学与技术大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.LG
AI总结 CurioSFT通过自适应自蒸馏和熵引导温度选择,提升大推理模型的探索能力,在监督微调和强化学习阶段均取得显著改进。