Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts
不要告诉答案,真正引导RL Rollout期间的推理
机构 * School of Data Science, Fudan University(复旦大学数据科学学院) ; Shanghai Institute of Artificial Intelligence for Education, East China Normal University(华东师范大学教育人工智能研究所) ; College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) ; Ant Group(蚂蚁集团)
AI总结 针对强化学习中任务难度超出模型能力导致奖励稀疏的问题,提出HINT框架,通过元提示和亲和力感知策略优化,提升模型推理能力并保持训练稳定性。
Comments Accepted to Findings of ACL 2026. 19 pages