Dynamics-Predictive Sampling for Active RL Finetuning of Large Reasoning Models
动态预测采样用于主动强化学习微调大推理模型
机构 * Department of Automation, Tsinghua University(清华大学自动化系)
专题命中 测试时计算 :reasoning(title,abstract);planning(abstract);分类 cs.AI、cs.LG
AI总结 DPS通过动态预测方法减少回放成本,提升大推理模型的强化学习微调效率与性能。
Comments Accepted to ICLR 2026