Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
近端策略优化区域:教师存在于提示中,而非梯度中
机构 * NVIDIA(英伟达)
专题命中 后训练与偏好优化 :LLM(abstract,abstract_cn);language model(abstract);分类 cs.CL
AI总结 提出ZPPO方法,通过将教师知识注入提示而非策略梯度,解决小模型知识蒸馏中教师梯度主导和强化学习策略漂移问题,在多种规模模型上超越现有方法。
Comments Project page: https://byungkwanlee.github.io/ZPPO-page/