Learning from the Right Rollouts: Data Attribution for PPO-based LLM Post-Training
从正确的回放中学习:基于PPO的LLM后训练的数据归因
机构 * Northwestern University(西北大学) ; Stevens Institute of Technology(史蒂文斯理工学院)
专题命中 后训练与偏好优化 :post-training(title,abstract);LLM(title);SFT(abstract);分类 cs.LG
AI总结 本文提出Influence-Guided PPO框架,通过计算影响分数筛选反向齐梯度的回放片段,提升PPO后训练效率并减少不忠实的推理。