Pave-GRPO: Beyond Instantaneous Guidance through Principled Average Velocity Decomposition
Pave-GRPO:通过原则性平均速度分解超越瞬时引导
机构 * University of Science and Technology of China(中国科学技术大学) ; Shanghai Jiao Tong University(上海交通大学) ; Fudan University(复旦大学) ; Harbin Institute of Technology(哈尔滨工业大学) ; Beihang University(北京航空航天大学) ; Shanghai AI Laboratory(上海人工智能实验室)
AI总结 提出Pave-GRPO方法,通过原则性平均速度分解将粗粒度过渡分解为细粒度子轨迹,在不增加生成成本的情况下将奖励反馈传播到更多中间步骤,实现更全面的偏好对齐。
Comments 18 pages,9 figures