Stepwise Credit Assignment for GRPO on Flow-Matching Models
分步信用分配用于流匹配模型中的GRPO
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Adobe Research(Adobe研究院)
AI总结 本文提出分步流GRPO,通过分步奖励改进提升样本效率和收敛速度,引入DDIM启发的SDE提升奖励质量。
Comments Accepted to the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026 Project page: https://stepwiseflowgrpo.com