Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View
为扩散模型设计强化学习:统一的路径空间视角
机构 * State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室) ; School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) ; ByteDance(字节跳动)
AI总结 本文从路径空间视角统一了扩散模型RL算法的原理,推导得到降方差值梯度形式,提出多样本KDE估计器和尺度受限权重族,在SD3.5-M等模型上验证了方法有效性并优于基线。
Comments 29 pages, 9 figures, 4 tables; work in progress