TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable Reward
TDM-R1: 通过非可微奖励强化少步扩散模型
机构 * Hong Kong University of Science and Technology(香港科技大学) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; hi-Lab, Xiaohongshu Inc(小红书实验室) ; Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
AI总结 TDM-R1通过非可微奖励强化少步扩散模型,提升文本到图像生成性能