Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models
Linear-DPO: 用于扩散和流匹配生成模型的线性直接偏好优化
机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) ; Alibaba Group(阿里巴巴集团)
AI总结 本文提出Linear-DPO,通过统一的反向时间SDE框架推导出涵盖扩散和流匹配的通用DPO目标,指出标准DPO目标在文本到图像生成中不最优,并通过定性定量实验验证了其在扩散模型和流匹配模型上的优越性。
Comments Code and models are available at: https://github.com/Whynot0101/Linear-DPO . Work done during an internship at Alibaba Group