Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy Biases
对抗扩散模型中的奖励过度优化:从归纳偏置和优先偏置的角度出发
机构 * Institute of Artificial Intelligence, School of Computer Science, Wuhan University, China ; Hubei Luojia Laboratory, Wuhan, China ; The University of Sydney, Australia ; JD Explore Academy, Beijing, China ; Nanyang Technological University, Singapore
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG
AI总结 本文提出TDPO-R算法,通过利用扩散模型的时间归纳偏置和抑制活跃神经元的优先偏置,有效缓解奖励过度优化问题。
Comments Accepted to ICML 2024
Journal ref International Conference on Machine Learning, pp. 60396-60413, 2024