Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
策展合成数据不会崩溃:具有多元偏好的生成式再训练的理论研究
机构 * University of Washington(华盛顿大学)
AI总结 通过理论分析证明,基于多个奖励函数进行策展的递归训练可以避免生成模型崩溃,并收敛到满足加权纳什议价解的稳定分布。
Comments Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026