Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback
Oracle-RLAIF:一种利用排名反馈强化学习改进多模态视频模型微调框架的方法
机构 * Stanford University(斯坦福大学) ; Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室) ; Microsoft(微软公司)
AI总结 提出Oracle-RLAIF框架,用通用排序器替代奖励模型,结合基于GRPO的排名损失函数GRPO_rank,实现更高效的多模态视频模型微调,在多个基准上优于现有方法。
Comments Proceedings of the 39th Annual Conference on Neural Information Processing Systems, ARLET Workshop (Aligning Reinforcement Learning Experimentalists and Theorists)
Journal ref Transactions on Machine Learning Research, Vol. 2026, June 2026