Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis
Omni-RRM:通过基于自动评分标准的偏好合成推进全模态奖励建模
机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ; Tsinghua University(清华大学) ; Institute of Artificial Intelligence (TeleAI)(人工智能研究所) ; ShiFang Technology Inc.(ShiFang科技公司)
专题命中 偏好对齐 :RLHF(abstract,abstract_cn);alignment(abstract);分类 cs.CL
AI总结 研究针对多模态大语言模型对齐难题,提出全模态基于评分标准的奖励模型Omni-RRM。通过自动评分标准偏好合成构建数据集,用特定训练方案提升奖励辨别力,在多模态基准测试中达最优精度,还能指导选择与迁移,提升模型对齐能力。
Comments ECCV 2026