DIVA-GRPO: Enhancing Multimodal Reasoning through Difficulty-Adaptive Variant Advantage
DIVA-GRPO:通过难度自适应变体优势增强多模态推理
机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, CAS, Beijing, China(人工智能安全国家重点实验室,计算技术研究所,中国科学院,北京,中国) ; University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国) ; Kuaishou Technology, Beijing, China(快手科技,北京,中国)
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI
AI总结 DIVA-GRPO通过难度自适应变体优势方法提升多模态推理能力,解决GRPO在困难问题上的奖励稀疏性和优势消失问题,提升训练稳定性与推理性能。
Comments Accepted to ICLR 2026. Code and models are available at https://github.com/Siaaaaaa1/DIVA-GRPO