Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
Fine-R1: 通过链式推理使多模态大语言模型在细粒度视觉识别中脱颖而出
机构 * Wangxuan Institute of Computer Technology(王轩计算机技术研究所)
专题命中 图文多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 Fine-R1通过链式推理监督微调和三元组增强策略优化,提升多模态大语言模型在细粒度视觉识别中的表现,仅用4次训练即超越现有模型。
Comments Published as a conference paper at ICLR 2026. The models are available at https://huggingface.co/collections/StevenHH2000/fine-r1