DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
DiG:通过差异 grounding 提升多模态大语言模型的细粒度感知
机构 * University of Science and Technology of China(中国科学技术大学) ; State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
AI总结 本文提出DiG框架,通过学习相似图像对的差异识别提升多模态大语言模型的细粒度感知能力,实验表明其在多个视觉感知基准上表现优异。