Visual Attention Drifts,but Anchors Hold:Mitigating Hallucination in Multimodal Large Language Models via Cross-Layer Visual Anchors
视觉注意力漂移,但锚点仍有效:通过跨层视觉锚点缓解多模态大语言模型的幻觉
机构 * School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院) ; School of Computer Science, Wuhan University(武汉大学计算机学院)
专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract)
AI总结 本文提出CLVA方法,通过跨层视觉锚点缓解多模态大语言模型的幻觉问题,强调中间层视觉锚点的重要性,无需额外训练,有效抑制深度层注意力漂移。