Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
通过细化文本嵌入缓解大型视觉语言模型中的幻觉
机构 * University of Maryland(马里兰大学) ; Dolby Laboratories(杜比实验室) ; Capital One
专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV
AI总结 针对大型视觉语言模型因过度依赖文本先验而忽视视觉线索导致的幻觉问题,提出一种简单有效的视觉特征融入方法,通过学习视觉信息化的文本嵌入来平衡注意力分布,显著降低幻觉并提升多模态推理能力。
Comments Accepted at The 64th Annual Meeting of the Association for Computational Linguistics