Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
全局上下文还是局部细节?面向幻觉缓解的自适应视觉 grounding
机构 * School of Astronautics, Beihang University(北航航天学院) ; Longcat Interaction Team, Meituan(美团Longcat交互团队) ; Tianmushan Laboratory, Beihang University(北航天门山实验室)
专题命中 视觉定位与Grounding :grounding(title,title_cn);VLM(abstract,abstract_cn);LLaVA(abstract,abstract_cn);InternVL(abstract,abstract_cn)
AI总结 本文提出PND框架,通过双路径对比在解码过程中增强视觉真实性,减少幻觉并提升描述细节,无需模型微调。
Comments 9 pages, 8 figures, Findings of ACL 2025