Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
测量测量者:大型视觉-语言模型幻觉基准的品质评估
机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, and also with University of Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院,以及中国科学院大学)
AI总结 本文提出HQM框架评估幻觉基准质量,发现现有基准存在可靠性与有效性不足问题,并提出HQQ基准以提升评估效果,揭示LVLMs在幻觉问题上的严重缺陷。