Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: From Evaluation to Diagnosis
大规模视觉-语言模型在细粒度图像任务上的基准测试:从评估到诊断
机构 * School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院,中国) ; Alibaba Group(阿里巴巴集团) ; School of Computer Science and Engineering, School of Intelligence Science and Engineering, and Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Southeast University, China(东南大学计算机科学与工程学院、智能科学与工程学院以及新一代人工智能技术及其交叉应用关键实验室,中国) ; Wangxuan Institute of Computer Technology, National Key Laboratory for Multimedia Information Processing, Peking University, China(北京大学王轩计算机技术研究所、多媒体信息处理国家重点实验室,中国) ; University of Copenhagen, Denmark(丹麦哥本哈根大学)
专题命中 医疗多模态 :diagnosis(title);分类 cs.CV
AI总结 提出FG-BMK基准,含101万问题和28万图像,通过人机双范式评估LVLM的细粒度语义识别与视觉判别能力,诊断失败原因,发现视觉表示、语义对齐等瓶颈。