Jury Duty: Calibration and Orientation Failures in MLLM-as-a-Judge Under Cultural Ambiguity
陪审团职责:文化模糊性下MLLM作为评判者的校准与定向失败
机构 * Salesforce AI Research(Salesforce AI 研究院) ; University of Cambridge(剑桥大学) ; University of Colorado Boulder(科罗拉多大学博尔德分校) ; Carnegie Mellon University(卡内基梅隆大学) ; UIUC(伊利诺伊大学厄巴纳-香槟分校)
专题命中 幻觉与鲁棒性 :MLLM(title,title_cn);分类 cs.CV、cs.AI
AI总结 针对MLLM作为评判者在跨文化场景中的偏差问题,提出VOIR DIRE基准,通过分析校准失败(压缩尺度)和定向失败(默认一种文化规范),揭示模型偏向于更宽容的文化解读。
Comments Under Review