Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias
剖析不公平的评判者:基于大语言模型作为评判者的偏差的机械可解释性分析
机构 * AMAP, Alibaba Group(阿里巴巴集团AMAP实验室) ; Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; University of Southern California(南加州大学) ; University of Michigan, Ann Arbor(密歇根大学安娜堡分校)
AI总结 研究大语言模型作为评判者的偏差,提出偏差在隐藏状态有表示层面解释。通过七位评判者等实验发现,偏差输入沿特定子空间移动,操纵隐藏状态可控制评分,线性投影能预测评判失败,统一多方面内容。
Comments 58 pages, 13 figures, 30 tables; project page: https://xzx34.github.io/unfair-judge/