Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
野外对比归因:对LLM在真实基准上失败的可解释性分析
机构 * Southern University of Science and Technology(南方科技大学) ; Microsoft(微软)
专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
AI总结 本文研究了对比归因作为分析LLM在现实场景中失败的实用工具,通过对比错误输出与正确替代输出的logit差异,揭示了归因方法在不同数据集、模型规模和训练检查点上的表现,展示了其在某些失败案例中的有效性及局限性。
Comments 45 pages, 16 figures, 16 tables