What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs
LLM 解释的内容并非其信念:基于模型自身输入信念评估解释充分性
机构 * New York University(纽约大学)
专题命中 评测与基准 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 提出自洽充分性(SCSuff)指标,利用 LLM 自身生成替代输入来评估自由文本解释的充分性,发现 LLM 解释通常不充分且与模型大小、准确率等弱相关。
Comments 23 pages, 9 figures, 13 tables, Forty-Third International Conference on Machine Learning (ICML 2026)
Journal ref Forty-Third International Conference on Machine Learning (ICML 2026)