Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
Faithfulness Serum:通过归因引导缓解文本解释中LLM决策的忠实性差距
机构 * Blavatnik School of Computer Science, Tel Aviv University(塔尔瓦夫大学比拉维克计算机科学学院)
专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
AI总结 本文通过归因引导方法提升LLM决策解释的忠实性,验证了现有解释的epistemic忠实性不足,并展示了改进方法在多个模型和基准上的有效性。
Comments 24 pages, multiple figures (e.g., at least 6 main figures), includes experiments across several benchmarks (MMLU, CommonsenseQA, SciQ, ARC, OpenBookQA); code available on GitHub