Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
真实还是编造?利用因果归因减轻解释中的奖励作弊
机构 * Institute for Logic, Language and Computation (ILLC), University of Amsterdam(逻辑、语言与计算研究所(ILLC),阿姆斯特丹大学) ; Institute for Language, Cognition and Computation (ILCC), University of Edinburgh(语言、认知与计算研究所(ILCC),爱丁堡大学)
专题命中 逻辑推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL
AI总结 研究大语言模型解释中因奖励模型导致的奖励作弊问题,提出用预测的因果归因丰富奖励模型输入的方法,该方法能减少大语言模型生成误导性解释的倾向,提升解释忠实度。
Comments ICLR 2026 Camera-ready