Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language Models
揭示后门:通过梯度-注意力异常评分对预训练语言模型进行可解释防御
机构 * Umeå University(乌梅学院) ; Nanyang Technological University(南洋理工大学)
AI总结 本文提出通过梯度-注意力异常评分对预训练语言模型进行可解释防御,有效降低后门攻击成功率。
Comments 17 pages total (9 pages main text + 6 pages appendix + references), 16 figures. Preprint version; the final camera-ready version may differ. Accepted to ICLR 2026