Which Neurons Detect Malicious Code? A Probing Study of LLM Security Knowledge
哪些神经元能检测恶意代码?对大语言模型安全知识的探索性研究
专题命中 安全训练 :alignment(abstract)
AI总结 研究探索大语言模型中检测恶意代码的神经元,应用机械可解释性方法定位相关神经元,通过对恶意和良性PyPI包实验发现放大促进神经元、抑制抑制神经元可提升准确率,有助于了解模型编码恶意概念方式,为可靠防御机制提供思路。
Comments The paper has been peer reviewed and accepted for publication in the 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)