The Geometry of Harmful Intent: Training-Free Anomaly Detection via Angular Deviation in LLM Residual Streams
有害意图的几何学:通过残差流中的角度偏差实现无训练异常检测
机构 * Independent Researcher(独立研究者)
专题命中 知识编辑与模型理解 :LLM(title,comments);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出LatentBiopsy方法,通过分析大语言模型残差流激活的几何特性,无需训练即可检测有害提示。该方法基于残差流激活的主成分分析和角度偏差,实现高精度的异常检测,具有低延迟和强鲁棒性。
Comments 20 pages, 10 figures, 3 tables. Training-free harmful-prompt detector via angular deviation in LLM residual streams. Evaluated on six Qwen variants (base / instruct / abliterated). Achieves AUROC over 0.937 (harmful-vs-normative) and 1.000 (harmful-vs-benign-aggressive) with no harmful training data