What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics
中间层知道什么:从熵动力学检测越狱
机构 * LMU Munich(慕尼黑大学) ; relAI – Konrad Zuse School of Excellence in Reliable AI(relAI – 康拉德·楚泽可靠人工智能卓越学校) ; University of Pisa(比萨大学) ; Munich Center for Machine Learning(慕尼黑机器学习中心) ; School of Engineering, Computing and Mathematics, Oxford Brookes University(牛津布鲁克斯大学工程、计算与数学学院)
专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 通过分析冻结LLM各层的token级预测熵轨迹,发现中间层的熵动力学特征(如基于排名的单调趋势分数)能有效检测越狱攻击,且无需额外训练。
Comments Accepted at the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD) 2026. A short version accepted at EIML@ICML 2026