ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
ShadowCoT:针对大语言模型内部推理机制的隐秘推理后门攻击
机构 * School of Communication and Information Engineering, Shanghai University(上海大学通信与信息工程学院) ; College of Computer Science and Information Technology, IAU(国际应用技术大学计算机科学与信息学院) ; Center for AI Research (CAIR), University of Agder (UiA)(阿格德大学人工智能研究中心)
专题命中 复杂问题求解 :reasoning(title,abstract);CoT(abstract,abstract_cn);chain-of-thought(abstract);分类 cs.CL
AI总结 ShadowCoT提出一种新型攻击框架,通过操控模型认知推理路径,实现多步推理链的劫持,产生逻辑连贯但对抗性的结果,实验显示其攻击成功率高达94.4%。
Comments Zhao et al., 16 pages, 2025, uploaded by Hanzhou Wu, Shanghai University
Journal ref IEEE Transactions on Information Forensics and Security (2026)