Know Thy Enemy: Securing LLMs Against Prompt Injection via Diverse Data Synthesis and Instruction-Level Chain-of-Thought Learning
知彼者智:通过多样化数据合成和指令级推理学习增强大语言模型抗提示注入能力
机构 * State Key Laboratory of Complex System Modeling and Simulation Technology(复杂系统建模与仿真技术国家重点实验室) ; Science and Technology on Integrated Information System Laboratory, Institute of Software Chinese Academy of Sciences(中国科学院软件研究所综合信息系统技术国家级重点实验室) ; University of Chinese Academy of Sciences(中国科学院大学) ; Nanyang Technological University(南洋理工大学) ; Beijing Forestry University(北京林业大学)
专题命中 其他推理 :chain-of-thought(title,abstract);分类 cs.AI
AI总结 本文提出InstruCoT方法,通过多样化数据合成和指令级推理学习提升大语言模型对提示注入攻击的防御能力,实验表明其在行为偏差、隐私泄露和有害输出方面显著优于基线方法。
Comments 19 pages, 6 figures; accepted by ACL 2026 Findings