Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
Reflect-Guard: 通过逻辑自我反思增强大语言模型对对抗性提示的防护
机构 * Yale University(耶鲁大学) ; Columbia University(哥伦比亚大学) ; Citigroup(摩根大通) ; Independent Researcher(独立研究者)
专题命中 指令微调 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
AI总结 提出Reflect-Guard方法,通过参数高效微调为大语言模型安全分类器注入链式思维自我反思能力,显著提升对对抗性越狱攻击的检测性能。
Comments 12 pages, 2 figures, and 4 tables