Breaking Customized LLMs for Coding: Automated Red Teaming for Instruction Backdoor Attacks
针对代码任务定制化大语言模型的突破:针对指令后门攻击的自动红队测试
专题命中 红队测试 :red teaming(title)
AI总结 本文提出ARIA自动红队测试框架,可高效生成针对代码定制化LLM的隐蔽带后门指令,攻击成功率达0.945且规避检测能力强,性能优于现有基线攻击。
Comments Accepted to the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026