Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
推理作为攻击面:针对大语言模型的自适应进化思维链越狱方法
机构 * Nanyang Technological University, Singapore(南洋理工大学) ; The University of Hong Kong, Hong Kong, China(香港大学) ; The Hong Kong Polytechnic University, Hong Kong, China(香港理工大学) ; Zhejiang University, Zhejiang, China(浙江大学) ; Renmin University of China, Beijing, China(中国人民大学) ; Sun Yat-sen University, Guangdong, China(中山大学) ; Northeastern University(东北大学) ; Hebei Key Laboratory of Data Science and Knowledge Management(河北省数据科学与知识管理重点实验室)
专题命中 规划推理 :CoT(title,summary_cn);reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI
AI总结 提出自适应进化思维链越狱框架AE-CoT,通过教师角色扮演重写有害目标、分解推理片段、多代进化搜索及自适应变异率控制,有效生成高破坏性越狱提示,在多个模型和数据集上超越现有方法。