One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
一个泄漏:预训练模型暴露如何放大微调LLM中的劫持风险
机构 * Institute of Science Tokyo(东京科学研究所) ; Riken AIP(理化学研究所AIP)
专题命中 知识编辑与模型理解 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI
AI总结 本研究揭示了预训练模型暴露如何放大微调LLM的劫持风险,提出PGP攻击方法,证明了预训练到微调过程中存在的安全漏洞。
Comments This paper has been accepted to the ACM SIGSAC Conference on Computer and Communications Security (ACM CCS)