RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation
RoguePrompt:用于自重构以规避大型语言模型(LLM)审核的双层编码
专题命中 安全训练 :safety(abstract);jailbreak(abstract);分类 cs.AI
AI总结 RoguePrompt是一种采用双层编码(维吉尼亚密码+ROT13)的LLM越狱流程,在黑盒威胁模型下对313个被拒提示词测试,实现93.93%的过滤绕过率,提供多阶段越狱失效的阶段级证据。
Comments This manuscript supersedes the preliminary version available as arXiv:2511.18790. The work has been substantially revised, expanded, and reorganized, with a refined threat model, revised methodology, clearer stage-level evaluation criteria, and expanded analysis of moderation bypass, instruction reconstruction, and execution