arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Empirical Methods in Natural Language Processing · 会议 · Natural Language Processing

2026-06-03 至 2026-06-03 共收录 1
2502.09755 2026-06-03 cs.CR cs.LG

Jailbreak Attack Initializations as Extractors of Compliance Directions

越狱攻击初始化作为合规方向的提取器

Amit Levi, Rom Himelstein, Yaniv Nemcovsky, Avi Mendelson, Chaim Baskin

机构 * Department of Computer Science, Technion - Israel Institute of Technology(技术学院计算机科学系) Department of Data and Decision Science, Technion - Israel Institute of Technology(技术学院数据与决策科学系) School of Electrical and Computer Engineering Engineering, Ben-Gurion University of the Negev(内盖夫本· Gurion大学电气与计算机工程学院)

AI总结 本文发现基于梯度的越狱攻击初始化会收敛到抑制拒绝的单一合规方向,并据此提出CRI框架,通过沿合规方向投影未见提示来提高攻击成功率并降低计算开销。

Comments Accepted to Findings of the Association for Computational Linguistics 2025 (EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏