Precision Knowledge Editing: Enhancing Safety in Large Language Models
专题命中 越狱攻击 :safety(title,abstract);分类 cs.CL、cs.AI
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 越狱攻击 :safety(title,abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI、cs.LG
专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :safety(title,abstract);分类 cs.CL、cs.AI
Comments ACL 2024
专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.AI
Comments Published at ICML 2024 Workshop on Foundation Models in the Wild
专题命中 越狱攻击 :prompt injection(title,abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :alignment(title,abstract);分类 cs.CL、cs.LG
Comments 8 pages, 6 figures
专题命中 越狱攻击 :safety(title,abstract);分类 cs.CL、cs.AI
Comments Submitted to EMNLP 2024
专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.AI
Comments 4 pages, 2 figures. This paper was submitted to The 7th Deep Learning Security and Privacy Workshop (DLSP 2024) and was accepted as extended abstract, see https://dlsp2024.ieee-security.org/
专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :safety(title,abstract);分类 cs.CL、cs.AI
Comments 11 pages, 2 figures
机构 * Alibaba Group(阿里巴巴集团) ; Beijing Electronic Science and Technology Institute(北京电子科技研究所) ; Nanjing University(南京大学) ; Renmin University of China(中国人民大学) ; Northeastern University(东北大学) ; BraneMatrix AI ; Nanyang Technological University(南洋理工大学)
专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.AI
Comments Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
大型音频语言模型综述:通用性、可信度与展望
机构 * Nanyang Technological University(南洋理工大学) ; Independent Researcher(独立研究者) ; The University of Melbourne(墨尔本大学) ; North China Electric Power University(华北电力大学) ; Beijing University of Posts and Telecommunications(北京邮电大学) ; University of Chinese Academy of Sciences(中国科学院大学) ; University of Science and Technology of China(中国科学技术大学) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; Shanghai AI Laboratory(上海人工智能实验室) ; Huazhong University of Science and Technology(华中科技大学) ; Tsinghua University(清华大学) ; Fortemedia Singapore(富媒体新加坡) ; Tencent(腾讯) ; Fudan University(复旦大学) ; Wuhan University(武汉大学) ; Chinese University of Hong Kong(香港中文大学) ; Chongqing University of Posts and Telecommunications(重庆邮电大学) ; University of Illinois Chicago(伊利诺伊大学芝加哥分校)
专题命中 越狱攻击 :trustworthy(abstract,abstract_cn);alignment(abstract);safety(abstract)
AI总结 本文综述了大型音频语言模型的通用性、可信度及未来发展方向,探讨了其架构创新、对齐算法及安全风险,并提出了防御深入、因果音频世界建模等策略以提升音频智能的可信度。
ARMOR:通过精细推理对齐安全的大语言模型
专题命中 越狱攻击 :alignment(abstract);RLHF(abstract);safety(abstract);jailbreak(abstract)
AI总结 研究大语言模型安全问题,针对现有方法不足,提出ARMOR,通过三步推理管道提取核心恶意意图并验证,开发ARMOR-Think提升性能,在防御越狱攻击上效果显著。
Comments ICLR 2026
TRACE: 任务感知的自适应自我进化代理越狱
专题命中 越狱攻击 :safety(abstract,abstract_cn);alignment(abstract);jailbreak(abstract)
AI总结 提出TRACE框架,通过任务分解、场景伪装和Q学习进化,实现LLM代理的高效越狱,在AgentHarm和AdvCUA上达到最高100%绕过率。
Comments 16 pages, 7 figures
大规模真实对话分析揭示LLM越狱的复杂性界限
机构 * Valencian Research Institute for Artificial Intelligence (VRAIN), Universitat Politècnica de València, Valencia, Spain.(瓦伦西亚人工智能研究 institute,瓦伦西亚理工大学,西班牙瓦伦西亚) ; Department of Computer Science, The University of Chicago, Chicago, USA(计算机科学系,芝加哥大学,美国芝加哥) ; Center for Automation and Robotics, Spanish National Research Council, Madrid, Spain(自动化与机器人中心,西班牙国家研究委员会,西班牙马德里)
专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY
AI总结 通过分析超过200万条真实对话,发现越狱尝试的复杂性并不显著高于正常对话,且攻击复杂性随时间保持稳定,表明LLM安全演化受人类创造力限制。
Comments Code: https://github.com/ACMCMC/risky-conversations Results: https://huggingface.co/risky-conversations Visualizer: https://huggingface.co/spaces/risky-conversations/Visualizer
稀疏令牌足矣:通过令牌感知梯度优化越狱音频语言模型
机构 * Wuhan University ; Institute for Math \& AI, Wuhan University ; Huazhong University of Science ; Shanghai Jiao Tong University ; Xidian University
专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出令牌感知梯度优化(TAGO)方法,通过仅保留高梯度能量的音频令牌对应的波形梯度,实现稀疏越狱攻击,在保持高成功率的同时大幅减少优化量。
Comments To appear in the 43rd International Conference on Machine Learning (ICML 2026)
组合式劫持:对齐大语言模型中突变链相互作用的实证分析
专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);AI safety(abstract)
AI总结 本文研究了对齐大语言模型中突变链的相互作用,通过十二个基线突变器评估有害提示的组合效果,揭示了突变器间的非均匀交互特性及安全对齐结构的潜在属性。
Comments 16 pages, 7 figures, 3 tables
方向嵌入平滑用于鲁棒视觉语言模型
机构 * Mitsubishi Electric Research Laboratories(三菱电机研究实验室)
专题命中 越狱攻击 :alignment(abstract);safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出通过方向嵌入平滑技术提升视觉语言模型的鲁棒性,针对多模态劫持攻击,验证了方向嵌入噪声在降低攻击成功率中的有效性。
Comments Accepted at ICLR 2026 Workshop on Agents in the Wild
通过校准劫持大语言模型
机构 * Department of Computer Science, Peking University, Beijing, China(计算机科学系,北京大学,北京,中国)
专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 通过校准劫持大语言模型,提出一种新的聚合框架,提升攻击成功率并降低劫持成本。
突破大语言模型与视觉语言模型:机制、评估与统一防御
专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);red teaming(abstract)
AI总结 本文提出三维框架系统回顾大语言模型和视觉语言模型的突破攻击与防御机制,提出统一防御原则并讨论未来研究方向。
AdvPrefix: 一种面向 nuanced LLM 模型 jailbreak 的目标
机构 * University of Maryland, College Park(马里兰大学 College Park分校) ; FAIR, Meta(Meta的FAIR)
专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 AdvPrefix 提出了一种新的目标,通过结合高预填充攻击成功率和低负对数似然来选择模型依赖的前缀,以提升 jailbreak 攻击的精确性和有效性。
专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.CY
专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 15 pages, 6 figures, 7 tables
专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Version 2 updates: Added comparison of three more evaluation methods and their reliability check using human labeling. Added results for jailbreaking Llama2 (individual behavior) and included complexity and hyperparameter analysis. Revised objectives for prompt leaking. Other minor changes made
机构 * University of Strathclyde(斯特拉思克莱德大学)
专题命中 越狱攻击 :alignment(abstract,comments);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI
Comments Published at Transaction of Machine Learning Research 08/2025, Large Language Models (LLMs), Interference-time activation shifting, Steerability, Explainability, AI alignment, Interpretability