A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents
专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments USENIX Security 2024 (https://www.usenix.org/conference/usenixsecurity24/presentation/dahiya)
专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.LG
专题命中 越狱攻击 :alignment(abstract);分类 cs.AI、cs.LG
Journal ref Conference on Neural Information Processing Systems (NeurIPS), Neural Information Processing Systems Foundation, Dec 2023, New Orleans (Louisiana), United States
专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.LG
专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.CY
Comments 13 pages, 9 figures, 7 tables, accepted to findings of EMNLP 2023
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments Technical report
专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.LG
Comments 12 pages
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments Published as a conference paper at the International Conference on Learning Representations (ICLR 2022). Code is available at https://sparseevoattack.github.io/
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.CY
Journal ref Published in ACM CCS 2022. Please cite the CCS version
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments Submitted to conference
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments Accepted at IEEE/CVF WACV 2022 MAP
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments 18 pages, 6 figures
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments Preprint
用于激光雷达距离图像合成的对抗引导扩散
机构 * School of Electrical and Computer Engineering, National Technical University of Athens(雅典国立技术大学电气与计算机工程学院) ; Industrial Systems Institute, Athena Research Center(雅典娜研究中心工业系统研究所)
专题命中 越狱攻击 :safety(abstract);分类 cs.LG;trustworthy(comments)
AI总结 研究针对二维距离图像分割的无限制对抗攻击,提出基于扩散并利用分割损失对抗引导的方法,在SemanticKITTI数据集实验,能跨架构转移,相比基线在有效性与现实性间有独特权衡,实现可控退化。
Comments Accepted at the 1st Workshop on Secure and Trustworthy AI (STAI 2026), co-located with the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD 2026)
机构 * Independent Researcher(独立研究者)
专题命中 越狱攻击 :prompt injection(abstract,comments);分类 cs.AI
Comments 33 pages, 3 figures, 6 tables. Keywords: LLM security; defense-in-depth; prompt injection; activation steering; multimodal sandbox; threat modeling
收敛式绕路劫持:基于技能的大语言模型智能体中的任务保留型资源放大
专题命中 越狱攻击 :safety(abstract);分类 cs.AI
AI总结 该研究提出收敛式绕路劫持攻击,耦合LLM智能体技能选择与规划阶段,可放大任务资源消耗,在DeepSeek-V4-Pro上80.02%的任务会被选中攻击者控制的协调器,任务完成率不变但成本显著提升。
关于智能体大语言模型漏洞的理解、识别与缓解
专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI
AI总结 本研究针对智能体大语言模型安全开展系统文献综述,提出四层漏洞分类法,发现攻击研究多于防御研究、感知层漏洞研究占比偏高等现状,识别7个开放问题及架构耦合的核心不安全根源。
永不停止说话:针对端到端语音语言模型的拒绝服务攻击
专题命中 越狱攻击 :alignment(abstract);分类 cs.AI
AI总结 本研究针对端到端语音语言模型提出基于扰动的拒绝服务攻击,通过优化声学扰动抑制EOS生成以延长解码,在三类开源模型上验证了其攻击有效性及安全风险。
评估部署于智能电网操作中的LLM jailbreaking漏洞:与NERC标准的基准测试
机构 * ECE Department, University of Toronto(多伦多大学电子工程系) ; CIISE, Concordia University(麦吉尔大学CIISE)
专题命中 越狱攻击 :safety(abstract);分类 cs.AI
AI总结 本文评估了智能电网操作中部署LLM的jailbreaking漏洞,通过与NERC标准的基准测试,发现DeepInception方法攻击成功率最高,Claude 3.5 Haiku完全免疫,Gemini 2.0 Flash-Lite最易受攻击。
Comments \c{opyright} 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
DataRx:面向更安全的大语言模型任务特定微调的缺失感知采样方法
机构 * Northwestern Polytechnical University(西北工业大学)
专题命中 越狱攻击 :safety(abstract);分类 cs.CL
AI总结 本文提出DataRx缺失感知采样方法,通过高维隐表示量化安全信号差距,仅用1% BeaverTails安全样本,即可大幅降低Llama3-8B-Instruct的攻击成功率,还可与现有安全数据合成方法结合提升防御效果。
能力悖论:更聪明的审计员如何使多智能体系统更不安全
机构 * University of Chinese Academy of Sciences(中国科学院大学) ; Max Planck Institute for Security and Privacy(马克斯·普朗克安全与隐私研究所) ; Henan Yinzhu Safety Technology Co., Ltd.(河南亿众安全技术有限公司) ; Harbin Institute of Technology, Faculty of Computing(哈尔滨工业大学计算机学院)
专题命中 越狱攻击 :safety(abstract);分类 cs.AI
AI总结 本文研究了多智能体系统中,随着工人能力的提升,系统级攻击成功率反而上升的现象,揭示了语言确定性在攻击传播中的作用,并提出异质性集成验证作为解决方案,以降低攻击成功率。
Comments 28 pages, 6 figures
视觉令牌压缩增强多模态大语言模型的鲁棒性
机构 * Hefei University of Technology(合肥工业大学)
专题命中 越狱攻击 :jailbreak(abstract);分类 cs.LG
AI总结 研究如何增强多模态大语言模型的鲁棒性,核心方法是通过测量视觉令牌与语言特征空间的距离来识别并剪枝分布外视觉令牌,该方法能有效防御越狱攻击、减轻幻觉,提升模型在通用数据集上的表现。
Comments 20 pages, 16 figures. Accepted at ACM Multimedia 2026. Code: https://github.com/Eurek001/OOD-VTP
重写响应路径:BYOK LLM 代理中的静默篡改与提供者签名防御
机构 * Fudan University(复旦大学) ; The Hong Kong University of Science and Technology(香港科学与技术大学) ; Huazhong University of Science and Technology(华中科技大学)
专题命中 越狱攻击 :alignment(abstract);分类 cs.AI
AI总结 研究BYOK LLM代理响应路径完整性问题,发现中继可篡改响应。提出sign-c服务器端方案,能认证执行承载字段和查询,本地垫片验证,加密保护机密性,防御有效拒绝篡改响应,误拒率为零,延迟开销低。
图形用户界面代理相信它们的眼睛吗?诊断状态信念对像素与结构的依赖
机构 * Shenzhen University(深圳大学) ; The Hong Kong University of Science and Technology(香港科技大学)
专题命中 越狱攻击 :safety(abstract);分类 cs.AI
AI总结 研究多模态GUI代理状态信念来源,通过对310个真实探针进行单通道干预形式化视觉状态依赖并测量,核心指标是感知融合差距,发现文本状态信念依赖结构,图像精度高,错误会导致行动失败。
Comments 17 pages, 3 figures
零权限操纵:我们能否信任由大型多模态模型驱动的GUI代理?
机构 * State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) ; Hornor Device Co., Ltd(Hornor设备有限公司) ; Institute of Dataspace, Hefei Comprehensive National Science Center(数据空间研究所,合肥综合性国家科学中心)
专题命中 越狱攻击 :alignment(abstract);分类 cs.AI
AI总结 研究揭示了由大型多模态模型驱动的GUI代理在Android中存在视觉原子性假设的漏洞,通过零权限攻击实现对代理执行的重新绑定,并利用意图对齐策略绕过验证门,展示了代理-操作系统集成中的安全缺陷。
当局部监测器遗漏组合性危害时:诊断多智能体系统中的分布式后门
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
AI总结 研究多智能体系统中分布式后门问题及局部监测器漏洞,提出可观测性边界概念,通过实验证明局部监测器在局部无害时会失效,还展示了特定监测器和门的效果,指出找到能暴露负载的表示形式是未解决问题。