Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH
隐于无形:使用DECOMPBENCH基准测试代理安全对抗分解攻击
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Simons Institute, UC Berkeley(Simons研究所,伯克利大学)
专题命中 安全训练 :safety(title,abstract);分类 cs.AI、cs.LG
AI总结 提出DeCompBench基准,通过分解攻击将有害任务拆分为良性子任务,揭示现有代理安全机制在对抗分解攻击时的脆弱性。