REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems
REDAgentBench:可执行的红队测试与LLM智能体系统的忠实测量
专题命中 红队测试 :red teaming(title);safety(abstract);分类 cs.AI
AI总结 研究针对LLM智能体安全评估的缺陷,推出REDAgentBench框架,经实验发现其宏平均ASR为65.69%,还揭示了识别-执行差距,无训练策略提醒可减少70%以上违规
Comments 6 figures, 4 tables. Supplementary material included