Efficient LLM Safety Evaluation through Multi-Agent Debate
通过多智能体辩论实现高效的LLM安全评估
机构 * Beijing Institute of AI Safety and Governance(北京人工智能安全与治理研究院) ; Beijing Key Laboratory of Safe AI and Super Alignment(北京安全人工智能与超对齐重点实验室) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; CSE, The Chinese University of Hong Kong(香港中文大学电子工程系) ; University of Chinese Academy of Sciences(中国科学院大学) ; Department of Mathematics, The Chinese University of Hong Kong(香港中文大学数学系) ; Long-term AI(长期人工智能)
专题命中 安全评测 :safety(title,abstract);jailbreak(abstract);分类 cs.AI
AI总结 本文提出多智能体辩论框架,利用HAJailBench基准测试,提升LLM安全评估的可靠性与经济性,验证了少量辩论轮次即可获得显著效果。
Comments 15 pages, 5 figures, 10 tables. Updated abstract to fix an incconsistency issue with the main paper: HAJailBench size (12,000 -> 11,100)