What AI Red-Team Evaluations Can and Cannot Prove
人工智能红队评估能证明什么以及不能证明什么
机构 * APIsec Research Labs(APIsec研究实验室)
专题命中 红队测试 :safety(abstract);red teaming(abstract);分类 cs.AI
AI总结 研究人工智能红队评估能证与不能证之事,通过定义证据上限确定界限,发现高于某危害率基准可证类别,低于则否,该界限不限于基准,审核评估套件发现当前基准对高频危害足够,对罕见灾难性不足。
Comments 21 pages, 4 figures, 5 tables. Code and data links provided in the manuscript. v2: corrected Figure 1(b); corrected required sample sizes in Table 4 and in Sections 4.2, 4.6 and 5.2, which had been rounded rather than taken to the ceiling; corrected the sample-size expression stated in Methods; minor corrections to Table 1 and the Figure 2 caption. No theorem, result or conclusion is affected