Empirical Validation of the Classification-Verification Dichotomy for AI Safety Gates
对AI安全闸门中分类-验证二元对立的实证验证
专题命中 安全训练 :safety(title,abstract);AI safety(title);分类 cs.AI、cs.LG
AI总结 研究通过实证表明,基于分类器的安全闸门在AI系统经过数百次迭代后无法维持可靠监管,且结构上无法实现。
Comments 21 pages, 9 figures. Companion theory paper: doi:10.5281/zenodo.19237451