BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems
BELLS-O:评估LLM监督系统的运营权衡
机构 * University of Graz, Graz, Austria(格拉茨大学) ; Supervised Program for Alignment Research (SPAR)(对齐研究监督计划 (SPAR)) ; Centre pour la Sécurité de l'IA (CeSIA), Paris, France(人工智能安全研究中心 (CeSIA),巴黎,法国)
专题命中 安全评测 :jailbreak(abstract);分类 cs.AI、cs.LG;trustworthy(comments)
AI总结 提出首个独立运营基准BELLS-O,评估28个LLM监督系统在检测率、误报率、延迟和成本上的权衡,发现专用护栏在内容审核中占优,而前沿通用模型在越狱检测中表现更好但成本更高。
Comments Accepted at the ICML 2026 Workshop on Trustworthy AI for Good (AI4GOOD). 2 figures; main text plus appendices