Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models
基准的基准:评估小型语言模型的自动化安全基准
机构 * University of Kansas(堪萨斯大学)
专题命中 评测与基准 :LLM(summary_cn,abstract);language model(title,abstract);small language model(title,abstract);SLM(abstract,abstract_cn)
AI总结 本研究通过评估5个基准套件在26个开源SLMs上的表现,发现现有以LLM为中心的安全基准无法可靠评估SLMs,模糊性会导致模型排名出现显著变化。
Comments This paper is accepted for publication at ESORICS 2026