Tail-Shape Estimation in LLM Evaluation Is Fragile: A Protocol for Diagnosing False Positives
LLM评估中的尾部形状估计是脆弱的:诊断假阳性的协议
机构 * Sapienza University of Rome(罗马大学)
专题命中 评测与基准 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG
AI总结 本文提出一个协议,用于检验LLM评估中尾部形状估计的假阳性,通过极值理论指标区分尾部重量和尾部质量,并在毒性评估中识别出三种假阳性模式。
Comments 9 pages of main paper, 4 figures and 4 tables in the main paper, more in the appendix