Whose Gold? Annotator-Pool Disagreement Is Large at the Item Level, and Hidden by Small Leaderboards
谁的“金标准”?标注者群体在条目层面分歧巨大,却被小型排行榜掩盖
AI总结 该研究发现标注者群体在条目层面分歧巨大,小型模型排行榜的一致性是假象,量化了其脆弱性,证明某数据集的标注者无差异假设错误,LLM评判器更贴合大众标注者群体。
Comments Submitted to the HAIC workshop at NeurIPS 2026