Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation
黄金标准的幻觉:长文本生成中人类评估协议的大规模分析
机构 * University of Washington(华盛顿大学) ; National Tsing Hua University(国立清华大学) ; Seoul National University(首尔大学) ; Mila - Québec AI Institute(米拉-魁北克人工智能研究所) ; Allen Institute for AI(艾伦人工智能研究所)
AI总结 通过分析2023-2025年*CL会议论文中的人类评估协议,发现报告不透明和可重复性差的问题,并提出改进建议。
Comments Accepted to ACL 2026 Main