Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth
重新思考文献检索评估:深度研究有帮助,且人类引用列表并非金标准
机构 * Mila – Quebec AI Institute(魁北克AI研究所) ; HEC Montréal(蒙特利尔HEC商学院) ; ServiceNow Research(ServiceNow研究) ; Canada CIFAR AI Chair(加拿大CIFAR人工智能主席) ; Université de Montréal(蒙特利尔大学) ; Polytechnique Montréal(蒙特利尔理工学院)
AI总结 本文通过改进检索流程和检验人类引用列表作为评估目标的可靠性,发现深度研究管道显著提升召回率,而人类引用中仅51%被判定为中等相关以上,建议采用多维度评估。