MedProbeBench: Systematic Benchmarking at Deep Evidence Integration for Expert-level Medical Guideline
MedProbeBench: 在深度证据整合中的系统性基准测试用于专家级医疗指南
机构 * Fudan University(复旦大学) ; Nanjing University(南京大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; University of Hong Kong(香港大学)
AI总结 本文提出MedProbeBench,首个利用高质量临床指南作为专家级参考的基准,通过1200+任务自适应评估标准和5130+原子性声明验证,评估17种LLM和深度研究代理在证据整合和指南生成中的能力差距。