PanCanBench: A Comprehensive Benchmark for Evaluating Large Language Models in Pancreatic Oncology
PanCanBench: 用于评估大型语言模型在胰腺肿瘤学中的综合基准
机构 * Department of Biostatistics, University of Washington(华盛顿大学生物统计学系) ; Clinical Research Division, Fred Hutch Cancer Center(Fred Hutch癌症中心临床研究部) ; Division of Hematology and Oncology, Department of Medicine, University of Washington(华盛顿大学医学系血液学与肿瘤学分会) ; Allen Institute for AI(Allen人工智能研究所) ; Department of Computer Science and Engineering, University of Washington(华盛顿大学计算机科学与工程系) ; Public Health Sciences, Biostatistics, Fred Hutchinson Cancer Center(Fred Hutchinson癌症中心公共卫生科学与生物统计学)
专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.AI
AI总结 PanCanBench通过评估22种LLM在胰腺肿瘤学问题上的表现,揭示了模型在事实准确性、临床完整性和网络搜索整合方面的差异。