Faithful or Just Plausible? Evaluating the Faithfulness of Closed-Source LLMs in Medical Reasoning
忠实还是只是合理?评估闭源LLM在医学推理中的忠实性
Halimat Afolabi, Zainab Afolabi, Elizabeth Friel, Jude Roberts, Antonio Ji-Xu, Lloyd Chen, Egheosa Ogbomo, Emiliomo Imevbore, Phil Eneje, Wissal El Ouahidi, Aaron Sohal, Alisa Kennan, Shreya Srivastava, Anirudh Vairavan, Laura Napitu, Katie McClure
机构
*
Stratified Precision
;
Harvard Medical School(哈佛医学院)
;
Imperial College London(帝国理工学院伦敦分校)
;
National Health Service(国家健康服务系统)
;
Ipsen France(Ipsen法国)
;
University College London(伦敦大学学院)
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
MedHopQA: 一种以疾病为中心的多跳推理基准和评估框架,用于基于大语言模型的生物医学问答
Rezarta Islamaj, Robert Leaman, Joey Chan, Nicholas Wan, Qiao Jin, Natalie Xie, John Wilbur, Shubo Tian, Lana Yeganova, Po-Ting Lai, Chih-Hsuan Wei, Yifan Yang, Yao Ge, Qingqing Zhu, Zhizheng Wang, Zhiyong Lu
机构
*
National Library of Medicine, Division of Intramural Research(国家医学图书馆,院内研究部)
;
University of Illinois at Urbana-Champaign, Department of Computer Science(伊利诺伊大学厄巴纳-香槟分校计算机科学系)
;
University of Michigan Medical School(密歇根大学医学院)
CheXTemporal: A Dataset for Temporally-Grounded Reasoning in Chest Radiography
CheXTemporal:用于胸部X光影像中时间感知推理的数据集
Eva Prakash, Yunhe Gao, Chong Wang, Justin Xu, Neal Prakash, Arne Michalson, Seena Dehkharghani, Eun Kyoung Hong, Julie Bauml, Roger Boodoo, Jean-Benoit Delbrouck, Sophie Ostmeier, Curtis Langlotz
机构
*
Stanford University(斯坦福大学)
;
University of Oxford(牛津大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
HOPPR
;
University Hospital Zurich(苏黎世大学医院)
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院 Gallagher 学院)
;
King Abdullah University of Science and Technology(国王 Abdullah 科学技术大学)
;
Dongbei University of Finance and Economics(东北财经大学)
;
University of Science and Technology of China(中国科学技术大学)
PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoning
PlantMarkerBench: 一个多物种证据导向植物标记推理基准
Sajib Acharjee Dip, Song Li, Liqing Zhang
机构
*
Department of Computer Science, Virginia Tech(弗吉尼亚理工学院计算机科学系)
;
School of Plant and Environmental Sciences, Virginia Tech(弗吉尼亚理工学院植物与环境科学学院)
;
Fralin Biomedical Research Institute, Virginia Tech(弗吉尼亚理工学院弗拉林生物医学研究学院)
;
FBRI Cancer Research Center, Washington, DC(华盛顿特区FBRI癌症研究中心)
Comments46 pages including appendices (two-column preprint format). Under review at JAMIA. Code, frozen evaluator, and benchmark released at https://huggingface.co/datasets/alexstinard/epikg-clinicalbench. ClinicalBench v2 is a 400-question MIMIC-IV stress test for assertion-aware retrieval
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
ExploitGym:AI代理能否将安全漏洞转化为实际攻击?
Zhun Wang, Nico Schiller, Hongwei Li, Srijiith Sesha Narayana, Milad Nasr, Nicholas Carlini, Xiangyu Qi, Eric Wallace, Elie Bursztein, Luca Invernizzi, Kurt Thomas, Yan Shoshitaishvili, Wenbo Guo, Jingxuan He, Thorsten Holz, Dawn Song
机构
*
UC Berkeley(加州大学伯克利分校)
;
Max Planck Institute for Security and Privacy(马克斯·普朗克安全与隐私研究所)
;
UC Santa Barbara(加州大学圣巴巴拉分校)
;
Arizona State University(亚利桑那州立大学)
;
Anthropic(Anthropic公司)
;
OpenAI
;
Google(谷歌)
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
Stargazer:一种可扩展的AI代理在天体物理约束下的模型拟合基准环境
Xinge Liu, Terry Jingchen Zhang, Bernhard Schölkopf, Zhijing Jin, Kristen Menou
机构
*
University of Toronto(多伦多大学)
;
Vector Institute(向量研究所)
;
Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)
;
ELLIS Institute Tübingen(图宾根ELLIS研究所)
RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild
RW-Post:可审计的证据导向多模态事实核查
Danni Xu, Shaojing Fan, Harry Cheng, Mohan Kankanhalli
机构
*
School of Computing (SoC), National University of Singapore (NUS)(新加坡国立大学计算机学院(SoC))
;
National University of Singapore (NUS)(新加坡国立大学)
;
Department of Electrical and Computer Engineering (ECE), National University of Singapore (NUS)(新加坡国立大学电子与计算机工程系(ECE))
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
不要Pass@k:大规模语言模型评估的贝叶斯框架
Mohsen Hariri, Amirhossein Samandar, Michael Hinczewski, Vipin Chaudhary
机构
*
Department of Computer and Data Sciences, Case Western Reserve University(计算机与数据科学系,凯斯西储大学)
;
Department of Physics, Case Western Reserve University(物理系,凯斯西储大学)
Overview of the MedHopQA track at BioCreative IX: track description, participation and evaluation of systems for multi-hop medical question answering
BioCreative IX MedHopQA 轨道概述:轨道描述、参与及多跳医学问答系统评估
Rezarta Islamaj, Joey Chan, Robert Leaman, Jongmyung Jung, Hyeongsoon Hwang, Quoc-An Nguyen, Hoang-Quynh Le, Harikrishnan Gurushankar Saisudha, Ganesh Chandrasekar, Rustam R. Taktashov, Nadezhda Yu. Bizyukova, Sofia I. R. Conceição, Paulo R. C. Lopes, Reem Abdel Salam, Mary Adewunmi, Zhiyong Lu
机构
*
National Library of Medicine (NLM), National Institutes of Health (NIH)(美国国家医学图书馆(NLM)、国家卫生研究院(NIH))
;
University of Illinois at Urbana Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Korea University(韩国大学)
;
VNU University of Engineering and Technology, Hanoi, Vietnam(越南河内工程大学)
;
Concordia University, Montreal, QC, CA(蒙特利尔大学)
;
Institute of Biomedical Chemistry (IBMC), 10 bld. 8, Pogodinskaya str., 119121 Moscow, Russia(俄罗斯生物医学化学研究所(IBMC))
;
LASIGE, Departamento de Informática, Faculdade de Ciências, Universidade de Lisboa, 1749-016 Lisbon, Portugal(葡萄牙里斯本大学 LASIGE 实验室)
;
Faculty of Engineering, Computer Engineering Department Cairo University(埃及开罗大学工程学院)
;
Menzies School of Health Research, Charles Darwin University, NT, Australia(澳大利亚查尔斯达尔文大学梅恩兹健康研究中心)
;
CaresAI, Australia(澳大利亚 CaresAI)
专题命中
推理评测
:reasoning(abstract);分类 cs.CL
AI总结
本文介绍了BioCreative IX MedHopQA共享任务,旨在评估大型语言模型在多跳推理中的表现,通过构建1000个挑战性问题,展示了检索增强生成策略的重要性,并提供了公开数据集和评估结果。