Benchmarking Multi-turn Medical Diagnosis: Hold, Lure, and Self-Correction
多轮医学诊断基准测试:Hold、Lure 和 Self-Correction
机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) ; New York University(纽约大学) ; The University of Texas Southwestern Medical Center(德克萨斯大学西南医学中心) ; University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
AI总结 本文通过MINT基准测试揭示了大语言模型在多轮证据积累中的行为模式,发现提前回答、自我修正和强诱因影响诊断决策,提出延迟提问和保留关键证据可提升诊断准确性。