Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
忠实性与安全性:在反事实医学证据下评估LLM行为
机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) ; Northeastern University(东北大学) ; MD Anderson Cancer Center(MD安德森癌症中心) ; Georgia Institute of Technology(佐治亚理工学院)
专题命中 幻觉与事实性 :safety(title,abstract);分类 cs.CL
AI总结 本文研究了在反事实医学证据下LLM的行为,构建了MedCounterFact数据集,发现模型在面对危险或不合理证据时仍提供自信回答,表明模型可能过度强调忠实性而忽视安全性。
Comments Accepted to Findings of ACL 2026