The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
大语言模型在医疗诊断中的可靠性:一致性、可操纵性与情境感知能力的检验
专题命中 诊断辅助 :diagnosis(title,abstract)
AI总结 该研究检验Google Gemini 2.0 Flash与OpenAI ChatGPT-4o的医疗诊断可靠性,发现二者虽一致性达100%,但易受无关信息操纵,且情境响应存在差异,提示其医疗应用需临床监督与结构化保障。