Quantifying and Mitigating Premature Closure in Frontier LLMs
对前沿大语言模型中过早闭合进行量化与缓解
机构 * Department of Medicine, Stanford University(斯坦福大学医学系) ; Department of Biomedical Data Science, Stanford University(斯坦福大学生物医学数据科学系)
AI总结 本文研究了大语言模型在医疗任务中过早闭合的问题,通过评估五个前沿模型发现其在不确定情况下仍频繁给出答案,安全提示虽能减少错误,但仍有残留问题,需进一步验证医疗LLM是否能判断何时不应回答。
Comments 14 pages, 3 figures, 1 table