Comments19 pages, 6 figures, 8 tables. Introduces the Deployment Wall framework, the Seam Index diagnostic instrument (with an evidence-anchored scoring protocol), and the Deployment Debt construct; includes six falsifiable propositions and a research agenda
Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support
现实世界临床护理中的推理:为何大型语言模型(LLM)尚不适合自主临床决策支持
Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem, Ryaan Sultan, Nicolas von Mallinckrodt, Max Solovyev, Alexey Matyushkin, Sumon Sadhu, Gabriele C DeLuca, Sanjeeva Jeyaretna, James Hillis, Manoj Ramachandran, Prakash Jayakumar
机构
*
Atman Labs(阿特曼实验室)
;
Oxford University Hospitals(牛津大学医院)
;
University of Oxford(牛津大学)
;
NIHR Oxford Biomedical Research Centre(英国国立卫生研究院牛津生物医学研究中心)
;
Imperial College London(伦敦帝国理工学院)
;
Technical University of Munich(慕尼黑工业大学)
;
Massachusetts General Hospital(麻省总医院)
;
Mass General Brigham(麻省总医院布里格姆医疗系统)
;
Harvard Medical School(哈佛医学院)
;
Barts Health NHS Trust(巴特保健国民保健信托基金会)
;
the University of Texas at Austin(德克萨斯大学奥斯汀分校)
The Cross-Domain Generalization Cost of Offensive Language Detection
攻击性语言检测的跨域泛化成本
Ruixing Ren, Junhui Zhao, Xiaoke Sun, Qiuping Li
机构
*
School of Electronic and Information Engineering, Beijing Jiaotong University(北京交通大学电子信息工程学院)
;
National Computer Network Emergency Response Technical Team/Coordination Center of China (CNCERT/CC)(国家计算机网络应急技术处理协调中心)