First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations
首先,不伤害:迈向临床安全的大语言模型
David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj, Priyank Jain, Jessica Tran, Matthew Gwiazdon, Arjun Rustagi, Jenelle Jindal, Jacob M. Koshy, Vinay Kadiyala, Anup Agarwal, Bassman Tappuni, Brianna French, Sirus Jesudasen, Christopher V. Cosgriff, Rebanta Chakraborty, Jillian Caldwell, Susan Ziolkowski, David J. Iberri, Robert Diep, Rahul S. Dalal, Kira L. Newman, Kristin Galetta, J. Carl Pallais, Nancy Wei, Kathleen M. Buchheit, David I. Hong, Vartan Pahalyants, Ernest Y. Lee, Allen Shih, Tamara B. Kaplan, Vishnu Ravi, Sarita Khemani, Thomas A. Buckley, April S. Liang, Daniel Shirvani, Advait Patil, Nicholas Marshall, Kanav Chopra, Joel Koh, Adi Badhwar, Anastasia Perez, Austin J. Schoeffler, Mahbuba Tusty, Chase M. Walton, Liam G. McCoy, David J. H. Wu, Yingjie Weng, Sumant Ranji, Kevin Schulman, Nigam H. Shah, Jason Hom, Arnold Milstein, Arjun K. Manrai, Adam Rodman, Jonathan H. Chen, Ethan Goh
机构
*
Harvard Combined Dermatology Program(哈佛联合皮肤科项目)
;
Department of Dermatology, Mass General Brigham(麻省总医院皮肤科)
;
Harvard Medical School(哈佛医学院)
;
Stanford Center for Biomedical Informatics Research(斯坦福生物医学信息学研究中心)
;
Stanford University(斯坦福大学)
;
Division of Hospital Medicine, Department of Medicine, Stanford University School of Medicine(斯坦福大学医学院医院医学科)
;
Department of Medicine, Cambridge Health Alliance(剑桥健康联盟医学科)
;
Beth Israel Deaconess Hospital–Plymouth(贝塞斯达德acons医院-普利茅斯)
;
Department of Medicine, University of California, San Francisco(加州大学旧金山分校医学科)
;
Department of Neurology, Stanford University School of Medicine(斯坦福大学医学院神经科)
;
Department of Medicine, Beth Israel Deaconess Medical Center(贝塞斯达德acons医学中心医学科)
;
Division of Cardiology, Department of Medicine, Cambridge Health Alliance(剑桥健康联盟心脏病科)
;
Department of Cardiovascular Medicine, Summa Health System(Summa健康系统心血管医学科)
;
Division of Allergy, Pulmonary, and Critical Care Medicine, Department of Medicine, University of Wisconsin-Madison(威斯康星大学麦迪逊分校医学科过敏、呼吸科和危重医学科)
;
Division of Pulmonary and Critical Care Medicine, Department of Medicine, Massachusetts General Hospital(麻省总医院呼吸科和危重医学科)
;
Center for Immunology and Inflammatory Diseases, Department of Medicine, Massachusetts General Hospital(麻省总医院免疫和炎症疾病中心)
;
Broad Institute of MIT and Harvard(MIT和哈佛Broad研究所)
;
Division of Pulmonary, Critical Care, and Sleep Medicine, Cambridge Health Alliance(剑桥健康联盟呼吸科、危重医学科和睡眠医学科)
Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap
迈向可信自主科学:两年社区路线图
Rafael Ferreira da Silva, Milad Abolhasani, Peter Beaucage, Laura Biven, Michael Bussmann, Kyle Chard, Ryan Coffee, Stephen DeWitt, Sagar Dolas, Carrie Eckert, David Elbert, Ian Foster, Tirthankar Ghosal, Anna Giannakou, Tom Gibbs, Leslie Hamilton, Glenn Lockwood, Theresa Mayer, Ben Mintz, Raffi Nazikian, Sal Nimer, Amanda Randles, Woong Shin, Sreenivas Rangan Sukumar, Frédéric Suter, Mitra Taheri, Michela Taufer, Draguna Vrabie
机构
*
U.S. Department of Energy, Office of Science, Office of Advanced Scientific Computing Research(美国能源部科学办公室高级科学计算研究办公室)
G-SHARE: A Guideline-Based Structured Reasoning Framework for Human-Factor Event Diagnosis
G-SHARE:一种基于指南的人为因素事件诊断结构化推理框架
Xingyu Xiao, Mao Du, Jiejuan Tong, Jingang Liang, Haitao Wang
机构
*
Institute of Nuclear and New Energy Technology, Tsinghua University(清华大学核能与新能源技术研究院)
;
National Key Laboratory of Human Factors Engineering(人因工程重点实验室)
;
Fujian Fuqing Nuclear Power Co., Ltd.(福建福清核电有限公司)
Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction
大语言模型能写出可靠的评分标准吗?实验复现的元评估
Hanhua Hong, Yizhi Li, Jiaoyan Chen, Luu Gia Huy, Sophia Ananiadou, Jung-jae Kim, Chenghua Lin
机构
*
The University of Manchester(曼彻斯特大学)
;
Institute for Infocomm Research (I²R), A*STAR(资讯通信研究院(I²R),新加坡科技研究局)
;
IQuest Research(IQuest研究公司)
;
ELLIS Manchester(ELLIS曼彻斯特)
;
University of Information Technology, VNU(越南国家大学信息技术大学)
CommentsThis preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution will be published in Computer Aided Systems Theory - EUROCAST 2026, Lecture Notes in Computer Science, Springer
Lost in Visual Translation: A VLM-Assisted Perceptual-Semantic Coherence Framework for EEG-to-Image Reconstruction
迷失在视觉翻译中:用于脑电到图像重建的基于视觉语言模型的感知语义连贯框架
Sukriti Tiwari, BHVSP Subrahmanyam, Nidhi Goyal, Sai Amrit Patnaik
机构
*
Mahindra University(马欣德拉大学)
;
MU-VT Interdisciplinary Advanced Research Centre for Transformative Technologies, Mahindra University(马欣德拉大学MU-VT变革性技术跨学科高级研究中心)
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
机构
*
Department of Computer Science, George Mason University(计算机科学系,乔治·马歇尔大学)
;
Department of Engineering Science, University of South Florida(工程科学系,佛罗里达州立大学)
;
Department of Computer Science, Rutgers University(计算机科学系,罗格斯大学)
SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation
SheetMind:一个由端到端大语言模型驱动的用于电子表格自动化的多智能体框架
Xi Cheng, Ruiyan Zhu, Ke Liu, Rakesh Chowdary Machineni, Lyuhao Chen, Brian Zhu, Daniel Jin, Zheng Qi, Neeraj Parihar, Zhoutian Xu, Oliver Gao
机构
*
Cornell University(康奈尔大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
University of Michigan(密歇根大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Hong Kong University of Science and Technology (GZ)(香港科学与技术大学)