A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis
一种用于人工智能辅助鉴别诊断的面向安全的假设-演绎框架
Fan Ma, Mauro Giuffrè, Donald Wright, Kent McCann, Mark Iscoe, Lingfei Qian, Mingyang Jiang, Chi Wing Ng, Na Hong, Huan He, Cathy Shyr, Qingyu Chen, Lee Schwamm, Lucila Ohno-Machado, Hua Xu
机构
*
Yale School of Medicine, Yale University(耶鲁大学医学院,耶鲁大学)
;
Università degli Studi di Trieste(的里雅斯特大学)
;
Vanderbilt University(范德堡大学)
Comments7 pages, 2 figures, 5 tables. Oral paper at the 2nd Workshop on Epistemic Intelligence in Machine Learning (EIML@ICML 2026), Seoul, South Korea
Comments12 pages, 5 figures. has a same-size non-reasoning-teacher control, a three-judge LLM-as-a-judge panel with a negative control, full-source faithfulness grading, and a per-field routing analysis
CommentsAccepted for publication at the 2026 IEEE International Conference on Cyber Security and Resilience (IEEE CSR), Lisbon, Portugal, August 3-5, 2026. 8 pages, 1 figure
Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs
通过内部归因图对大语言模型越狱进行机制性可解释性研究
Anupam Wagle, Ifrat Ikhtear Uddin, Chaowei Zhang, Longwei Wang
机构
*
Department of Computer Science, University of South Dakota(南达科他大学计算机科学系)
;
School of Information and Artificial Intelligence, Yangzhou University(扬州大学信息与人工智能学院)