CommentsAfter further review, we identified substantive issues that materially affect the validity of the manuscript's core results and conclusions. Addressing these would require a fundamental reworking of the analysis and framing. To maintain the integrity of the public record, we request withdrawal of this version
Scaling Clinician-Grade Feature Generation from Clinical Notes with Multi-Agent Language Models
通过多智能体语言模型实现临床笔记中临床级特征生成的扩展
Jiayi Wang, Jacqueline Jil Vallon, Nikhil V. Kotha, Neil Panjwani, Xi Ling, Margaret Redfield, Sushmita Vij, Sandy Srinivas, John Leppert, Mark K. Buyyounouski, Mohsen Bayati
机构
*
Department of Management Science and Engineering, Stanford University School of Engineering(管理科学与工程系,斯坦福大学工程学院)
;
Department of Radiation Oncology, Stanford University School of Medicine(放射肿瘤学系,斯坦福大学医学院)
;
Operations, Information and Technology, Stanford University Graduate Business School(运营、信息与技术,斯坦福大学商学院)
;
Graduate Business School Research Hub, Stanford University Graduate Business School(商学院研究中心,斯坦福大学商学院)
;
Department of Medicine (Oncology), Stanford University School of Medicine(医学系(肿瘤学),斯坦福大学医学院)
;
Department of Medicine, Stanford University School of Medicine(医学系,斯坦福大学医学院)
;
Department of Urology, Stanford University School of Medicine(泌尿学系,斯坦福大学医学院)
;
Veterans Affairs Palo Alto Health Care System(退伍军人事务帕洛阿尔托医疗系统)
;
Department of Electrical Engineering, Stanford University School of Engineering(电气工程系,斯坦福大学工程学院)
Predicting LLM Correctness in Prosthodontics Using Metadata and Hallucination Signals
利用元数据和幻觉信号预测牙科修复学中大语言模型的正确性
Lucky Susanto, Anasta Pranawijayana, Cortino Sukotjo, Soni Prasad, Derry Wijaya
机构
*
1 Department of Data Science, Monash University Indonesia, Tangerang, Indonesia
;
2 Independent Researcher
;
3 Department of Prosthodontics, University of Pittsburgh, Pittsburgh, Pennsylvania
;
4 Department of Restorative Sciences, University of North Carolina Adams School of Dentistry, Chapel Hill, North Carolina
;
5 Department of Computer Science, Boston University, Boston, Massachusetts
CommentsAccepted at The 25th International Conference on Autonomous Agents and Multi-Agent Systems (AAMAS 2026). 21 pages including references, with 7 figures and 8 tables. Code is publicly available at the authors GitHub repository: https://github.com/S-Forouzandeh/MACLA-LLM-Agents-AAMAS-Conference
机构
*
Department of Computer Science(计算机科学系)
;
School of Computing and Information Systems(计算与信息学系)
;
Information Systems Technology and Design(信息系统技术与设计)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
思维链可监控性:AI安全的新且脆弱的机会
Tomek Korbak, Mikita Balesni, Elizabeth Barnes, Yoshua Bengio, Joe Benton, Joseph Bloom, Mark Chen, Alan Cooney, Allan Dafoe, Anca Dragan, Scott Emmons, Owain Evans, David Farhi, Ryan Greenblatt, Dan Hendrycks, Marius Hobbhahn, Evan Hubinger, Geoffrey Irving, Erik Jenner, Daniel Kokotajlo, Victoria Krakovna, Shane Legg, David Lindner, David Luan, Aleksander Mądry, Julian Michael, Neel Nanda, Dave Orr, Jakub Pachocki, Ethan Perez, Mary Phuong, Fabien Roger, Joshua Saxe, Buck Shlegeris, Martín Soto, Eric Steinberger, Jasmine Wang, Wojciech Zaremba, Bowen Baker, Rohin Shah, Vlad Mikulik
机构
*
UK AI Security Institute(英国人工智能安全研究所)
;
Apollo Research(阿波罗研究)
;
METR
;
University of Montreal(蒙特利尔大学)
;
Mila
;
Anthropic
;
OpenAI(开放人工智能研究所)
;
Google DeepMind(谷歌DeepMind)
;
Truthful AI
;
UC Berkeley(伯克利大学)
;
Center for AI Safety(人工智能安全中心)
;
AI Futures Project(人工智能未来项目)
;
Amazon(亚马逊)
;
Scale AI
;
Magic
;
Meta
;
Redwood Research(红木研究)
CommentsThe paper has been accepted to EMNLP 2024 (Main Conference) there is a follow up paper: Efficiently Selecting Response Generation Strategies for Synthetic Data Construction by Self-Aligned Perplexity Note: This is a revised version of arXiv:2402.11192 (v1, submitted 17 Feb 2024)
Dan Hendrycks, Dawn Song, Christian Szegedy, Honglak Lee, Yarin Gal, Erik Brynjolfsson, Sharon Li, Andy Zou, Lionel Levine, Bo Han, Jie Fu, Ziwei Liu, Jinwoo Shin, Kimin Lee, Mantas Mazeika, Long Phan, George Ingebretsen, Adam Khoja, Cihang Xie, Olawale Salaudeen, Matthias Hein, Kevin Zhao, Alexander Pan, David Duvenaud, Bo Li, Steve Omohundro, Gabriel Alfour, Max Tegmark, Kevin McGrew, Gary Marcus, Jaan Tallinn, Eric Schmidt, Yoshua Bengio
机构
*
Center for AI Safety(AI安全中心)
;
University of California, Berkeley(加州大学伯克利分校)
;
Virtue AI
;
Morph Labs(Morph实验室)
;
University of Michigan(密歇根大学)
;
LG AI Research(LG人工智能研究)
;
University of Oxford(牛津大学)
;
Stanford University(斯坦福大学)
;
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
Gray Swan AI
;
Carnegie Mellon University(卡内基梅隆大学)
;
Cornell University(康奈尔大学)
;
Hong Kong Baptist University(香港 Baptist大学)
;
HKUST(香港科技大学)
;
Nanyang Technological University(南洋理工大学)
;
KAIST(韩国科学技术院)
;
University of California, Santa Cruz(加州大学圣克鲁兹分校)
;
Massachusetts Institute of Technology(麻省理工学院)
;
University of Tübingen(图宾根大学)
;
University of Washington(华盛顿大学)
;
University of Toronto(多伦多大学)
;
Vector Institute(向量研究所)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Beneficial AI Research(有益AI研究)
;
Conjecture
;
Institute for Applied Psychometrics(应用心理测量研究所)
;
New York University(纽约大学)
;
CSER
;
Université de Montréal(蒙特利尔大学)
;
LawZero
CommentsEMNLP 2025 Workshop PALS. Additional note: There is a citation error on Evoke. The paper we are referring to is "Evoking critical thinking abilities in LLMs via reviewer-author prompt editing."