Comments15 pages, 5 main figures, 6 appendix figures. Spotlight presentation at Forecasting as a New Frontier of Intelligence / Workshop on AI Forecasting, ICML 2026
PreAct-Bench: Benchmarking Predictive Monitoring in LLMs
PreAct-Bench:大语言模型中的预测性监控基准
Hainiu Xu, Italo Luis da Silva, Jiangnan Ye, Yuhao Wang, Wei Liu, Linyi Yang, Jonathan Richard Schwarz, Nicola Paoletti, Yulan He, Hanqi Yan
机构
*
King’s College London(伦敦国王学院)
;
National University of Singapore(新加坡国立大学)
;
Southern University of Science and Technology(南方科技大学)
;
Thomson Reuters Foundational Research(汤姆森路透基础研究)
;
Imperial College London(伦敦帝国学院)
;
The Alan Turing Institute(艾伦·图灵研究所)
机构
*
State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学)
;
iFLYTEK AI Research (Central China), iFLYTEK Co., Ltd(iFLYTEK中央中国AI研究院,iFLYTEK公司)
机构
*
College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院)
;
School of Data Science, Fudan University(复旦大学数据科学学院)
;
Ant Group(蚂蚁集团)
;
School of Information and School of Smart Governance, Renmin University of China(中国人民大学信息学院与智慧治理学院)
ResidencyRL: Reinforcement Learning in Simulated Clinical Environments
ResidencyRL:在模拟临床环境中开展的强化学习
Valentin Liévin, Samuel Schmidgall, Tim Strother, Alex Bijamov, Akshay Goel, Anil Palepu, Chunjong Park, Vahid Balazadeh, Min Woo Sun, Marius Guerard, Justin Chen, Dave Steiner, Vikram Dhillon, Ibrahim Azar, Akhil Mehta, Nicholas Spetsieris, Shilpan Shah, Maen Abdelrahim, Amit Dahiya, Yun Liu, Katherine Chou, Yossi Matias, Avinatan Hassidim, Dale R. Webster, Quoc V. Le, Raia Hadsell, Joelle Barral, Carey Radebaugh, Aleksandra Faust, Shekoofeh Azizi, Mike Schaekermann, Po-Hsuan Cameron Chen, Tao Tu, David Racz, Lin Yang
机构
*
Google DeepMind(谷歌DeepMind)
;
Google Research(谷歌研究院)
;
Houston Methodist Hospital(休斯顿卫理公会医院)
;
Trinity Health Group(三一健康集团)
;
Stanford Oncology Partners(斯坦福肿瘤学伙伴)
;
St. Luke Hospital(圣卢克医院)
专题命中
推理评测
:reasoning(abstract);分类 cs.CL、cs.AI
AI总结
本研究提出 ResidencyRL,通过多轮强化学习训练临床 AI 智能体,在模拟临床环境中提升诊断准确性、降低漏报率,且能力可迁移至多个医学基准测试,为临床 AI 发展提供了新路径。