IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
IatroBench: AI安全措施中意外伤害的预注册证据
机构 * Harvard T.H. Chan School of Public Health(哈佛大学T.H. 洪学校公共卫生学院)
专题命中 安全训练 :safety(title,abstract);AI safety(title);分类 cs.CL、cs.AI、cs.CY
AI总结 该研究通过IatroBench评估了AI安全措施在医疗决策中的意外伤害风险,发现不同模型在身份相关性上的隐瞒行为存在显著差异,尤其在高度安全训练的模型中表现更明显。
Comments 30 pages, 3 figures, 11 tables. Pre-registered on OSF (DOI: 10.17605/OSF.IO/G6VMZ). Code and data: https://github.com/davidgringras/iatrobench. v2: Fix bibliography entries (add arXiv IDs, published venues); correct p-value typo in Limitations section; add AI Assistance Statement v3: Correct Figure 1 (decoupling scatter accidentally reverted to earlier draft in v2)