机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Qatar Computing Research Institute(卡塔尔计算研究所)
;
Hamad Bin Khalifa University(哈马德·本·卡伊夫大学)
Commentsv2: substantially expanded and retitled. Adds unpublished results on the dynamic (self-modifying) case, deriving the persistence barrier from Rice's Theorem one level up; a supervisory-regress theorem linking the results to scalable oversight and Yampolskiy's verifier theory; and a unified treatment of all four barriers as one obstruction, the Expressivity Invariant
Positive Alignment: Artificial Intelligence for Human Flourishing
积极对齐:人工智能促进人类繁荣
Ruben Laukkonen, Seb Krier, Chloé Bakalar, Shamil Chandaria, Morten Kringelbach, Adam Elwood, Daniel Ford, Fernando Rosas, Maty Bohacek, Matija Franklin, Nenad Tomašev, Stephanie Chan, Verena Rieser, Roma Patel, Michael Levin, Arun Rao
机构
*
Department of Psychiatry, University of Oxford(牛津大学精神病学系)
;
Flourishing Intelligence Program, Centre for Eudaimonia and Human Flourishing, Linacre College, University of Oxford(牛津大学幸福智能计划、幸福与人类繁荣中心、林acre学院)
;
Google DeepMind(谷歌DeepMind)
;
LIFE
;
OpenAI
;
Anthropic
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
Aily Labs(Aily实验室)
;
Stanford University(斯坦福大学)
;
Tufts University(塔夫茨大学)
;
Positive AI Labs(积极AI实验室)
;
Department of Informatics, University of Sussex(Sussex大学信息学系)
;
Department of Brain Sciences, Imperial College London(伦敦帝国理工学院脑科学系)
Comments8 pages, 2 figures, 7 tables. Accepted at the ICML 2026 Mechanistic Interpretability Workshop and the ICML 2026 Failure Modes in Agentic AI Workshop
SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives
SteeringSafety:针对大语言模型在多安全视角下的表征引导基准测试
Vincent Siu, Nicholas Crispino, David Park, Nathan W. Henry, Zhun Wang, Yang Liu, Dawn Song, Chenguang Wang
机构
*
University of California, Santa Cruz(加州大学圣克鲁兹分校)
;
Washington University in St. Louis(华盛顿大学圣路易斯分校)
;
University of California, Berkeley(加州大学伯克利分校)
Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load
安全关键环境下符合欧盟AI法案要求的短期负荷预测:德国输电电网聚合负荷41天在线挑战赛结果
Thomas Bartz-Beielstein, Inalbek Akiev, Lalo Mohamad
Time-Series Forecasting in Safety-Critical Environments: An Open-Source Package for EU-AI-Act-Compliant Development / Zeitreihenprognose in sicherheitskritischen Umgebungen: Ein Open-Source-Paket für die KI-VO-konforme Entwicklung