XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
XL-SafetyBench: 一个基于国家的跨文化安全基准,用于LLM安全性和文化敏感性
Dasol Choi, Eugenia Kim, Jaewon Noh, Sang Seo, Eunmi Kim, Myunggyo Oh, Yunjin Park, Brigitta Jesica Kartono, Josef Pichlmeier, Helena Berndt, Sai Krishna Mendu, Glenn Johannes Tungka, Özlem Gökçe, Suresh Gehlot, Katherine Pratt, Amanda Minnich, Haon Park
机构
*
AIM Intelligence(AIM智能研究院)
;
Microsoft(微软公司)
;
Korea AISI(韩国人工智能研究所)
;
KT Corporation(KT公司)
;
BMW Group(宝马集团)
;
Coinbase(Coinbase公司)
;
Technical University of Munich(慕尼黑技术大学)
;
Ankara University(安卡拉大学)
;
Cyril Amarchand Mangaldas(Cyril Amarchand Mangaldas法律事务所)
;
Seoul National University(首尔国立大学)
Comments74 pages, 2 figures, 4 tables. Hybrid systematic survey and conceptual framework on LLM evaluation and AI-safety failures, synthesizing 373 primary studies (2018-2026). Introduces the EvalSafetyGap framework (Instability Decomposition, Alignment Trilemma) and reports an exploratory ten-model audit. Submitted as a review/survey article; not currently under consideration elsewhere
Causal Path Alignment: Anchoring the Optimization Trajectory for Controllable In-Parameter Knowledge Editing
因果路径对齐:为可控的参数知识编辑锚定优化轨迹
Xiyu Liu, Zhengxiao Liu, Naibin Gu, Zheng Lin, Weiping Wang
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
Deliberative Alignment: Reasoning Enables Safer Language Models
Melody Y. Guan, Manas Joglekar, Eric Wallace, Saachi Jain, Boaz Barak, Alec Helyar, Rachel Dias, Andrea Vallone, Hongyu Ren, Jason Wei, Hyung Won Chung, Sam Toyer, Johannes Heidecke, Alex Beutel, Amelia Glaese
机构
*
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
;
Beijing Academy of Artificial Intelligence (BAAI)(北京智源人工智能研究院)
;
Beihang University(北京航空航天大学)
;
Eastern Institute of Technology, Ningbo(宁波东方理工大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Microsoft Research Asia (MSRA)(微软亚洲研究院)
BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format
The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies
Sam Coggins, Alexander K. Saeri, Katherine A. Daniell, Lorenn P. Ruster, Jessie Liu, Jenny L. Davis
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
Frederik Pahde, Thomas Wiegand, Sebastian Lapuschkin, Wojciech Samek
机构
*
Fraunhofer Heinrich Hertz Institut(弗劳恩霍夫 Heinrich Hertz 研究所)
;
Technische Universität Berlin(柏林技术大学)
;
Berlin Institute for the Foundations of Learning and Data (BIFOLD)(柏林学习与数据基础研究所(BIFOLD))