When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs
当技能与安全相遇:对技能合并大语言模型的自适应越狱鲁棒性进行基准测试与表征
Yu Ma, Hongli Shi, Jing Li, Xinran Xu, Weiwei Hou
机构
*
Google(谷歌公司)
;
University of New South Wales(新南威尔士大学)
;
University of Technology Sydney(悉尼科技大学)
;
Zhejiang University(浙江大学)
;
Australian National University(澳大利亚国立大学)
机构
*
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
;
Beijing Academy of Artificial Intelligence (BAAI)(北京智源人工智能研究院)
;
Beihang University(北京航空航天大学)
;
Eastern Institute of Technology, Ningbo(宁波东方理工大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Microsoft Research Asia (MSRA)(微软亚洲研究院)
CommentsEmail by the arXiv Support Team: "Dear Gjergji, Thank you for your patience. Your appeal was accepted. You are welcome to resubmit this work at your convenience to cs.CY (cs.AI, cs.HC) Regards, arXiv Support"
Journal refAI and Ethics (2026) 6:295, Springer Nature
RiskNet: A large-scale dataset of AI risk incidents from news with alignment and multi-dimensional annotations
RiskNet:一个来自新闻的大规模AI风险事件数据集,包含对齐和多维标注
Leihan Zhang, Wecheng Ye, Xianlong Ma, Haochuan Liu, Yang Li, Qianyu Zhang, Jinliang Chen, Qiang Yan
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Beijing Key Laboratory of Multimodal Data Intelligent Perception and Governance(多模态数据智能感知与治理北京市重点实验室)
The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification