Journal refAddepalli, S., Varun, Y., Suggala, A., Shanmugam, K., & Jain, P. (2025). Does safety training of LLMs generalize to semantically related natural prompts? In The Thirteenth International Conference on Learning Representations 2025
机构
*
SKLP, ICT, CAS & UCAS(SKLP、信息科技研究院、中国科学院及中国科学院大学)
;
University of Aberdeen(阿伯丁大学)
;
University of Leeds(利兹大学)
;
SKLP, ICT, CAS(SKLP、信息科技研究院、中国科学院)
机构
*
ELLIS Institute Tübingen(图宾根ELLIS研究所)
;
Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)
;
Intesa Sanpaolo(Intesa Sanpaolo公司)
;
University of Southern California(南加州大学)
Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing
通过语义触发和心理框架针对大推理模型的推理定向劫持攻击
Zehao Wang, Lanjun Wang
机构
*
College of Intelligence and Computing(智能与计算学院)
;
School of New Media and Communication(新媒体与传播学院)
;
Shanghai Key Laboratory of Data Science(上海数据科学 key laboratory)
机构
*
Deakin University(德克萨斯大学)
;
Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究院)
;
Shanghai Key Laboratory of Multimodal Embodied AI(上海多模态具身人工智能重点实验室)
;
City University of Hong Kong(香港城市大学)
;
The University of Melbourne(墨尔本大学)
;
Singapore Management University(新加坡管理大学)
;
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构
*
University of Science
;
United Arab Emirates University
;
Nanyang Technological University
;
Zayed University
;
Intelligent Science \& Technology Academy of CASIC