机构
*
Department of Electrical Engineering, National Taiwan University(国立台湾大学电子工程系)
;
GARMIN (ASIA) CORPORATION(GARMIN(亚洲)公司)
;
Institute of Artificial Intelligence Innovation, National Yang Ming Chiao Tung University(国家阳明交通大学人工智能创新研究所)
;
Department of Information Management, National Dong Hwa University(国立东吴大学资讯管理系)
Anyone Can Jailbreak: Prompt-Based Attacks on LLMs and T2Is
Ahmed B Mustafa, Zihan Ye, Yang Lu, Michael P Pound, Shreyank N Gowda
机构
*
School of Computer Science, University of Nottingham(诺丁汉大学计算机科学学院)
;
Department of Intelligent Science, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学智能科学系)
;
School of Informatics, Xiamen University(厦门大学信息学院)
Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025
Zonghao Ying, Siyang Wu, Run Hao, Peng Ying, Shixuan Sun, Pengyu Chen, Junze Chen, Hao Du, Kaiwen Shen, Shangkun Wu, Jiwei Wei, Shiyuan He, Yang Yang, Xiaohai Xu, Ke Ma, Qianqian Xu, Qingming Huang, Shi Lin, Xun Wang, Changting Lin, Meng Han, Yilei Jiang, Siqi Lai, Yaozhi Zheng, Yifei Song, Xiangyu Yue, Zonglei Jing, Tianyuan Zhang, Zhilei Zhu, Aishan Liu, Jiakai Wang, Siyuan Liang, Xianglong Kong, Hainan Li, Junjie Mu, Haotong Qin, Yue Yu, Lei Chen, Felix Juefei-Xu, Qing Guo, Xinyun Chen, Yew Soon Ong, Xianglong Liu, Dawn Song, Alan Yuille, Philip Torr, Dacheng Tao
机构
*
Beihang University(北航)
;
Zhongguancun Laboratory(中关村实验室)
;
Hefei Comprehensive National Science Center(合肥综合性国家科学中心)
;
ETH Zürich(苏黎世联邦理工学院)
;
Pengcheng Laboratory(鹏城实验室)
;
Tsinghua University(清华大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
Johns Hopkins University(约翰霍普金斯大学)
;
University of Oxford(牛津大学)
;
Meta
;
Google Brain(谷歌脑)
;
Nanyang Technological University(南洋理工大学)
;
Aarhus University(哥本哈根大学)
;
China University of Mining(中国矿业大学)
;
University of Electronic Science(电子科学大学)
;
School of Electronic, Electrical(电子、电气与通信工程学院)
;
Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, CAS(智能信息处理国家重点实验室)
;
School of Computer Science(计算机科学学院)
;
Key Laboratory of Big Data Mining(大数据挖掘国家重点实验室)
;
Zhejiang Gongshang University(浙江工商大学)
;
Binjiang Institute of Zhejiang University(浙江大学滨江学院)
;
Zhejiang University(浙江大学)
;
The Chinese University of Hong Kong(香港中文大学)
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
Yang Yao, Xuan Tong, Ruofan Wang, Yixu Wang, Lujundong Li, Liang Liu, Yan Teng, Yingchun Wang
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
The University of Hong Kong(香港大学)
;
Fudan University(复旦大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
Taiye Chen, Zeming Wei, Ang Li, Yisen Wang
机构
*
School of EECS, Peking University(电子工程系,北京大学)
;
School of Mathematical Sciences, Peking University(数学系,北京大学)
;
State Key Lab of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(通用人工智能国家重点实验室,智能科学与技术学校,北京大学)
;
Institute for Artificial Intelligence, Peking University(人工智能研究院,北京大学)
机构
*
State Key Laboratory of AI Safety, Institute of Computing Technology, CAS(人工智能安全国家重点实验室,计算技术研究所,中国科学院)
;
School of Computer Science & Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学)
CommentsAccepted at ICLR 2025. Updates in the v3: GPT-4o and Claude 3.5 Sonnet results, improved writing. Updates in the v2: more models (Llama3, Phi-3, Nemotron-4-340B), jailbreak artifacts for all attacks are available, evaluation with different judges (Llama-3-70B and Llama Guard 2), more experiments (convergence plots, ablation on the suffix length for random search), examples of jailbroken generation
Journal refICLR 2025, BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks. In Proceedings of the International Conference on Learning Representations (ICLR), 2025