arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1717 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 1717 篇

2602.17837 2026-02-23 cs.CR cs.CL cs.LG 62%

TFL: Targeted Bit-Flip Attack on Large Language Model

TFL:针对大语言模型的定向位翻转攻击

Jingkai Guo, Chaitali Chakrabarti, Deliang Fan

机构 * School of Electrical, Computer and Energy Engineering(电气、计算机与能源工程学院) Arizona State University(亚利桑那州立大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.LG

AI总结 TFL是一种新型定向位翻转攻击框架,通过精确操控LLM输出并减少对无关输入的影响,提升攻击的隐蔽性和针对性。

Comments 13 pages, 11 figures. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00038 2026-02-12 cs.CL cs.AI cs.CR 62%

from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors

从无害到有害:通过对抗隐喻 jailbreak 语言模型

Yu Yan, Sheng Sun, Zenghao Duan, Teli Liu, Min Liu, Zhiyi Yin, Jingyu Lei, Qi Li

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学) People’s Public Security University of China(中国人民公安大学) Tsinghua University(清华大学)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

AI总结 通过对抗隐喻诱导语言模型生成有害内容,提出AVATAR框架以提升jailbreak攻击效果。

Comments arXiv admin note: substantial text overlap with arXiv:2412.12145

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08859 2026-02-06 cs.CL cs.AI cs.CR 62%

Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models

模式增强的多轮劫持:在大语言模型中利用结构漏洞

Ragib Amin Nihal, Rui Wen, Kazuhiro Nakadai, Jun Sakuma

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出PE-CoA框架,通过五个对话模式构建多轮劫持攻击,揭示了大语言模型在不同危害类别下的漏洞及防御局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02395 2026-02-03 cs.LG cs.AI cs.CR cs.MA 62%

David vs. Goliath: Verifiable Agent-to-Agent Jailbreaking via Reinforcement Learning

大卫与歌利亚:通过强化学习实现可验证的代理到代理劫持

Samuel Nellessen, Tal Kachman

机构 * Department of Artificial Intelligence, Radboud University, Nijmegen, The Netherlands(人工智能系,拉德堡德大学,尼姆韦根,荷兰)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

AI总结 通过强化学习框架Slingshot,验证了Tag-Along攻击作为可验证的威胁模型,并展示了如何通过环境交互从现成模型中引发有效代理攻击。

Comments Under review. 8 main pages, 2 figures, 2 tables. Appendix included

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19375 2026-01-28 cs.LG cs.AI 62%

Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection

选择性引导:通过判别层选择实现的范数保持控制

Quy-Anh Dang, Chris Ngo

机构 * VNU University of Science(越南国家科学大学) Knovel Engineering Lab(Knovel工程实验室)

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 选择性引导通过判别层选择和范数保持旋转,实现对LLM行为的可控稳定修改。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03420 2026-01-08 cs.LG cs.AI cs.CR 62%

Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks

无需梯度或先验知识的LLM劫持:有效且可转移的攻击

Zhakshylyk Nurlanov, Frank R. Schmidt, Florian Bernard

机构 * Department of Visual Computing, University of Bonn(视觉计算系,波恩大学) Bosch Center for Artificial Intelligence(博世人工智能中心) University of Bonn(波恩大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

AI总结 RAILS通过无需梯度或先验知识的令牌级迭代优化,实现了对LLM的有效且可转移的对抗性攻击。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00867 2026-01-06 cs.CR cs.AI cs.CY cs.HC 62%

The Silicon Psyche: Anthropomorphic Vulnerabilities in Large Language Models

硅之心:大型语言模型中的人格化脆弱性

Giuseppe Canale, Kashyap Thimmaraju

机构 * CPF3.org(CPF3组织) Flowguard Institute(Flowguard研究所)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI、cs.CY

AI总结 本文提出LLMs继承人类心理脆弱性,通过心理测量框架揭示其对权威操纵等攻击的易受性,呼吁开发心理防火墙以保护AI安全。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16962 2025-12-22 cs.CR cs.AI cs.LG 62%

MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval

MemoryGraft: 通过中毒经验检索持续破坏LLM代理

Saksham Sahai Srivastava, Haoyu He

机构 * School of Computing University of Georgia(计算机学院 佐治亚大学)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI、cs.LG

AI总结 MemoryGraft 通过在LLM代理的长期记忆中植入恶意经验,实现对代理行为的持续破坏。

Comments 14 pages, 1 figure, includes appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15081 2025-12-18 cs.CR cs.AI cs.CL 62%

Quantifying Return on Security Controls in LLM Systems

对LLM系统中安全控制的回报进行量化

Richard Helder Moulton, Austin O'Brien, John D. Hastings

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种量化LLM系统安全控制回报的方法,通过模拟攻击和风险评估,比较了ABAC、NER删除和NeMo Guardrails三种安全措施的效果,发现ABAC显著降低风险,RoC达到9.83。

Comments 13 pages, 9 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16689 2025-12-16 cs.CL cs.AI 62%

Concept-Based Interpretability for Toxicity Detection

基于概念的可解释性用于毒性检测

Samarth Garg, Divya Singh, Deeksha Varshney, Mamta

机构 * ABV–IIITM(ABV-IIITM) IIT Patna(印度帕纳大学) IIT Jodhpur(印度朱道普尔大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出基于概念梯度的方法,通过分析有毒语言的归因机制,改进毒性检测的可解释性和准确性。

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17904 2025-12-16 cs.CR cs.AI cs.CL 62%

BreakFun: Jailbreaking LLMs via Schema Exploitation

BreakFun: 通过模式利用对LLM进行劫持

Amirkia Rafiei Oskooei, Mehmet S. Aktas

机构 * Department of Computer Engineering, Yildiz Technical University(计算机工程系,伊兹密尔技术大学)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

AI总结 BreakFun通过利用LLM对结构模式的遵守能力,揭示了其在劫持攻击中的脆弱性,并提出对抗性提示解构作为缓解策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12648 2025-11-18 cs.CR cs.AI cs.LG 62%

Scalable Hierarchical AI-Blockchain Framework for Real-Time Anomaly Detection in Large-Scale Autonomous Vehicle Networks

Rathin Chandra Shit, Sharmila Subudhi

机构 * organization= Dept. of Computer Science \& Engg., International Institute of Information Technology , city= Bhubaneswar , postcode= 751003 , state= Odisha , country= India organization= Dept. of Computer Science, Maharaja Sriram Chandra Bhanja Deo University , city= Baripada , postcode= 757003 , state= Odisha , country= India

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments Submitted to the Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10088 2025-11-14 cs.LG cs.AI cs.CV 62%

eXIAA: eXplainable Injections for Adversarial Attack

Leonardo Pesce, Jiawen Wei, Gianmarco Mengaldo

机构 * Department of Mechanical Engineering National University of Singapore(机械工程系国立新加坡大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08597 2025-11-13 cs.CL cs.AI 62%

Self-HarmLLM: Can Large Language Model Harm Itself?

Heehwan Kim, Sungjune Park, Daeseon Choi

机构 * Soongsil University(首尔大学)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03299 2025-11-10 cs.LG cs.CL cs.CV 62%

GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models

Haibo Jin, Ruoxi Chen, Peiyan Zhang, Andy Zhou, Haohan Wang

机构 * School of Information Sciences University of Illinois at Urbana-Champaign(信息科学学院伊利诺伊大学厄巴纳-香槟分校) Independent Researcher, Starc Institute(Starc研究所独立研究者) Computer Science and Engineering HKUST(HKUST计算机科学与工程学院) Computer Science Lapis Labs University of Illinois Urbana-Champaign(计算机科学Lapis Labs伊利诺伊大学厄巴纳-香槟分校) School of Information Sciences University of Illinois Urbana-Champaign(信息科学学院伊利诺伊大学厄巴纳-香槟分校)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.LG

Comments 28 papges

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19360 2025-10-21 cs.CL cs.AI 62%

Semantic Representation Attack against Aligned Large Language Models

Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang, Shaohui Mei, Lap-Pui Chau

机构 * Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(香港理工大学电子与电气工程系) School of Electronics and Information, Northwestern Polytechnical University(西北工业大学电子与信息学院)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19304 2025-10-21 cs.CY cs.AI cs.CE 62%

Epistemic Trade-Off: An Analysis of the Operational Breakdown and Ontological Limits of "Certainty-Scope" in AI

Generoso Immediato

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.CY

Comments Preprint V3 (October 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10601 2025-10-15 cs.CL cs.AI 62%

When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers

Divij Handa, Zehua Zhang, Amir Saeidi, Shrinidhi Kumbhar, Md Nayem Uddin, Aswin RRV, Chitta Baral

机构 * Arizona State University(亚利桑那州立大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

Comments Published in Reliable ML from Unreliable Data workshop @ NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21947 2025-10-07 cs.LG cs.AI 62%

Active Attacks: Red-teaming LLMs via Adaptive Environments

Taeyoung Yun, Pierre-Luc St-Charles, Jinkyoo Park, Yoshua Bengio, Minsu Kim

机构 * KAIST(韩国科学技术院) Mila – Québec AI Institute(魁北克人工智能研究所) LawZero Omelet Université de Montréal(蒙特利尔大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments 22 pages, 7 figures, 18 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13115 2025-10-01 cs.CL cs.AI cs.CR 62%

Dagger Behind Smile: Fool LLMs with a Happy Ending Story

Xurui Song, Zhixin Xie, Shuo Huai, Jiayi Kong, Jun Luo

机构 * S-Lab, Nanyang Technological University, Singapore(南洋理工大学新加坡分校S实验室) College of Computing and Data Science, Nanyang Technological University, Singapore(南洋理工大学计算机与数据科学学院)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17881 2025-09-30 cs.CL cs.AI 62%

GRAF: Multi-turn Jailbreaking via Global Refinement and Active Fabrication

Hua Tang, Lingyong Yan, Yukun Zhao, Shuaiqiang Wang, Jizhou Huang, Dawei Yin

机构 * The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen)(香港中文大学(深圳)) Baidu Inc.(百度公司)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19100 2025-09-24 cs.LG cs.AI 62%

Algorithms for Adversarially Robust Deep Learning

Alexander Robey

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17265 2025-09-23 cs.LG cs.AI 62%

SUA: Stealthy Multimodal Large Language Model Unlearning Attack

Xianren Zhang, Hui Liu, Delvin Ce Zhang, Xianfeng Tang, Qi He, Dongwon Lee, Suhang Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Amazon(亚马逊) University of Sheffield(谢菲尔德大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI、cs.LG

Comments EMNLP25

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16792 2025-09-23 cs.CL cs.AI 62%

MIST: Jailbreaking Black-box Large Language Models via Iterative Semantic Tuning

Muyang Zheng, Yuanzhi Yao, Changting Lin, Caihong Kai, Yanxiang Chen, Zhiquan Liu

机构 * School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) College of Cyber Security, Jinan University(暨南大学网络安全学院)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13266 2025-09-17 cs.LG cs.AI 62%

JANUS: A Dual-Constraint Generative Framework for Stealthy Node Injection Attacks

Jiahao Zhang, Xiaobing Pei, Zhaokun Zhong, Wenqiang Hao, Zhenghao Tang

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10931 2025-09-16 cs.AI cs.CL 62%

Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding

Seongho Joo, Hyukhun Koh, Kyomin Jung

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17674 2025-09-10 cs.CR cs.AI cs.LG 62%

Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models

Qiming Guo, Jinwen Tang, Xingran Huang

机构 * Department of Computer Science Texas A\&M University–Corpus Christi Corpus Christi, TX, USA EECS Department University of Missouri Columbia, MO, USA Department of Computer Engineering University of California–Riverside Riverside, CA, USA

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10722 2025-09-03 cs.CL cs.AI 62%

MEGen: Generative Backdoor into Large Language Models via Model Editing

Jiyang Qiu, Xinbei Ma, Zhuosheng Zhang, Hai Zhao, Yun Li, Qianren Wang

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) Key Laboratory of Shanghai Education Commission for Intelligent Interaction and Cognitive Engineering, Shanghai Jiao Tong University(上海交通大学智能交互与认知工程重点实验室) Shanghai Key Laboratory of Trusted Data Circulation and Governance in Web3(上海Web3可信数据流通与治理重点实验室) Cognitive AI Lab(认知人工智能实验室)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11727 2025-09-03 cs.CR cs.AI cs.CL cs.SE 62%

Efficient Detection of Toxic Prompts in Large Language Models

Yi Liu, Junzhe Yu, Huijia Sun, Ling Shi, Gelei Deng, Yuqi Chen, Yang Liu

机构 * Nanyang Technological University(南洋理工大学) ShanghaiTech University(上海科技大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted by the 39th IEEE/ACM International Conference on Automated Software Engineering (ASE 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04196 2025-08-07 cs.CL cs.AI cs.CR 62%

Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models

Siddhant Panpatil, Hiskias Dingeto, Haon Park

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏