arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1717 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 1717 篇

2507.17922 2025-07-25 cs.LG cs.AI 62%

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models

Jessica Quaye, Charvi Rastogi, Alicia Parrish, Oana Inel, Minsuk Kahng, Lora Aroyo, Vijay Janapa Reddi

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17946 2025-07-10 cs.CR cs.AI cs.CL 62%

Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs

Shuai Zhao, Leilei Gan, Zhongliang Guo, Xiaobao Wu, Yanhao Jia, Luwei Xiao, Cong-Duy Nguyen, Luu Anh Tuan

机构 * Nanyang Technological University(南洋理工大学) Zhejiang University(浙江大学) University of St Andrews(圣安德鲁大学) East China Normal University(华东师范大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00406 2025-07-09 cs.AI cs.CL 62%

Agents Are All You Need for LLM Unlearning

Debdeep Sanyal, Murari Mandal

机构 * RespAI Lab, School of Computer Engineering, KIIT Bhubaneswar(RespAI实验室,计算机工程学院,KIIT大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16327 2025-07-09 cs.CR cs.AI cs.CL 62%

Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs

Rui Pu, Chaozhuo Li, Rui Ha, Zejian Chen, Litian Zhang, Zheng Liu, Lirong Qiu, Zaisheng Ye

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Hangzhou Dianzi University(杭州电子科技大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Fujian Cancer Hospital(福建癌症医院)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00239 2025-07-02 cs.CL cs.AI 62%

Linearly Decoding Refused Knowledge in Aligned Language Models

Aryan Shrivastava, Ari Holtzman

机构 * University of Chicago(芝加哥大学)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04202 2025-06-27 cs.CR cs.AI cs.LG 62%

TracLLM: A Generic Framework for Attributing Long Context LLMs

Yanting Wang, Wei Zou, Runpeng Geng, Jinyuan Jia

机构 * Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI、cs.LG

Comments To appear in USENIX Security Symposium 2025. The code and data are at: https://github.com/Wang-Yanting/TracLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01633 2025-06-26 cs.LG cs.AI 62%

Adversarial Reasoning at Jailbreaking Time

Mahdi Sabbaghi, Paul Kassianik, George Pappas, Yaron Singer, Amin Karbasi, Hamed Hassani

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted to the 42nd International Conference on Machine Learning (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22960 2025-06-23 cs.AI cs.LG 62%

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Yongjin Yang, Euiin Yi, Jongwoo Ko, Kimin Lee, Zhijing Jin, Se-Young Yun

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments Preprint, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13726 2025-06-17 cs.AI cs.CR cs.LG 62%

Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models

Arjun Krishna, Aaditya Rastogi, Erick Galinkin

机构 * University of Waterloo(滑铁卢大学) NVIDIA(NVIDIA公司)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to LLMSEC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00382 2025-06-06 cs.LG cs.CL 62%

Spectral Insights into Data-Oblivious Critical Layers in Large Language Models

Xuyuan Liu, Lei Hsiung, Yaoqing Yang, Yujun Yan

机构 * Dartmouth College(达特茅斯学院)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted by Findings of ACL2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01926 2025-06-04 cs.CL cs.AI 62%

Unnatural Languages Are Not Bugs but Features for LLMs

Keyu Duan, Yiran Zhao, Zhili Feng, Jinjie Ni, Tianyu Pang, Qian Liu, Tianle Cai, Longxu Dou, Kenji Kawaguchi, Anirudh Goyal, J. Zico Kolter, Michael Qizhe Shieh

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00548 2025-06-03 cs.CR cs.CL cs.LG 62%

Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities

Jiahui Geng, Thy Thy Tran, Preslav Nakov, Iryna Gurevych

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05962 2025-06-03 cs.CL cs.AI 62%

Effective faking of verbal deception detection with target-aligned adversarial attacks

Bennett Kleinberg, Riccardo Loconte, Bruno Verschuere

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to Legal and Criminological Psychology (author version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24232 2025-06-02 cs.CV cs.AI cs.CL 62%

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models

Haibo Jin, Peiyan Zhang, Peiran Wang, Man Luo, Haohan Wang

机构 * School of Information Sciences University of Illinois at Urbana-Champaign(信息科学学院伊利诺伊大学厄巴纳-香槟分校) Computer Science and Engineering HKUST(计算机科学与工程香港科技大学) Computer Science Department University of California, Los Angeles(计算机科学系加州大学洛杉矶分校) Research Scientist, Intel Labs(英特尔实验室研究员) School of Information Sciences University of Illinois Urbana-Champaign(信息科学学院伊利诺伊大学厄巴纳-香槟分校)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17541 2025-05-30 cs.AI cs.CL 62%

Dataset Featurization: Uncovering Natural Language Features through Unsupervised Data Reconstruction

Michal Bravansky, Vaclav Kubon, Suhas Hariharan, Robert Kirk

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16359 2025-05-30 cs.CL cs.AI 62%

Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context

Nilanjana Das, Edward Raff, Aman Chadha, Manas Gaur

机构 * University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校) Amazon Web Services(亚马逊网络服务)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: text overlap with arXiv:2407.14644

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10619 2025-05-29 cs.AI cs.CL cs.CR 62%

Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search

Andy Zhou, Ron Arel

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17040 2025-05-26 cs.LG cs.CL 62%

Generalizing Large Language Model Usability Across Resource-Constrained

Yun-Da Tsai

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.LG

Comments Doctoral disstertation

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14425 2025-05-21 cs.CL cs.AI cs.CR 62%

Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation

Shuai Zhao, Xiaobao Wu, Cong-Duy Nguyen, Yanhao Jia, Meihuizi Jia, Yichao Feng, Luu Anh Tuan

机构 * Nanyang Technological University(南洋理工大学) Northwest Normal University(西北师范大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04578 2025-05-08 cs.LG cs.AI 62%

Fight Fire with Fire: Defending Against Malicious RL Fine-Tuning via Reward Neutralization

Wenjun Cao

机构 * Independent Researcher(独立研究者)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.00108 2025-05-02 cs.CR cs.AI cs.CL 62%

LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystem

Hongyi Liu, Shaochen Zhong, Xintong Sun, Minghao Tian, Mohsen Hariri, Zirui Liu, Ruixiang Tang, Zhimeng Jiang, Jiayi Yuan, Yu-Neng Chuang, Li Li, Soo-Hyun Choi, Rui Chen, Vipin Chaudhary, Xia Hu

机构 * Rice University(里士大学) Case Western Reserve University(凯斯西储大学) University of Minnesota(明尼苏达大学) Rutgers University(罗格斯大学) Texas A&M University(德克萨斯阿姆大学) Samsung Electronics America(三星电子美国公司)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01094 2025-04-03 cs.SD cs.AI cs.CL cs.CR eess.AS 62%

Multilingual and Multi-Accent Jailbreaking of Audio LLMs

Jaechul Roh, Virat Shejwalkar, Amir Houmansadr

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 6 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21464 2025-03-28 cs.CL cs.AI cs.PF 62%

Harnessing Chain-of-Thought Metadata for Task Routing and Adversarial Prompt Detection

Ryan Marinelli, Josef Pichlmeier, Tamas Bisztray

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.06608 2025-03-25 cs.LG cs.AI cs.CR 62%

MF-CLIP: Leveraging CLIP as Surrogate Models for No-box Adversarial Attacks

Jiaming Zhang, Lingyu Qiu, Qi Yi, Yige Li, Jitao Sang, Changsheng Xu, Dit-Yan Yeung

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08226 2025-03-12 cs.CL cs.AI 62%

A Grey-box Text Attack Framework using Explainable AI

Esther Chiramal, Kelvin Soh Boon Kai

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19038 2025-03-11 cs.CL cs.LG 62%

DIESEL -- Dynamic Inference-Guidance via Evasion of Semantic Embeddings in LLMs

Ben Ganon, Alon Zolfi, Omer Hofman, Inderjeet Singh, Hisashi Kojima, Yuval Elovici, Asaf Shabtai

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14628 2025-02-21 cs.LG cs.CL 62%

PEARL: Towards Permutation-Resilient LLMs

Liang Chen, Li Shen, Yang Deng, Xiaoyan Zhao, Bin Liang, Kam-Fai Wong

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.LG

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09039 2025-01-17 cs.CR cs.AI cs.CY 62%

Playing Devil's Advocate: Unmasking Toxicity and Vulnerabilities in Large Vision-Language Models

Abdulkadir Erol, Trilok Padhi, Agnik Saha, Ugur Kursuncu, Mehmet Emin Aktas

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14795 2025-01-09 cs.CL cs.CR cs.LG 62%

Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models

Jiaming He, Wenbo Jiang, Guanyu Hou, Wenshu Fan, Rui Zhang, Hongwei Li

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.LG

Comments The paper has been accepted to AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11208 2024-10-30 cs.CR cs.AI cs.CL 62%

Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents

Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen, Jie Zhou, Xu Sun

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at NeurIPS 2024, camera ready version. Code and data are available at https://github.com/lancopku/agent-backdoor-attacks

详情

展开后加载摘要…

URL PDF HTML 收藏