arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1717 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 1717 篇

2508.03125 2025-08-06 cs.CR cs.AI cs.MA 57%

Attack the Messages, Not the Agents: A Multi-round Adaptive Stealthy Tampering Framework for LLM-MAS

Bingyu Yan, Ziyi Zhou, Xiaoming Zhang, Chaozhuo Li, Ruilin Zeng, Yirui Qi, Tianbo Wang, Litian Zhang

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03110 2025-08-06 cs.CL 57%

Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation

Zizhong Li, Haopeng Zhang, Jiawei Zhang

机构 * University of California, Davis(加州大学戴维斯分校) University of Hawaii at Mānoa(夏威夷大学马诺阿分校)

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19880 2025-07-29 cs.CR cs.AI 57%

Trivial Trojans: How Minimal MCP Servers Enable Cross-Tool Exfiltration of Sensitive Data

Nicola Croce, Tobin South

机构 * Pivotal Research(Pivotal研究机构) Stanford University(斯坦福大学)

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.AI

Comments Abstract submitted to the Technical AI Governance Forum 2025 (https://www.techgov.ai/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18656 2025-07-28 cs.CV cs.LG 57%

ShrinkBox: Backdoor Attack on Object Detection to Disrupt Collision Avoidance in Machine Learning-based Advanced Driver Assistance Systems

Muhammad Zaeem Shahzad, Muhammad Abdullah Hanif, Bassem Ouni, Muhammad Shafique

机构 * eBRAIN Lab, New York University Abu Dhabi (NYUAD), UAE(eBRAIN实验室,纽约大学阿布扎赫尔分校(NYUAD),阿联酋) AI and Digital Science Research Center, Technology Innovation Institute (TII), Abu Dhabi, UAE(人工智能与数字科学研究中心,技术创新研究所(TII),阿布扎赫尔,阿联酋)

专题命中 越狱攻击 :safety(abstract);分类 cs.LG

Comments 8 pages, 8 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13761 2025-07-21 cs.CL 57%

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models

Palash Nandi, Maithili Joshi, Tanmoy Chakraborty

机构 * Department of Electrical Engineering(电气工程系) Indian Institute of Technology Delhi(印度理工学院德里)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11117 2025-07-16 cs.AI 57%

AI Agent Architecture for Decentralized Trading of Alternative Assets

Ailiya Borjigin, Cong He, Charles CC Lee, Wei Zhou

机构 * Centre for Sustainable Development, University of Newcastle (Australia), Singapore(可持续发展中心,新南威尔士大学(澳大利亚),新加坡)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

Comments 8 Pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06323 2025-07-10 cs.CR cs.AI 57%

Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms

Tarek Gasmi, Ramzi Guesmi, Ines Belhadj, Jihene Bennaceur

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10273 2025-07-09 cs.CR cs.AI cs.NI 57%

AttentionGuard: Transformer-based Misbehavior Detection for Secure Vehicular Platoons

Hexu Li, Konstantinos Kalogiannis, Ahmed Mohamed Hussain, Panos Papadimitratos

机构 * Networked Systems Security (NSS) Group\ Royal Institute of Technology Stockholm Sweden Networked Systems Security (NSS) Group\ Royal Institute of Technology

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

Comments Author's version; Accepted for presentation at the ACM Workshop on Wireless Security and Machine Learning (WiseML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02057 2025-07-04 cs.CR cs.AI 57%

MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation

Lu Yan, Zhuo Zhang, Xiangzhe Xu, Shengwei An, Guangyu Shen, Zhou Xuan, Xuan Chen, Xiangyu Zhang

机构 * Purdue University(普渡大学) Columbia University(哥伦比亚大学) Virginia Tech(弗吉尼亚理工大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00052 2025-07-02 cs.CV cs.AI 57%

VSF-Med:A Vulnerability Scoring Framework for Medical Vision-Language Models

Binesh Sadanandan, Vahid Behzadan

机构 * SAIL Lab, University of New Haven, West Haven, CT, USA(SAIL实验室,新罕布什尔大学,西哈文,康涅狄格州,美国)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21842 2025-06-30 quant-ph cs.CR cs.LG 57%

Adversarial Threats in Quantum Machine Learning: A Survey of Attacks and Defenses

Archisman Ghosh, Satwik Kundu, Swaroop Ghosh

机构 * Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.LG

Comments 23 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17318 2025-06-24 cs.CR cs.AI 57%

Context manipulation attacks : Web agents are susceptible to corrupted memory

Atharv Singh Patlan, Ashwin Hebbar, Pramod Viswanath, Prateek Mittal

机构 * princeton(普林斯顿大学)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11934 2025-06-17 cs.AI 57%

Stepwise Reasoning Error Disruption Attack of LLMs

Jingyu Peng, Maolin Wang, Xiangyu Zhao, Kai Zhang, Wanyu Wang, Pengyue Jia, Qidong Liu, Ruocheng Guo, Qi Liu

机构 * University of Science and Technology of China(中国科学技术大学) City University of Hong Kong(香港城市大学) Independent Researcher(独立研究者)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11521 2025-06-16 cs.CR cs.AI cs.MM 57%

Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models

Jinming Wen, Xinyi Wu, Shuai Zhao, Yanhao Jia, Yuwen Li

机构 * Jilin University(吉林大学) Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) Northeastern University(东北大学)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07948 2025-06-10 cs.LG cs.CR 57%

TokenBreak: Bypassing Text Classification Models Through Token Manipulation

Kasimir Schulz, Kenneth Yeung, Kieran Evans

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07987 2025-06-06 cs.AI 57%

Universal Adversarial Attack on Aligned Multimodal LLMs

Temurbek Rahmatullaev, Polina Druzhinina, Nikita Kurdiukov, Matvey Mikhalchuk, Andrey Kuznetsov, Anton Razzhigaev

机构 * AIRI MSU(莫斯科国立大学) HSE University(俄罗斯高等经济大学) Skoltech(斯克里普钦科技大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI

Comments Added benchmarks, baselines, author, appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02859 2025-06-04 cs.CR cs.AI 57%

ATAG: AI-Agent Application Threat Assessment with Attack Graphs

Parth Atulbhai Gandhi, Akansha Shukla, David Tayouri, Beni Ifland, Yuval Elovici, Rami Puzis, Asaf Shabtai

机构 * Dept. of Software and Information Systems Engineering(软件与信息系统工程系) Ben-Gurion University of the Negev(贝叶尔-加利利大学)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13715 2025-06-03 cs.CR cs.CY cs.HC 57%

Digital Deception: Generative Artificial Intelligence in Social Engineering and Phishing

Marc Schmitt, Ivan Flechais

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.CY

Comments Submitted to CHI 2024

Journal ref Artificial Intelligence Review, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23706 2025-05-30 cs.NI cs.AI cs.DC cs.IT eess.SP math.IT 57%

Distributed Federated Learning for Vehicular Network Security: Anomaly Detection Benefits and Multi-Domain Attack Threats

Utku Demir, Yalin E. Sagduyu, Tugba Erpek, Hossein Jafari, Sastry Kompella, Mengran Xue

机构 * Nexcepta Inc.(Nexcepta公司) RTX BBN Technologies

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16555 2025-05-30 cs.CL 57%

Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models

Yanxu Mao, Peipei Liu, Tiehan Cui, Zhaoteng Yan, Congying Liu, Datao You

机构 * School of Software, Henan University(河南大学软件学院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19773 2025-05-27 cs.CL cs.CR 57%

What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs

Sangyeop Kim, Yohan Lee, Yongwoo Song, Kimin Lee

机构 * Coxwave Seoul National University(首尔国立大学) Kyung Hee University(庆熙大学) KAIST(韩国科学技术院)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL

Comments Accepted by ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09546 2025-05-26 cs.RO cs.AI 57%

How Secure Are Large Language Models (LLMs) for Navigation in Urban Environments?

Congcong Wen, Jiazhao Liang, Shuaihang Yuan, Hao Huang, Geeta Chandra Raju Bethala, Yu-Shen Liu, Mengyu Wang, Anthony Tzes, Yi Fang

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16957 2025-05-23 cs.CR cs.AI 57%

Invisible Prompts, Visible Threats: Malicious Font Injection in External Resources for Large Language Models

Junjie Xiong, Changjia Zhu, Shuhang Lin, Chong Zhang, Yongfeng Zhang, Yao Liu, Lingyao Li

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13076 2025-05-20 cs.CR cs.AI 57%

The Hidden Dangers of Browsing AI Agents

Mykyta Mudryi, Markiyan Chaklosh, Grzegorz Wójcik

机构 * Polish-Japanese Academy of Information Technology(波兰-日本信息科技学院) University of the National Education Commission in Kraków(克拉科夫国家教育委员会大学) Maria Curie-Sklodowska University in Lublin(利沃夫玛丽亚·克里沃斯卡大学)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12567 2025-05-20 cs.CR cs.AI 57%

A Survey of Attacks on Large Language Models

Wenrui Xu, Keshab K. Parhi

机构 * Department of Electrical and Computer Engineering, University of Minnesota(电气与计算机工程系,明尼苏达大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09602 2025-05-15 cs.LG cs.CR 57%

Adversarial Suffix Filtering: a Defense Pipeline for LLMs

David Khachaturov, Robert Mullins

机构 * Department of Computer Science and Technology(计算机科学与技术系)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09500 2025-05-15 cs.LG 57%

Layered Unlearning for Adversarial Relearning

Timothy Qian, Vinith Suriyakumar, Ashia Wilson, Dylan Hadfield-Menell

机构 * MIT(麻省理工学院)

专题命中 越狱攻击 :alignment(abstract);分类 cs.LG

Comments 37 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01294 2025-05-13 cs.CL 57%

Endless Jailbreaks with Bijection Learning

Brian R. Y. Huang, Maximilian Li, Leonard Tang

机构 * Haize Labs(哈伊兹实验室)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16721 2025-05-05 cs.CV cs.AI 57%

Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks

Han Wang, Gang Wang, Huan Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20774 2025-05-01 cs.CR cs.AI 57%

Can We Trust Embodied Agents? Exploring Backdoor Attacks against Embodied LLM-based Decision-Making Systems

Ruochen Jiao, Shaoyuan Xie, Justin Yue, Takami Sato, Lixu Wang, Yixuan Wang, Qi Alfred Chen, Qi Zhu

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

Comments Accepted paper at ICLR 2025, 31 pages, including main paper, references, and appendix

详情

展开后加载摘要…

URL PDF HTML 收藏