arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 250 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 红队测试 250 篇

2512.00412 2026-04-15 cs.CR cs.AI 79%

Red Teaming Large Reasoning Models

对大型推理模型进行红队测试

Jiawei Chen, Yang Yang, Chao Yu, Yu Tian, Zhi Cao, Xue Yang, Linghao Li, Hang Su, Zhaoxia Yin

机构 * Shanghai Key Laboratory of Multidimensional Information Processing, East China Normal University(上海多维信息处理关键实验室,东华大学) Zhongguancun Academy(中关村学院) Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院,清华大学) Dept. of Comp. Sci. and Tech., THBI Lab, Tsinghua University(计算机科学与技术系,清华THBI实验室) Beihang University(北京航空航天大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 红队测试 :red teaming(title);safety(abstract);分类 cs.AI

AI总结 本文提出RT-LRM基准,评估大型推理模型的可信度,揭示其在面对推理风险时的脆弱性,并发布工具箱支持未来研究。

Comments 30 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05379 2026-03-16 cs.CR cs.AI 79%

ThreatGPT: An Agentic AI Framework for Enhancing Public Safety through Threat Modeling

ThreatGPT: 一个通过威胁建模增强公共安全的代理AI框架

Sharif Noor Zisad, Ragib Hasan

机构 * Department of Computer Science(计算机科学系) University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)

专题命中 红队测试 :safety(title,abstract);分类 cs.AI

AI总结 ThreatGPT通过整合AI与人类判断,帮助工程师、安全官和政策制定者分析公共安全系统的威胁,利用STRIDE、MITRE ATT&CK等框架生成智能威胁模型,提升安全防护能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21267 2026-02-26 cs.CR cs.AI 79%

A Systematic Review of Algorithmic Red Teaming Methodologies for Assurance and Security of AI Applications

对AI应用保障和安全性的算法红队方法论系统综述

Shruti Srivastava, Kiranmayee Janardhan, Shaurya Jauhari

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI

AI总结 本文系统综述了自动化红队方法的现有研究,探讨其方法、工具、优势与局限,并指出未来改进方向,以提升AI应用的安全保障能力。

Comments 39 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15998 2025-11-25 cs.CR cs.AI 79%

Hiding in the AI Traffic: Abusing MCP for LLM-Powered Agentic Red Teaming

AI流量中的隐藏:利用MCP进行LLM驱动的代理红队攻击

Strahinja Janjusevic, Anna Baron Garcia, Sohrob Kazerounian

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI

AI总结 本文提出一种基于MCP的C2架构,用于LLM驱动的红队攻击,通过异步并行操作和实时情报共享,减少检测足迹并提升系统整体效能。

Comments 23 pages, 9 figures, 3 tables. Submitted as a full paper for review

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16466 2025-11-25 cs.CR cs.AI 79%

Incalmo: An Autonomous LLM-assisted System for Red Teaming Multi-Host Networks

Incalmo:一种自主的LLM辅助系统用于多主机网络红队测试

Brian Singer, Keane Lucas, Lakshmi Adiga, Meghna Jain, Lujo Bauer, Vyas Sekar

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI

AI总结 Incalmo是一种基于LLM的自主红队测试系统,通过高层次任务规划和专用任务代理实现多主机网络攻击,显著提高了红队测试的效率和成功率。

Comments 18 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22757 2025-09-30 cs.CR cs.AI cs.NI cs.SY eess.SY 79%

Red Teaming Quantum-Resistant Cryptographic Standards: A Penetration Testing Framework Integrating AI and Quantum Security

Petar Radanliev

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI

Journal ref The Journal of Defense Modeling and Simulation. 2025;0(0)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21011 2025-09-26 cs.CR cs.AI cs.SE 79%

Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools

Ping He, Changjiang Li, Binbin Zhao, Tianyu Du, Shouling Ji

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Palo Alto Networks(帕洛阿尔托网络公司) School of Software Technology, Zhejiang University(浙江大学软件技术学院)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22133 2025-07-31 cs.CR cs.CL 79%

Prompt Optimization and Evaluation for LLM Automated Red Teaming

Michael Freenor, Lauren Alvarez, Milton Leal, Lily Smith, Joel Garrett, Yelyzaveta Husieva, Madeline Woodruff, Ryan Miller, Erich Kummerfeld, Rafael Medeiros, Sander Schulhoff

机构 * Fuel iX Applied Research(Fuel iX应用研究) North Carolina State University(北卡罗来纳州立大学) University of Minnesota(明尼苏达大学) TELUS Digital(TELUS数字) Learn Prompting

专题命中 红队测试 :red teaming(title,abstract);分类 cs.CL

Comments 9 pages, 5 Figures, and 1 Appendix item

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06253 2025-06-19 cs.LG 79%

MAD-MAX: Modular And Diverse Malicious Attack MiXtures for Automated LLM Red Teaming

Stefan Schoepf, Muhammad Zaid Hameed, Ambrish Rawat, Kieran Fraser, Giulio Zizzo, Giandomenico Cornacchia, Mark Purcell

机构 * Univ. of Cambridge, Cambridge, UK.(剑桥大学) IBM Research Europe, Dublin, Ireland(IBM欧洲研究院)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.LG

Comments Data in Generative Models Workshop: The Bad, the Ugly, and the Greats (DIG-BUGS) at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10047 2025-06-13 cs.CR cs.CL 79%

GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models

Zilong Wang, Xiang Zheng, Xiaosen Wang, Bo Wang, Xingjun Ma, Yu-Gang Jiang

专题命中 红队测试 :red teaming(title);safety(abstract);分类 cs.CL

Comments 27 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04302 2025-06-06 cs.LG 79%

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming

Xiang Zheng, Xingjun Ma, Wei-Bin Lee, Cong Wang

机构 * City University of Hong Kong(香港城市大学) Fudan University(复旦大学) Hon Hai Research Institute(鸿海研究有限公司)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01278 2025-04-03 cs.AI 79%

Strategize Globally, Adapt Locally: A Multi-Turn Red Teaming Agent with Dual-Level Learning

Si Chen, Xiao Yu, Ninareh Mehrabi, Rahul Gupta, Zhou Yu, Ruoxi Jia

专题命中 红队测试 :red teaming(title);jailbreak(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.06237 2024-12-12 cs.CL cs.CR cs.HC 79%

Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming

Nanna Inie, Jonathan Stray, Leon Derczynski

专题命中 红队测试 :red teaming(title,abstract);分类 cs.CL

Journal ref PLoS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16738 2024-10-23 cs.LG 79%

LLM-Assisted Red Teaming of Diffusion Models through "Failures Are Fated, But Can Be Faded"

Som Sagar, Aditya Taparia, Ransalu Senanayake

专题命中 红队测试 :red teaming(title);alignment(abstract);分类 cs.LG

Comments 13 pages, 11 figures. arXiv admin note: substantial text overlap with arXiv:2406.07145

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10701 2024-08-21 cs.CL 79%

Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique

Tej Deep Pala, Vernon Y. H. Toh, Rishabh Bhardwaj, Soujanya Poria

专题命中 红队测试 :red teaming(title);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08017 2024-03-18 cs.CV cs.AI 79%

Red Teaming Models for Hyperspectral Image Analysis Using Explainable AI

Vladimir Zaigrajew, Hubert Baniecki, Lukasz Tulczyjew, Agata M. Wijata, Jakub Nalepa, Nicolas Longépé, Przemyslaw Biecek

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI

Comments 14 pages, 9 figures, ICLR 2024 Machine Learning for Remote Sensing (ML4RS) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11500 2023-12-20 cs.CR cs.AI 79%

A Red Teaming Framework for Securing AI in Maritime Autonomous Systems

Mathew J. Walter, Aaron Barrett, Kimberly Tam

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12867 2023-05-30 cs.CL cs.SE 79%

Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity

Terry Yue Zhuo, Yujin Huang, Chunyang Chen, Zhenchang Xing

专题命中 红队测试 :red teaming(title,abstract);分类 cs.CL

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20910 2025-04-30 cs.CY cs.AI cs.HC 79%

When Testing AI Tests Us: Safeguarding Mental Health on the Digital Frontlines

Sachin R. Pendse, Darren Gergle, Rachel Kornfield, Jonah Meyerhoff, David Mohr, Jina Suh, Annie Wescott, Casey Williams, Jessica Schleider

机构 * Northwestern University(西北大学) Microsoft Research(微软研究院) Williams Research Consulting(威廉斯研究咨询)

专题命中 红队测试 :safety(abstract);red teaming(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments Accepted to ACM Conference on Fairness, Accountability, and Transparency (FAccT 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06140 2026-06-05 cs.CR 78%

RedEdit: Agentic Red-Teaming of Image Safety Classifiers via MCTS-Guided Photo-Editing

RedEdit: 基于MCTS引导的照片编辑的图像安全分类器智能红队测试

Weilin Lin, Ziqi Lin, Zhenxing Zhou, Jianze Li, Tong Zhang, Hui Xiong, Li Liu

专题命中 红队测试 :safety(title,abstract)

AI总结 提出RedEdit,一种基于视觉语言模型提议和蒙特卡洛树搜索规划的黑盒红队代理,通过组合编辑工具序列使不安全图像逃避检测,平均不到两次编辑即可使76.2%的不安全图像绕过分类器,同时保留93.0%的恶意语义。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17075 2026-05-19 cs.CR 78%

A Red Teaming Framework for Evaluating Robustness of AI-enabled Security Orchestration, Automation, and Response Systems

一种用于评估AI增强的安全编排、自动化和响应系统鲁棒性的红队框架

Ayan Javeed Shaikh, Nathaniel D. Bastian, Ankit Shah

专题命中 红队测试 :red teaming(title,abstract)

AI总结 本文提出了一种结合大语言模型和强化学习的自主红队框架,用于评估企业网络中自主防御系统的鲁棒性,展示了单一LLM代理在多阶段攻击中的不足以及领域特定安全模型的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10546 2026-04-28 cs.CL cs.AI cs.LG 78%

Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain

学习隐藏风险:面向金融领域的可控多轮红队测试框架

Gang Cheng, Haibo Jin, Wenbin Zhang, Haohan Wang, Jun Zhuang

机构 * Bloomberg(彭博社) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Florida International University(佛罗里达国际大学) Boise State University(博伊西州立大学)

专题命中 红队测试 :red teaming(title);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出CoRT框架,通过可控的多轮红队测试方法,针对金融领域潜在风险进行隐蔽攻击,提升LLM在监管合规方面的安全性。

Comments Accepted for ACL'26 (Main). TL;DR: We propose a controllable multi-turn risk-concealed red-teaming framework, CoRT, that progressively conceals surface-level risk while exploiting regulatory-violating behaviors on a proposed new benchmark, FinRisk-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08638 2025-09-11 cs.RO 78%

AutoODD: Agentic Audits via Bayesian Red Teaming in Black-Box Models

Rebecca Martin, Jay Patrikar, Sebastian Scherer

机构 * Carnegie Mellon University(卡内基梅隆大学) Field AI

专题命中 红队测试 :red teaming(title);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12504 2025-08-22 cs.HC 78%

Organization Matters: A Qualitative Study of Organizational Dynamics in Red Teaming Practices for Generative AI

Bixuan Ren, EunJeong Cheon, Jianghui Li

专题命中 红队测试 :red teaming(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10603 2025-05-01 cs.CR 78%

Demo: ViolentUTF as An Accessible Platform for Generative AI Red Teaming

Tam n. Nguyen

专题命中 红队测试 :red teaming(title,abstract)

Comments 3 pages, 1 figure, 1 table. This is a demo paper for CyberWarrior2025. The video demo is at https://youtu.be/c-UCYXq0rfY. Codes will be shared when the competition concludes in June 2025 due to embargo requirements

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.09442 2023-10-12 cs.CL cs.AI cs.LG 78%

Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Stephen Casper, Jason Lin, Joe Kwon, Gatlen Culp, Dylan Hadfield-Menell

专题命中 红队测试 :red teaming(title);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31639 2026-07-01 cs.CR cs.AI cs.GT cs.LO 新提交 77%

A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems

大语言模型漏洞的生命周期与应用栈调查:攻击、风险、防御与开放问题

Seyed Bagher Hashemi Natanzi, Bo Tang

机构 * Worcester Polytechnic Institute(伍斯特理工学院)

专题命中 红队测试 :alignment(abstract);safety(abstract);red teaming(abstract);分类 cs.AI

AI总结 本文通过生命周期与应用栈视角系统化梳理大语言模型系统的漏洞,涵盖数据收集、预训练、对齐、供应链、检索、推理、工具执行和部署等阶段,分析攻击、风险与防御,并提出安全研究议程。

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00322 2024-07-30 cs.CL cs.GT 77%

Evolving Diverse Red-team Language Models in Multi-round Multi-agent Games

Chengdong Ma, Ziran Yang, Hai Ci, Jun Gao, Minquan Gao, Xuehai Pan, Yaodong Yang

专题命中 红队测试 :alignment(abstract);safety(abstract);harmlessness(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15398 2024-09-25 cs.CR cs.AI cs.LG 76%

Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI

Ambrish Rawat, Stefan Schoepf, Giulio Zizzo, Giandomenico Cornacchia, Muhammad Zaid Hameed, Kieran Fraser, Erik Miehling, Beat Buesser, Elizabeth M. Daly, Mark Purcell, Prasanna Sattigeri, Pin-Yu Chen, Kush R. Varshney

专题命中 红队测试 :red teaming(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11079 2023-10-18 cs.CL cs.AI 76%

Learning from Red Teaming: Gender Bias Provocation and Mitigation in Large Language Models

Hsuan Su, Cheng-Chu Cheng, Hua Farn, Shachi H Kumar, Saurav Sahay, Shang-Tse Chen, Hung-yi Lee

专题命中 红队测试 :red teaming(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏