arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 250 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 红队测试 250 篇

2305.17444 2023-05-30 cs.AI cs.CL cs.CR cs.LG 82%

Query-Efficient Black-Box Red Teaming via Bayesian Optimization

Deokjae Lee, JunYeong Lee, Jung-Woo Ha, Jin-Hwa Kim, Sang-Woo Lee, Hwaran Lee, Hyun Oh Song

专题命中 红队测试 :red teaming(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2023 Long Paper - Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.03286 2022-02-08 cs.CL cs.AI cs.CR cs.LG 82%

Red Teaming Language Models with Language Models

Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, Geoffrey Irving

专题命中 红队测试 :red teaming(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00290 2024-01-02 cs.CL cs.AI 82%

Red Teaming for Large Language Models At Scale: Tackling Hallucinations on Mathematics Tasks

Aleksander Buszydlik, Karol Dobiczek, Michał Teodor Okoń, Konrad Skublicki, Philip Lippmann, Jie Yang

专题命中 红队测试 :red teaming(title,abstract);分类 cs.CL、cs.AI;safety(comments)

Comments Accepted to The ART of Safety: Workshop on Adversarial testing and Red-Teaming for generative AI (IJCNLP-AACL 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04018 2026-08-06 cs.CY cs.AI 新提交 81%

Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming

管控智能体AI系统的执行风险:一种基于轨迹的红队测试框架

Zhihao Zhu, Yi Yang

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI、cs.CY

AI总结 本研究提出基于轨迹的红队测试框架TrajRed和运行时管控层TrajGuard,在AgentDojo实验中,TrajRed识别的漏洞更强,TrajGuard可将攻击成功率降至近零且保留任务效用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00801 2026-06-02 cs.CR cs.CL cs.ET cs.LG cs.NE 81%

Quality-Diversity Evolution for Discovering Diverse Vulnerabilities in LLM Safety

用于发现LLM安全中多样漏洞的质量-多样性进化

Subhadip Mitra

机构 * Rota Labs(Rota实验室)

专题命中 红队测试 :safety(title,abstract);分类 cs.CL、cs.LG

AI总结 提出基于质量-多样性进化框架(MAP-Elites)在语义层面生成可解释攻击策略,发现不同LLM的特定漏洞模式。

Comments 9 pages, 6 figures. Accepted at the ICLR 2026 Workshop on Agents in the Wild (AIWILD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10807 2026-03-12 q-fin.CP cs.AI cs.CY 81%

Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services

针对金融服务业LLM自动红队测试的风险调整危害评分

Fabrizio Dimino, Bhaskarjit Sarmah, Stefano Pasquali

专题命中 红队测试 :red teaming(title);jailbreak(abstract);分类 cs.AI、cs.CY

AI总结 本文提出了一种针对金融服务业LLM安全风险的评估框架,通过风险敏感的RAHS度量标准,评估LLM在长期对抗压力下的安全性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01908 2025-11-13 cs.CR cs.AI cs.LG 81%

UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning

Jiawei Zhang, Shuang Yang, Bo Li

机构 * Department of Computer Science, University of Chicago(芝加哥大学计算机科学系) Meta Department of Computer Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05538 2025-11-03 cs.AI cs.CR cs.CY 81%

Red Teaming AI Red Teaming

Subhabrata Majumdar, Brian Pendleton, Abhishek Gupta

机构 * Vijil / AI Risk and Vulnerability Alliance(Vijil / AI风险与漏洞联盟) AI Risk and Vulnerability Alliance(AI风险与漏洞联盟) Montreal AI Ethics Institute(蒙特利尔人工智能伦理研究所)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI、cs.CY

Comments Conference on Applied Machine Learning for Information Security (CAMLIS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20061 2025-10-24 cs.CY cs.AI cs.CR 81%

Ask What Your Country Can Do For You: Towards a Public Red Teaming Model

Wm. Matthew Kennedy, Cigdem Patlak, Jayraj Dave, Blake Chambers, Aayush Dhanotiya, Darshini Ramiah, Reva Schwartz, Jack Hagen, Akash Kundu, Mouni Pendharkar, Liam Baisley, Theodora Skeadas, Rumman Chowdhury

机构 * Oxford Internet Institute(牛津互联网研究所) University of Oxford(牛津大学) Amazon(亚马逊) Department of Computer Science(计算机科学系) University of Wisconsin - Eau Claire(威斯康星大学欧克莱尔分校) Carnegie Mellon University(卡内基梅隆大学) Harvard University(哈佛大学)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04451 2025-08-07 cs.LG cs.AI 81%

Automatic LLM Red Teaming

Roman Belaire, Arunesh Sinha, Pradeep Varakantham

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00467 2025-07-14 cs.CL cs.AI 81%

Red Teaming Large Language Models for Healthcare

Vahid Balazadeh, Michael Cooper, David Pellow, Atousa Assadi, Jennifer Bell, Mark Coatsworth, Kaivalya Deshpande, Jim Fackler, Gabriel Funingana, Spencer Gable-Cook, Anirudh Gangadhar, Abhishek Jaiswal, Sumanth Kaja, Christopher Khoury, Amrit Krishnan, Randy Lin, Kaden McKeen, Sara Naimimohasses, Khashayar Namdar, Aviraj Newatia, Allan Pang, Anshul Pattoo, Sameer Peesapati, Diana Prepelita, Bogdana Rakova, Saba Sadatamin, Rafael Schulman, Ajay Shah, Syed Azhar Shah, Syed Ahmar Shah, Babak Taati, Balagopal Unnikrishnan, Iñigo Urteaga, Stephanie Williams, Rahul G Krishnan

机构 * University of Toronto(多伦多大学) Vector Institute for AI(人工智能向量研究所) Johns Hopkins Medical Institutions(约翰霍普金斯医学机构) University Health Network(医疗网络) Algoma University(阿尔戈马大学) University of Iowa Hospitals & Clinics(爱荷华大学医院与诊所) University of Edinburgh(爱丁堡大学)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01931 2025-06-03 cs.CY cs.AI 81%

Red Teaming AI Policy: A Taxonomy of Avoision and the EU AI Act

Rui-Jie Yew, Bill Marino, Suresh Venkatasubramanian

机构 * Brown University(布朗大学) University of Cambridge(剑桥大学)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI、cs.CY

Comments Forthcoming at the 2025 ACM Conference on Fairness, Accountability, and Transparency

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18003 2025-05-26 cs.LG cs.AI 81%

An Example Safety Case for Safeguards Against Misuse

Joshua Clymer, Jonah Weinbaum, Robert Kirk, Kimberly Mai, Selena Zhang, Xander Davies

专题命中 红队测试 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10128 2024-10-08 cs.CL cs.AI 81%

Red Teaming Language Models for Processing Contradictory Dialogues

Xiaofei Wen, Bangzheng Li, Tenghao Huang, Muhao Chen

专题命中 红队测试 :red teaming(title,abstract);分类 cs.CL、cs.AI

Comments 20 pages, 5 figures, 11 tables. EMNLP2024 (main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02828 2024-10-07 cs.CR cs.AI cs.CL 81%

PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System

Gary D. Lopez Munoz, Amanda J. Minnich, Roman Lutz, Richard Lundeen, Raja Sekhar Rao Dheekonda, Nina Chikanov, Bolor-Erdene Jagdagdorj, Martin Pouliot, Shiven Chawla, Whitney Maxwell, Blake Bullwinkel, Katherine Pratt, Joris de Gruyter, Charlotte Siska, Pete Bryan, Tori Westerhoff, Chang Kawaguchi, Christian Seifert, Ram Shankar Siva Kumar, Yonatan Zunger

专题命中 红队测试 :red teaming(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07786 2024-09-12 cs.HC cs.AI cs.CY 81%

The Human Factor in AI Red Teaming: Perspectives from Social and Collaborative Computing

Alice Qian Zhang, Ryland Shaw, Jacy Reese Anthis, Ashlee Milton, Emily Tseng, Jina Suh, Lama Ahmad, Ram Shankar Siva Kumar, Julian Posada, Benjamin Shestakofsky, Sarah T. Roberts, Mary L. Gray

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI、cs.CY

Comments Updated with camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17576 2024-06-26 cs.CR cs.AI cs.LG 81%

Leveraging Reinforcement Learning in Red Teaming for Advanced Ransomware Attack Simulations

Cheng Wang, Christopher Redino, Ryan Clark, Abdul Rahman, Sal Aguinaga, Sathvik Murli, Dhruv Nandakumar, Roland Rao, Lanxiao Huang, Daniel Radke, Edward Bowen

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16247 2024-01-30 cs.CL cs.CY 81%

Towards Red Teaming in Multimodal and Multilingual Translation

Christophe Ropers, David Dale, Prangthip Hansanti, Gabriel Mejia Gonzalez, Ivan Evtimov, Corinne Wong, Christophe Touret, Kristina Pereyra, Seohyun Sonia Kim, Cristian Canton Ferrer, Pierre Andrews, Marta R. Costa-jussà

专题命中 红队测试 :red teaming(title,abstract);分类 cs.CL、cs.CY

Comments arXiv admin note: substantial text overlap with arXiv:2312.05187

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19713 2023-10-20 cs.CL cs.LG 81%

Red Teaming Language Model Detectors with Language Models

Zhouxing Shi, Yihan Wang, Fan Yin, Xiangning Chen, Kai-Wei Chang, Cho-Jui Hsieh

专题命中 红队测试 :red teaming(title);safety(abstract);分类 cs.CL、cs.LG

Comments Preprint. Accepted by TACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04425 2023-10-10 cs.CY cs.CR cs.ET cs.LG 81%

Red Teaming Generative AI/NLP, the BB84 quantum cryptography protocol and the NIST-approved Quantum-Resistant Cryptographic Algorithms

Petar Radanliev, David De Roure, Omar Santos

专题命中 红队测试 :red teaming(title,abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08877 2026-03-09 cs.LG 80%

Stress-Testing Alignment Audits With Prompt-Level Strategic Deception

通过提示级策略欺骗进行对齐审计压力测试

Oliver Daniels, Perusha Moodley, Benjamin M. Marlin, David Lindner

机构 * MATS University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Google Deepmind(谷歌DeepMind)

专题命中 红队测试 :alignment(title,abstract);分类 cs.LG;trustworthy(comments)

AI总结 本文提出自动红队流水线,通过生成针对白盒和黑盒审计方法的欺骗提示,揭示了当前对齐审计方法在面对强大非对齐模型时的脆弱性。

Comments Accepted at the ICLR 2026 Workshop on Principled Design for Trustworthy AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17259 2025-09-23 cs.AI 80%

Mind the Gap: Comparing Model- vs Agentic-Level Red Teaming with Action-Graph Observability on GPT-OSS-20B

Ilham Wicaksono, Zekun Wu, Rahul Patel, Theo King, Adriano Koshiyama, Philip Treleaven

机构 * University College London(伦敦大学学院) Holistic AI(整体AI)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI

Comments Winner of the OpenAI GPT-OSS-20B Red Teaming Challenge (Kaggle, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11516 2026-02-12 cs.LG cs.AI cs.CL 80%

Building Production-Ready Probes For Gemini

为Gemini构建生产级探针

János Kramár, Joshua Engels, Zheng Wang, Bilal Chughtai, Rohin Shah, Neel Nanda, Arthur Conmy

机构 * Google(谷歌)

专题命中 红队测试 :safety(abstract);red teaming(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出新的探针架构以应对长上下文分布变化,并通过实验验证其在网络安全领域的有效性,同时展示了自动化AI安全研究的初步成果。

Comments v4 (another minor acknowledgements fix)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10669 2026-08-12 cs.AI 新提交 79%

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

REDAgentBench:可执行的红队测试与LLM智能体系统的忠实测量

Zixing Chen, Xingyuan Liu, Jie Zhu, Huaixia Dou, Shuo Jiang, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang

专题命中 红队测试 :red teaming(title);safety(abstract);分类 cs.AI

AI总结 研究针对LLM智能体安全评估的缺陷,推出REDAgentBench框架,经实验发现其宏平均ASR为65.69%,还揭示了识别-执行差距,无训练策略提醒可减少70%以上违规

Comments 6 figures, 4 tables. Supplementary material included

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20852 2026-07-24 cs.AI 新提交 79%

Code Monitor Red Teaming for Public-Test-Passing Code

用于通过公开测试的代码的代码监控红队

Junchi Liao, Jiawen Deng, Fuji Ren

机构 * University of Electronic Science and Technology of China(电子科技大学)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI

AI总结 研究公开测试通过后代码隐藏错误监控问题,引入代码监控红队协议并实例化为代码监控基准。通过大量实验发现弱验证器虽能改进但仍易漏错,对抗性压力影响验证效果,GLM - 5.1验证器缩小部分差距,揭示了验证器故障与证据限制的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20408 2026-07-07 cs.CR cs.AI 新提交 79%

NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms

LLM智能体安全性、多轮红队测试、越狱基准、对抗鲁棒性、安全关键系统

Hanwool Lee, Dasol Choi, Bokyeong Kim, Haon Park, Seung Geun Kim

机构 * AIM Intelligence(AIM智能公司) KAERI(韩国原子能研究所)

专题命中 红队测试 :safety(title,abstract);分类 cs.AI

AI总结 提出NRT-Bench基准,通过模拟核电站控制室的多轮红队测试,评估LLM智能体在安全关键系统中的对抗鲁棒性,发现不同模型的漏洞几乎不重叠,且防御效果高度依赖模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01913 2026-07-03 cs.CY 新提交 79%

From Battlefield to Boardroom: Strategic Red Teaming as an Epistemic Governance Instrument in the Age of AI

从战场到董事会:战略红队作为人工智能时代的知识治理工具

Jeroen Janssen

专题命中 红队测试 :red teaming(title,abstract);分类 cs.CY

AI总结 提出战略红队作为董事会层级的治理工具,通过六组件模型测试AI战略假设,将不确定性转化为可审查对象。

Comments 7 pages, technical report; arXiv edition of Apparens working paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05952 2026-06-05 cs.RO cs.AI 79%

Learning of Robot Safety Policies via Adversarial Synthetic Scenarios

通过对抗性合成场景学习机器人安全策略

Nikolai Dorofeev, Alexey Odinokov, Rostislav Yavorskiy

机构 * National Research Institute of Automation and Applied Mathematics(国家自动化与应用数学研究所)

专题命中 红队测试 :safety(title,abstract);分类 cs.AI

AI总结 提出一个基于对抗性游戏的框架,通过红蓝两队对抗生成危险场景并迭代优化安全策略,以高效发现高风险边缘案例。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05101 2026-06-04 cs.SD cs.LG 79%

FoeGlass: Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors

FoeGlass: 简单的上下文学习足以对音频深度伪造检测器进行红队测试

Sepehr Dehdashtian, Jacob H Seidman, Vishnu N Boddeti, Gaurav Bharaj

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.LG

AI总结 提出FoeGlass,一种基于大语言模型上下文学习的黑盒自动红队方法,通过生成音频样本发现深度伪造检测器的盲点,将假阴性率降低高达94%。

Comments Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03134 2026-05-29 cs.CL 79%

The Anatomy of Conversational Scams: A Topic-Based Red Teaming Analysis of Multi-Turn Interactions in LLMs

对话式诈骗的剖析:基于主题的LLM多轮交互红队分析

Xiangzhe Yuan, Zhenhao Zhang, Haoming Tang, Siying Hu

机构 * Department of Computer Science, University of Iowa(爱荷华大学计算机科学系) Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系)

专题命中 红队测试 :red teaming(title);safety(abstract);分类 cs.CL

AI总结 通过LLM间模拟框架研究多轮社交工程对话中的对抗动态,分析攻击与防御策略,发现跨模型和跨语言的结果差异及策略转换的结构性变化。

详情

展开后加载摘要…

URL PDF HTML 收藏