arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 250 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 红队测试 250 篇

2209.02167 2023-10-17 cs.AI cs.CR cs.LG 76%

Red Teaming with Mind Reading: White-Box Adversarial Policies Against RL Agents

Stephen Casper, Taylor Killian, Gabriel Kreiman, Dylan Hadfield-Menell

专题命中 红队测试 :red teaming(title);分类 cs.AI、cs.LG

Comments Code is available at https://github.com/thestephencasper/lm_white_box_attacks

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10894 2023-09-25 cs.LG cs.AI cs.CV 76%

Red Teaming Deep Neural Networks with Feature Synthesis Tools

Stephen Casper, Yuxiao Li, Jiawei Li, Tong Bu, Kevin Zhang, Kaivalya Hariharan, Dylan Hadfield-Menell

专题命中 红队测试 :red teaming(title);分类 cs.AI、cs.LG

Comments In Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.07476 2022-08-17 cs.CR cs.AI cs.LG 76%

CTI4AI: Threat Intelligence Generation and Sharing after Red Teaming AI Models

Chuyen Nguyen, Caleb Morgan, Sudip Mittal

专题命中 红队测试 :red teaming(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22880 2026-05-25 cs.CL cs.AI cs.CY 75%

How Far Will They Go? Red-Teaming Online Influence with Large Language Models

它们会走多远?使用大型语言模型对在线影响力进行红队测试

Daniel C. Ruiz, Anna Serbina, Ashwin Rao, Emilio Ferrara, Luca Luceri

机构 * Information Sciences Institute University of Southern California(信息科学研究所 乌德穆尔特国立大学)

专题命中 红队测试 :alignment(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出一个红队测试框架,通过测量大型语言模型在争议话题上的政治观点表达范围(Overton Window),并量化简单自然语言越狱如何扩展该范围,发现开源LLM在政治表达上存在系统性不对称,且越狱效果因模型系列而异。

Comments 30 pages, 8 figures, submitted to COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16222 2025-06-12 cs.LG cs.AI cs.CL cs.CR 75%

An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks

Valentyn Boreiko, Alexander Panfilov, Vaclav Voracek, Matthias Hein, Jonas Geiping

机构 * University of Tübingen(图宾根大学) Tübingen AI Center(图宾根人工智能中心) Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所) ELLIS Institute Tübingen(图宾根ELLIS研究所)

专题命中 红队测试 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18212 2025-05-27 cs.CY cs.AI cs.CL 75%

Towards medical AI misalignment: a preliminary study

Barbara Puccio, Federico Castagna, Allan Tucker, Pierangelo Veltri

专题命中 红队测试 :jailbreak(abstract);red teaming(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15221 2024-09-05 cs.LG cs.CL cs.CR cs.CY 75%

LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Nathaniel Li, Ziwen Han, Ian Steneker, Willow Primack, Riley Goodside, Hugh Zhang, Zifan Wang, Cristina Menghini, Summer Yue

专题命中 红队测试 :jailbreak(abstract);red teaming(abstract);分类 cs.CL、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00069 2025-01-03 cs.CL cs.AI 74%

Adversarial Negotiation Dynamics in Generative Language Models

Arinbjörn Kolbeinsson, Benedikt Kolbeinsson

专题命中 红队测试 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI;red teaming(comments)

Comments Paper at NeurIPS 2024 Workshop on Red Teaming GenAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08876 2026-08-04 cs.LG 版本更新 74%

OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents

OTora:一种用于LLM代理推理层面拒绝服务攻击的统一红队框架

Xinyu Li, Ronghui Mu, Lin Li, Tianjin Huang, Gaojie Jin

机构 * Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系) Department of Computer Science, University of Oxford(牛津大学计算机科学系) Department of Mathematics and Computer Science, Eindhoven University of Technology(埃因霍温理工大学数学与计算机科学系)

专题命中 红队测试 :red teaming(title);分类 cs.LG

AI总结 OTora是首个统一的两阶段红队框架,用于实现推理层面拒绝服务攻击,通过优化对抗触发器和生成代理感知的推理负载,提升推理token数量和延迟,同时保持任务准确性。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22063 2026-07-03 cs.SE cs.AI 版本更新 74%

RedCoder: Automated Multi-Turn Red Teaming for Code LLMs

RedCoder: 面向代码大语言模型的自动化多轮红队测试

Wenjie Jacky Mo, Qin Liu, Xiaofei Wen, Dongwon Jung, Hadi Askari, Wenxuan Zhou, Zhe Zhao, Muhao Chen

机构 * University of California, Davis(加州大学戴维斯分校) University of Southern California(南加州大学)

专题命中 红队测试 :red teaming(title);分类 cs.AI

AI总结 提出RedCoder,一个通过多轮对话诱导代码大模型生成漏洞代码的自动化红队测试智能体,采用多智能体博弈生成原型对话和攻击策略库,并微调LLM作为骨干,实验表明其优于现有单轮和多轮方法。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19626 2025-03-26 cs.CR cs.LG 74%

Red Teaming with Artificial Intelligence-Driven Cyberattacks: A Scoping Review

Mays Al-Azzawi, Dung Doan, Tuomo Sipola, Jari Hautamäki, Tero Kokkonen

专题命中 红队测试 :red teaming(title);分类 cs.LG

Comments An earlier version first published in Good Practices and New Perspectives in Information Systems and Technologies (pp. 129-138), 2024 by Springer Nature

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09408 2026-06-09 cs.CY cs.AI cs.HC 新提交 73%

Can Data Work be Reparative?

数据工作能否具有修复性?

Srravya Chandhiramowuli, Ding Wang, Alex Taylor

机构 * University of Edinburgh(爱丁堡大学) Google Research(谷歌研究院)

专题命中 红队测试 :safety(abstract);red teaming(abstract);分类 cs.AI、cs.CY

AI总结 通过民族志研究,探讨公民科技倡议如何从女性主义视角协作构建安全数据集,旨在将数据工作重塑为修复与补救的场所,并分析其中遇到的挑战与张力。

Comments To be presented at ACM FAccT, Montréal, Canada, June 25 to June 28, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23694 2026-06-04 cs.AI cs.CL cs.CR 73%

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents

SafeSearch: 基于LLM的搜索代理的自动化红队测试

Jianshuo Dong, Sheng Guo, Hao Wang, Xun Chen, Zhuotao Liu, Tianwei Zhang, Ke Xu, Minlie Huang, Han Qiu

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 红队测试 :safety(abstract);prompt injection(abstract);分类 cs.CL、cs.AI

AI总结 提出SafeSearch自动化红队框架,系统评估基于LLM的搜索代理在五个风险类别中的安全性,发现GPT-4.1-mini在搜索工作流中攻击成功率高达90.5%,且常见防御措施效果有限。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17380 2026-05-19 cs.AI cs.CR cs.LG 73%

ADR: An Agentic Detection System for Enterprise Agentic AI Security

ADR:一种用于企业代理AI安全的代理检测系统

Chenning Li, Pan Hu, Justin Xu, Baris Ozbas, Olivia Liu, Caroline Van, Manxue Li, Wei Zhou, Mohammad Alizadeh, Pengyu Zhang, KK Sriramadhesikan, Ming Zhang

机构 * Uber

专题命中 红队测试 :red teaming(abstract);prompt injection(abstract);分类 cs.AI、cs.LG

AI总结 本文提出ADR系统,一种大规模、经过生产验证的企业框架,用于安全地管理通过模型上下文协议(MCP)运行的AI代理。该系统解决了三个关键问题:观测有限、鲁棒性不足和检测成本高,并通过三个组件实现了这些目标:ADR传感器、ADR探索器和ADR检测器。

Comments Accepted at MLSys 2026 (Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05682 2026-05-12 cs.HC cs.AI cs.CY 73%

PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI

PersonaTeaming: 支持基于人设的生成AI红队测试

Wesley Hanwen Deng, Mingxi Yan, Sunnie S. Y. Kim, Akshita Jha, Lauren Wilcox, Kenneth Holstein, Motahhare Eslami, Leon A. Gatys

机构 * Carnegie Mellon University(卡内基梅隆大学) Apple(苹果公司)

专题命中 红队测试 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

AI总结 本文提出PersonaTeaming方法,通过整合人设提升红队测试的自动化与人机协作能力,实验显示其在攻击成功率和提示多样性方面优于现有方法,并通过用户界面促进人机协同创新。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10675 2026-01-07 cs.RO cs.AI cs.CV cs.LG 73%

Evaluating Gemini Robotics Policies in a Veo World Simulator

在Veo世界模拟器中评估Gemini机器人策略

Gemini Robotics Team, Krzysztof Choromanski, Coline Devin, Yilun Du, Debidatta Dwibedi, Ruiqi Gao, Abhishek Jindal, Thomas Kipf, Sean Kirmani, Isabel Leal, Fangchen Liu, Anirudha Majumdar, Andrew Marmon, Carolina Parada, Yulia Rubanova, Dhruv Shah, Vikas Sindhwani, Jie Tan, Fei Xia, Ted Xiao, Sherry Yang, Wenhao Yu, Allan Zhou

机构 * Gemini Robotics Team(Gemini机器人团队) Google DeepMind(谷歌DeepMind)

专题命中 红队测试 :safety(abstract);red teaming(abstract);分类 cs.AI、cs.LG

AI总结 本研究提出基于Veo视频模型的生成式评估系统,用于在Veo世界模拟器中全面评估Gemini机器人策略的性能,包括名义性能、分布外泛化及安全约束测试。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19461 2025-08-28 cs.AI cs.CR cs.LG 73%

Reliable Weak-to-Strong Monitoring of LLM Agents

Neil Kale, Chen Bo Calvin Zhang, Kevin Zhu, Ankit Aich, Paula Rodriguez, Scale Red Team, Christina Q. Knight, Zifan Wang

机构 * Scale AI Carnegie Mellon University(卡内基梅隆大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 红队测试 :red teaming(abstract);prompt injection(abstract);分类 cs.AI、cs.LG

Comments 18 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18479 2025-06-13 cs.CR cs.AI cs.LG 73%

SoK: Watermarking for AI-Generated Content

Xuandong Zhao, Sam Gunn, Miranda Christ, Jaiden Fairoze, Andres Fabrega, Nicholas Carlini, Sanjam Garg, Sanghyun Hong, Milad Nasr, Florian Tramer, Somesh Jha, Lei Li, Yu-Xiang Wang, Dawn Song

机构 * University of California, Berkeley(加州大学伯克利分校) Columbia University(哥伦比亚大学) Cornell University(康奈尔大学) Anthropic Oregon State University(俄勒冈州立大学) Google DeepMind(谷歌DeepMind) ETH Zurich(苏黎世联邦理工学院) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Carnegie Mellon University(卡内基梅隆大学) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 红队测试 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments IEEE S&P 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15058 2024-04-24 cs.CY cs.AI 73%

A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI

Seliem El-Sayed, Canfer Akbulut, Amanda McCroskery, Geoff Keeling, Zachary Kenton, Zaria Jalan, Nahema Marchal, Arianna Manzini, Toby Shevlane, Shannon Vallor, Daniel Susser, Matija Franklin, Sophie Bridgers, Harry Law, Matthew Rahtz, Murray Shanahan, Michael Henry Tessler, Arthur Douillard, Tom Everitt, Sasha Brown

专题命中 红队测试 :red teaming(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17419 2026-03-19 cs.CR cs.AI 72%

Caging the Agents: A Zero Trust Security Architecture for Autonomous AI in Healthcare

将代理笼中:面向医疗AI自主系统的零信任安全架构

Saikat Maiti

机构 * VP of Trust, Commure(Commure 职务负责人) Founder & CEO, nFactor Technologies(nFactor Technologies 创始人及首席执行官)

专题命中 红队测试 :prompt injection(abstract,comments);red teaming(abstract);分类 cs.AI

AI总结 本文提出零信任安全架构,用于医疗AI自主代理,通过多层防御机制解决代理在医疗环境中存在的关键安全漏洞,包括凭证泄露、执行能力滥用等,并开源相关配置和工具。

Comments Keywords: agentic AI security, autonomous agents, healthcare cybersecurity, zero trust, prompt injection, HIPAA, Kubernetes security, OpenClaw

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05659 2026-08-07 cs.CR 新提交 71%

Breaking Customized LLMs for Coding: Automated Red Teaming for Instruction Backdoor Attacks

针对代码任务定制化大语言模型的突破:针对指令后门攻击的自动红队测试

Yuchen Chen, Wei Cheng, Yuan Xiao, Wising Sun, Chunrong Fang, Yang Liu, Zhenyu Chen, Baowen Xu

专题命中 红队测试 :red teaming(title)

AI总结 本文提出ARIA自动红队测试框架,可高效生成针对代码定制化LLM的隐蔽带后门指令,攻击成功率达0.945且规避检测能力强,性能优于现有基线攻击。

Comments Accepted to the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09185 2026-05-12 cs.CE 71%

AutoRedTrader: Autonomous Red Teaming of Trading Agents through Synthetic Misinformation Injection

AutoRedTrader: 通过合成虚假信息注入实现交易代理的自主红队测试

Zhiwei Liu, Yangyang Yu, Yupeng Cao, Yuechen Jiang, Haohang Li, Zhuoran Lu, Yuyan Wang, Yixiang Zheng, Xiaorui Guo, Calvin Yixiang Cheng, Sophia Ananiadou

专题命中 红队测试 :red teaming(title)

AI总结 本文提出AutoRedTrader框架,通过行为偏差操控、微小文本扰动和重写策略生成金融领域虚假信息,评估其对交易代理的影响及历史市场证据的稳定性。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04989 2026-04-08 cs.CR 71%

SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement

SkillAttack:通过攻击路径细化实现对代理技能的自动化红队测试

Zenghao Duan, Yuxin Tian, Zhiyi Yin, Liang Pang, Jingcheng Deng, Zihao Wei, Shicheng Xu, Yuyao Ge, Xueqi Cheng

专题命中 红队测试 :red teaming(title)

AI总结 本文提出SkillAttack框架,通过对抗性提示动态验证技能漏洞可利用性,实验显示其在对抗性技能和真实场景中均显著优于基线方法,揭示了即使善意技能在实际代理交互中也存在严重安全风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22489 2026-03-25 cs.CR cs.SE 71%

Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning

模型上下文协议威胁建模及针对提示注入的漏洞分析

Charoes Huang, Xin Huang, Ngoc Phu Tran, Amin Milani Fard

专题命中 红队测试 :prompt injection(title)

AI总结 本文基于STRIDE和DREAD框架对MCP五个关键组件进行威胁建模,发现工具污染是主要客户端漏洞,提出多层防御策略以提升AI代理生态的安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25894 2025-10-01 cs.SE 71%

Red Teaming Program Repair Agents: When Correct Patches can Hide Vulnerabilities

Simin Chen, Yixin He, Suman Jana, Baishakhi Ray

专题命中 红队测试 :red teaming(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19855 2025-04-30 cs.CR 71%

The Automation Advantage in AI Red Teaming

Rob Mulla, Ads Dawson, Vincent Abruzzon, Brian Greunke, Nick Landers, Brad Palm, Will Pearce

专题命中 红队测试 :red teaming(title)

Comments 15 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21735 2026-07-31 cs.AI cs.CR 版本更新 70%

What AI Red-Team Evaluations Can and Cannot Prove

人工智能红队评估能证明什么以及不能证明什么

Bandana Kaur

机构 * APIsec Research Labs(APIsec研究实验室)

专题命中 红队测试 :safety(abstract);red teaming(abstract);分类 cs.AI

AI总结 研究人工智能红队评估能证与不能证之事,通过定义证据上限确定界限,发现高于某危害率基准可证类别,低于则否,该界限不限于基准,审核评估套件发现当前基准对高频危害足够,对罕见灾难性不足。

Comments 21 pages, 4 figures, 5 tables. Code and data links provided in the manuscript. v2: corrected Figure 1(b); corrected required sample sizes in Table 4 and in Sections 4.2, 4.6 and 5.2, which had been rounded rather than taken to the ceiling; corrected the sample-size expression stated in Methods; minor corrections to Table 1 and the Figure 2 caption. No theorem, result or conclusion is affected

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18263 2026-07-28 cs.AI 版本更新 70%

Position: AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII)

立场:人工智能/机器学习领域对深度伪造的研究与人工智能生成的非自愿亲密图像(AIG-NCII)不一致

Li Qiwei, Wells Lucas Santo, Sarita Schoenebeck, Eric Gilbert

专题命中 红队测试 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 研究指出人工智能/机器学习领域对深度伪造的研究与AIG-NCII不一致,现有技术干预忽略AIG-NCII,现有干预只关注观众认知危害,忽略主体尊严危害,建议更新威胁模型,在安全研究中解决AIG-NCII,同时警告研究人员涉足该领域需谨慎并做好防护。

Comments ICML 2026. Outstanding position paper honorable mention

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08892 2026-06-19 cs.LG 新提交 70%

Diffuse AI Control on Fuzzy Tasks

模糊任务上的扩散AI控制

Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar, Joe Benton

机构 * Anthropic Fellows Program (via MATS)(Anthropic 研究员计划(通过 MATS)) EPFL(洛桑联邦理工学院) Redwood Research(红木研究) Anthropic

专题命中 红队测试 :safety(abstract);AI safety(abstract);分类 cs.LG

AI总结 针对AI在模糊任务上的长期扩散威胁,提出蓝队与红队对抗框架,通过弱模型评分训练强模型,并发现红队可利用多目标进化提示优化找到评分高但性能差的子版本行为,蓝队则通过对抗优化提升鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07833 2026-06-09 cs.CR cs.AI 新提交 70%

Beyond Pass/Fail: Using Process Mining to Understand How LLMs Resist (and Fail) Red Team Attacks

超越通过/失败:使用过程挖掘理解LLM如何抵抗(和失败)红队攻击

Zvi Topol

机构 * MuyVentive LLC

专题命中 红队测试 :jailbreak(abstract);red teaming(abstract);分类 cs.AI

AI总结 提出将过程挖掘应用于红队攻击轨迹,通过分析事件日志提取直接跟随图和状态转移矩阵,揭示GPT-OSS和Llama 3.3在防御结构上的差异,发现传统攻击成功率指标无法捕捉的模型防御模式。

详情

展开后加载摘要…

URL PDF HTML 收藏