arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-04-28 至 2026-04-28 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 红队测试 3 篇

2604.23067 2026-04-28 cs.CR cs.CL 83%

Training a General Purpose Automated Red Teaming Model

训练通用自动化红队模型

Aishwarya Padmakumar, Leon Derczynski, Traian Rebedea, Christopher Parisien

机构 * NVIDIA

专题命中 红队测试 :red teaming(title,abstract);safety(abstract);分类 cs.CL

AI总结 本文提出一种通用红队模型训练方法,能适应任意对抗目标,无需依赖预训练评估器,通过微调小模型显著提升生成攻击的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10546 2026-04-28 cs.CL cs.AI cs.LG 78%

Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain

学习隐藏风险:面向金融领域的可控多轮红队测试框架

Gang Cheng, Haibo Jin, Wenbin Zhang, Haohan Wang, Jun Zhuang

机构 * Bloomberg(彭博社) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Florida International University(佛罗里达国际大学) Boise State University(博伊西州立大学)

专题命中 红队测试 :red teaming(title);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出CoRT框架,通过可控的多轮红队测试方法,针对金融领域潜在风险进行隐蔽攻击,提升LLM在监管合规方面的安全性。

Comments Accepted for ACL'26 (Main). TL;DR: We propose a controllable multi-turn risk-concealed red-teaming framework, CoRT, that progressively conceals surface-level risk while exploiting regulatory-violating behaviors on a proposed new benchmark, FinRisk-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23425 2026-04-28 cs.CR 67%

When the Agent Is the Adversary: Architectural Requirements for Agentic AI Containment After the April 2026 Frontier Model Escape

当代理成为对手:2026年4月前沿模型逃脱后代理AI约束的架构要求

Richard Joseph Mitchell

专题命中 红队测试 :alignment(abstract);safety(abstract)

AI总结 本文分析了四种现有约束方法的失效模式,提出五项架构要求,强调架构约束是应对代理AI安全威胁的唯一持久策略。

Comments 17 pages, 30 references, 5 tables. Derives five architectural requirements (R1-R5) for agentic AI containment from the April 2026 Mythos Preview incidents. Assesses AEGIS, Microsoft AGT, NVIDIA OpenShell, and other current systems; finds none satisfies all five requirements. Patent pending

详情

展开后加载摘要…

URL PDF HTML 收藏