Training a General Purpose Automated Red Teaming Model
训练通用自动化红队模型
机构 * NVIDIA
专题命中 红队测试 :red teaming(title,abstract);safety(abstract);分类 cs.CL
AI总结 本文提出一种通用红队模型训练方法,能适应任意对抗目标,无需依赖预训练评估器,通过微调小模型显著提升生成攻击的能力。
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
训练通用自动化红队模型
机构 * NVIDIA
专题命中 红队测试 :red teaming(title,abstract);safety(abstract);分类 cs.CL
AI总结 本文提出一种通用红队模型训练方法,能适应任意对抗目标,无需依赖预训练评估器,通过微调小模型显著提升生成攻击的能力。
学习隐藏风险:面向金融领域的可控多轮红队测试框架
机构 * Bloomberg(彭博社) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Florida International University(佛罗里达国际大学) ; Boise State University(博伊西州立大学)
专题命中 红队测试 :red teaming(title);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出CoRT框架,通过可控的多轮红队测试方法,针对金融领域潜在风险进行隐蔽攻击,提升LLM在监管合规方面的安全性。
Comments Accepted for ACL'26 (Main). TL;DR: We propose a controllable multi-turn risk-concealed red-teaming framework, CoRT, that progressively conceals surface-level risk while exploiting regulatory-violating behaviors on a proposed new benchmark, FinRisk-Bench
当代理成为对手:2026年4月前沿模型逃脱后代理AI约束的架构要求
专题命中 红队测试 :alignment(abstract);safety(abstract)
AI总结 本文分析了四种现有约束方法的失效模式,提出五项架构要求,强调架构约束是应对代理AI安全威胁的唯一持久策略。
Comments 17 pages, 30 references, 5 tables. Derives five architectural requirements (R1-R5) for agentic AI containment from the April 2026 Mythos Preview incidents. Assesses AEGIS, Microsoft AGT, NVIDIA OpenShell, and other current systems; finds none satisfies all five requirements. Patent pending