Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control
面向AI安全评估的对抗语用学:指令冲突、嵌入命令与策略模糊性基准
机构 * Humber Polytechnic(汉博理工学院) ; University of Toronto(多伦多大学)
专题命中 安全评测 :safety(title,abstract);AI safety(title);分类 cs.CL、cs.AI
AI总结 提出对抗语用学基准和标注协议,通过语言学控制的分类法评估模型在指令冲突、嵌入命令等场景下的行为,为安全评估提供实证和方法论工具。
Comments 32-page main paper plus 13-page supplement; 6 figures and 17 tables total; code and data artifact available at the linked repository