arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-03-03 至 2026-03-03 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 提示注入 4 篇

2506.02456 2026-03-03 cs.AI cs.CR 79%

VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents

VPI-Bench:用于计算机使用代理的视觉提示注入攻击

Tri Cao, Bennett Lim, Yue Liu, Yuan Sui, Yuexin Li, Shumin Deng, Lin Lu, Nay Oo, Shuicheng Yan, Bryan Hooi

机构 * National University of Singapore(国立新加坡大学) Cyber Emerging Tech and R&D Co(网络安全与研发公司)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

AI总结 本文提出VPI-Bench基准,用于评估计算机使用代理和浏览器使用代理在视觉提示注入攻击下的鲁棒性,发现当前代理在某些平台上被欺骗率达51%和100%。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04281 2026-03-03 cs.CY cs.CR 74%

Prompt Injection Vulnerability of Consensus Generating Applications in Digital Democracy

共识生成应用在数字民主中的提示注入漏洞

Jairo Gudiño-Rosero, Clément Contet, Umberto Grandi, César A. Hidalgo

专题命中 提示注入 :prompt injection(title);分类 cs.CY

AI总结 研究揭示了共识生成LLMs在数字民主中的提示注入漏洞,并提出鲁棒性方法以减少方向性失败。

Comments 33 pages, 11 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20666 2026-03-03 cs.CL cs.AI 62%

Cognitive models can reveal interpretable value trade-offs in language models

认知模型可以揭示语言模型中的可解释价值权衡

Sonia K. Murthy, Rosie Zhao, Jennifer Hu, Sham Kakade, Markus Wulfmeier, Peng Qian, Tomer Ullman

机构 * Kempner Institute for Natural and Artificial Intelligence, Harvard University(哈佛大学自然与人工智能研究所) Google DeepMind(谷歌DeepMind) Department of Psychology, Harvard University(哈佛大学心理学系)

专题命中 提示注入 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出利用认知模型评估语言模型中的价值权衡,揭示其行为特征变化及对社会行为的影响。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00472 2026-03-03 cs.AI cs.SE 57%

From Goals to Aspects, Revisited: An NFR Pattern Language for Agentic AI Systems

从目标到方面,重新审视:面向智能体AI系统的NFR模式语言

Yijun Yu

机构 * The Open University(开放大学)

专题命中 提示注入 :prompt injection(abstract);分类 cs.AI

AI总结 本文提出面向智能体AI系统的NFR模式语言,通过目标驱动的方法系统发现和模块化交叉切面问题,提升系统可靠性。

Comments 12 pages, submitted

详情

展开后加载摘要…

URL PDF HTML 收藏