arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-05-25 至 2026-05-25 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 4 篇

2605.23157 2026-05-25 cs.CL 85%

Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs

相同模型,不同弱点:语言与模态如何重塑前沿多模态大语言模型的越狱攻击面

Casey Ford, Madison Van Doren, Sicheng Jin, Emily Dix

机构 * Appen(Appliance)

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract);分类 cs.CL

AI总结 通过跨语言多模态红队测试,发现语言和模态通过不同机制影响越狱漏洞,导致安全排名跨语言不守恒,需重新设计评估框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20102 2026-05-25 cs.LG cs.AI 82%

BarrierSteer: LLM Safety via Learning Barrier Steering

BarrierSteer: 通过学习障碍引导实现大语言模型安全

Thanh Q. Tran, Arun Verma, Kiwan Wong, Bryan Kian Hsiang Low, Daniela Rus, Wei Xiao

机构 * Department of Computer Science, National University of Singapore(新加坡国立大学计算机科学系) Singapore-MIT Alliance for Research and Technology Centre(新加坡-麻省理工联合研究中心) CSAIL, Massachusetts Institute of Technology(麻省理工学院计算机科学与人工智能实验室) Worcester Polytechnic Institute(沃斯堡理工学院)

专题命中 越狱攻击 :safety(title,abstract);分类 cs.AI、cs.LG

AI总结 提出 BarrierSteer 框架,通过将学习到的非线性安全约束作为控制障碍函数嵌入潜在表示空间,在推理时引导模型生成安全内容,理论保证且不修改模型参数。

Comments This paper introduces SafeBarrier, a framework that enforces safety in large language models by steering their latent representations with control barrier functions during inference, reducing adversarial and unsafe outputs

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00979 2026-05-25 cs.CR cs.AI cs.CL 62%

GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents

GradingAttack: 揭示基于LLM的教育评分代理中的安全漏洞

Xueyi Li, Zhuoneng Zhou, Zitao Liu, Yongdong Wu

机构 * Guangdong Institute of Smart Education(广东智能教育研究院) Jinan University(济南大学)

专题命中 越狱攻击 :trustworthy(abstract);分类 cs.CL、cs.AI

AI总结 提出GradingAttack框架,通过token级和prompt级对抗攻击策略,系统评估基于LLM的教育评分代理的安全漏洞,发现其缺乏鲁棒防御。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22842 2026-05-25 cs.CR cs.AI cs.LG 62%

The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems

归因偏差:当记忆中毒在自主AI系统中看起来像模型失败时

Tanzim Ahad, Ismail Hossain, Md Jahangir Alam, Sai Puppala, Syed Bahauddin Alam, Sajedul Talukder

机构 * Department of Computer Science, University of Texas at El Paso(德克萨斯大学埃尔帕索分校计算机科学系) School of Computing, Southern Illinois University Carbondale(南方伊利诺伊大学卡本代尔分校计算机学院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文识别了多智能体AI系统中的“归因偏差”问题,即记忆层攻击导致的行为与模型失败无法区分,并形式化了“语义规范漂移”(SND)作为第三种不当行为路径,提出了反事实组合测试和记忆持久信息流控制等防御方法。

Comments This paper is presently under review at a top-tier security venue

详情

展开后加载摘要…

URL PDF HTML 收藏