Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models
自引导防御:通过合成指南实现推理模型的自适应安全对齐
机构 * Beijing Key Lab of Traffic Data Analysis and Mining(北京交通大数据分析与挖掘重点实验室) ; Beijing Jiaotong University(北京交通大学) ; School of Information Technology and Management(信息科技与管理学院) ; University of International Business and Economics(国际经济贸易大学)
专题命中 测试时计算 :reasoning(title,abstract);分类 cs.CL、cs.AI
AI总结 本文提出SGASA框架,通过合成指南实现推理模型的自适应安全对齐,提升对抗性提示防御能力并减少对良性请求的拒绝。