Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models
自适应且显式安全:触发大型推理模型中的潜在安全意识
机构 * The State Key Laboratory of Blockchain and Data Security, Zhejiang University(浙江大学区块链与数据安全全国重点实验室) ; Hangzhou HighTech Zone (Binjiang) Blockchain and Data Security Research Institute, China(杭州高新区(滨江)区块链与数据安全研究院) ; Li Auto Inc.(理想汽车) ; Tsinghua University(清华大学) ; King Abdullah University Of Science And Technology(阿卜杜拉国王科技大学)
AI总结 针对大型推理模型易受越狱攻击的问题,提出Safe Trigger方法,通过SFT显式诱导安全标签触发安全分析,并用DPO优化,显著降低攻击成功率而不影响通用性能。