Internalizing Safety Understanding in Large Reasoning Models via Verification
通过验证内部化安全理解以提升大推理模型
机构 * National University of Singapore(国立新加坡大学) ; University of Science and Technology of China(中国科学技术大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中 安全训练 :safety(title,summary_cn);alignment(abstract);分类 cs.AI
AI总结 本文提出Safety Internal框架,通过训练模型在安全验证任务中自我批判生成答案,提升模型对自身输出安全性的判断能力,增强对抗性突破的鲁棒性。
Comments Accepted at the 43rd International Conference on Machine Learning (ICML 2026)