UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases
UnsafeChain: 通过困难案例增强推理模型安全性
机构 * Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学) ; Cluster Innovation Centre, University of Delhi(德里大学集群创新中心) ; INSAIT
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL
AI总结 UnsafeChain通过构建包含困难提示的对齐数据集,提升模型安全性并保持推理能力,实验显示其在多个基准测试中表现优异。
Journal ref The Asian Federation of Natural Language Processing and The Association for Computational Linguistics 2025