Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints
通过耦合权重和激活约束防止大语言模型中的安全漂移
机构 * Hunan Normal University(湖南师范大学) ; The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,中国科学院自动化研究所) ; National University of Defense Technology(国防科技大学)
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.AI
AI总结 本文提出耦合权重和激活约束方法,通过同时约束权重更新和关键特征正则化,有效提升大语言模型的安全性,实验显示其在多种任务中表现优于现有方法。
Comments 17 pages, 6 figures, 6 tables, The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)