Graph-Regularized Sparse Autoencoders for LLM Safety Steering
图正则化稀疏自编码器用于LLM安全引导
机构 * ELLIS Institute Tübingen(图宾根ELLIS研究所) ; Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所) ; Intesa Sanpaolo(Intesa Sanpaolo公司) ; University of Southern California(南加州大学)
AI总结 本文提出图正则化稀疏自编码器,通过在神经元共激活图上平滑解码器向量并应用方向库,提升安全引导效果,在多个基准测试中显著提高有害请求拒绝率。