MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration
MMAligner:通过表示校准保护多模态大语言模型
专题命中 幻觉与事实性 :alignment(abstract);safety(abstract)
AI总结 MMAligner通过校准多模态大语言模型的表示,将不安全多模态输入的拒绝率提升至99%,仅造成不足2%的效用下降,显著优化了安全与效用的权衡。
Comments To Appear in the Proceedings of The ACM Conference on Computer and Communications Security (CCS), 2026