Atlas-Alignment: Making Interpretability Transferable Across Language Models
Atlas-Alignment:使语言模型间的可解释性可迁移
机构 * Department of Artificial Intelligence, Fraunhofer Heinrich Hertz Institute(人工智能系,弗劳恩霍夫 Heinrich Hertz 研究所) ; Department of Electrical Engineering and Computer Science, Technische Universität Berlin(电气工程与计算机科学系,柏林技术大学) ; Centre of eXplainable Artificial Intelligence, Technological University Dublin(可解释人工智能中心,都柏林技术大学) ; BIFOLD - Berlin Institute for the Foundations of Learning and Data(BIFOLD - 柏林学习与数据基础研究所)
专题命中 安全训练 :alignment(title,title_cn);分类 cs.CL、cs.AI、cs.LG
AI总结 Atlas-Alignment通过在新模型的潜在空间与预训练的Concept Atlas对齐,实现无需标注数据的可解释性迁移,降低可解释AI的成本。