Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence
通过潜在激活一致性测量稀疏自编码器中的单语义性
专题命中 VLM训练与架构 :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 该研究针对可解释人工智能中评估稀疏自编码器单语义性的挑战,提出无标签的Tversky单语义性分数(TMS),通过激活集一致性衡量。在多种模型和设置下评估,结果显示TMS受编码器各向异性影响小,能揭示训练动态,还能体现概念删除有效性。
Comments This is a preprint version. A shorter version of this paper has been accepted for presentation and publication in the post-workshop proceedings of the 8th International Workshop on eXplainable Knowledge Discovery in Data Mining (XKDD 2026), co-located with ECML PKDD 2026. The appendix is included only in this preprint and is not part of the peer-reviewed proceedings paper