Dissociating the Internal Representations of Sycophancy in LLMs
区分大语言模型中谄媚行为的内部表征
专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.LG
AI总结 研究大语言模型谄媚行为,将其表征分为事实和观点子类型,通过训练线性探针等方法评估不同模型对两子类型表征差异,为研究复杂模型行为表征结构提供了新框架。
Comments Accepted to Mechanistic Interpretability Workshop at ICML 2026