Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
标签问题与45种语言模型中谄媚现象的代际反转
机构 * Cornell Tech(康奈尔科技)
AI总结 研究在语言模型决策问题中附加标签的效应,通过特定实验设置在45个模型上测量,发现效应范围广且随代际反转,定位了抗性来源,表明标签极性更关键,工具能从模型行为读出反谄媚训练情况。
Comments 18 pages, 4 figures. Data, code, and raw model replies: https://github.com/tap2k/modelun. Interactive explorer: https://tap2k.github.io/modelun/suggestibility/