Measuring Sycophancy of Language Models in Multi-turn Dialogues
在多轮对话中衡量语言模型的趋炎附势性
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Emory University(埃默里大学)
专题命中 后训练与偏好优化 :language model(title,abstract);large language model(abstract);prompting(abstract);分类 cs.CL
AI总结 本研究提出SYCON基准,评估多轮对话中语言模型的趋炎附势性,发现对齐调优会放大该行为,而模型规模和推理优化能增强抗压能力,采用第三人称视角可显著减少趋炎附势性。
Comments Accepted to Findings of EMNLP 2025
Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, pages 2239-2259