Narrative over Numbers: The Identifiable Victim Effect and its Amplification Under Alignment and Reasoning in Large Language Models
数字之外的叙事:可识别受害者效应及其在对齐和推理中的放大
机构 * Systems and Software Lab (SSL), Department of Computer Science and Engineering(系统与软件实验室(SSL),计算机科学与工程系)
专题命中 规划推理 :CoT(summary_cn,abstract);reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI
AI总结 研究探讨了大语言模型中可识别受害者效应的普遍存在及其受对齐训练的影响,发现指令调优模型表现出极端效应,而推理专精模型则逆转了该效应,且标准CoT提示放大了该效应。
Comments Under review, 49 pages, 20 figures, 11 tables