The Alignment Tax: Response Homogenization in Aligned LLMs and Its Implications for Uncertainty Estimation
对齐税:对齐语言模型中的响应同质化及其对不确定性估计的影响
专题命中 幻觉与事实性 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 研究发现对齐语言模型存在响应同质化现象,影响不确定性估计方法的效果,提出通过叠加不确定性信号提升模型性能。
Comments 25 pages, 3 figures, 10 tables, 24 experiments across 5 benchmarks. v2: added SINdex head-to-head (Exp 27), NLI validation (Exp 28), decoding protocol analysis. Code: https://github.com/DigitLion/ucbd-experiment