LLMs Capture Emotion Labels, Not Emotion Uncertainty: Distributional Analysis and Calibration of Human-LLM Judgment Gaps
大型语言模型捕捉情绪标签而非情绪不确定性:人类与LLM判断差距的分布分析与校准
机构 * Faculty of Business and Commerce(商科学院) ; Faculty of Business Data Science(商务数据科学学院) ; RIKEN Center for Advanced Intelligence Project(先进智能项目研究所) ; Faculty of Data Science(数据科学学院) ; Japan Safety Society Research Center(日本安全协会研究中心)
专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL
AI总结 本文通过对比人类标注与LLM在GoEmotions和EmoBank基准上的情绪判断分布,发现零样本模型与人类分布差异显著,需领域微调而非模型规模扩大以缩小差距,并提出三种轻量校准方法以提升情绪标注的准确性。