Large language models show fragile cognitive reasoning about human emotions
大型语言模型在人类情感认知推理中的脆弱性
AI总结 本文探讨大型语言模型是否能通过认知维度而非标签进行情感推理,提出CoRE基准测试,发现模型在情感认知结构上有系统性关系,但与人类判断存在偏差且在不同上下文中不稳定。
Comments Under Review, a version was presented at WiML Workshop @ NeurIPS 2025