Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language Models
思维链作为镜像:评估大语言模型与人类偏好之间结构化推理的对齐
机构 * School of Computer Science and Informatics, University of Liverpool, United Kingdom(利物浦大学计算机科学与信息学学院)
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title);分类 cs.AI
AI总结 本文提出一种量化评估大语言模型多步结构化推理与人类偏好对齐的方法,引入对齐分数指标,通过构建基于语义熵的矩阵比较模型生成的思维链与人类偏好参考,发现对齐分数在2步推理时达到峰值,支持其作为诊断信号。
Comments Accepted to ACL 2026 (Main Conference)