Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
一致性并非对齐:人类与大语言模型(LLM)伦理判断中的道德依据分歧
机构 * University of Ljubljana(卢布尔雅那大学) ; Faculty of Computer and Information Science(计算机与信息科学学院) ; Faculty of Theology(神学院) ; Faculty of Arts(艺术学院)
专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI
AI总结 该研究以ETHICS基准的500项道德判断条目为对象,发现LLM与人类的最终伦理判断常一致,但道德依据存在系统性分歧,表明一致性不等同于对齐性,仅靠标签评估易产生误导。
Comments 9 pages, 4 figures, 3 tables. Accepted and presented at the AI Transparency Conference (AITC 2026), Nuremberg, Germany, June 5-6, 2026