DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values
DiverValue-Bench:用于使大语言模型与多元人类价值观对齐的基准及微调框架
专题命中 偏好对齐 :DPO(summary_cn,abstract);alignment(abstract);分类 cs.CL、cs.AI
AI总结 该研究推出覆盖74个国家/地区的人群感知基准DiverValue-Bench,发现现有LLMs存在地理与人口统计学价值观对齐差异,结合LoRA与DPO的轻量微调可提升对齐效果,为全球公平AI开发提供实用基础。
Comments 11 pages, 5 figures. Accepted to IJCAI-ECAI 2026 (Human-Centred AI Special Track). v2: Updated to the camera-ready version