Eliminating Biased Length Reliance of Direct Preference Optimization via Down-Sampled KL Divergence
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL
Comments EMNLP 2024 Main, Final Version
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL
Comments EMNLP 2024 Main, Final Version
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.AI
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.AI
Journal ref EMNLP 2024
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL
Comments 19 pages, 6 figures
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.LG
Comments COLM 2024
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL
Comments published at ICLR 2024
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.LG
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.LG
Comments 27 pages, 7 figures, 2 tables
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL
Comments Accepted in ICLR 2024
前沿模型的成长之痛:当排行榜不再区分以及接下来衡量什么
机构 * Zehen Labs(泽亨实验室)
专题命中 偏好对齐 :RLHF(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG;alignment(comments)
AI总结 本文通过分解SWE-bench和GPQA Diamond分数为种群耦合趋势和每版本残差(h场),诊断前沿模型能力之间的协作与权衡,并提供三步诊断法、每实验室测量优先级表及七个可证伪预测。
Comments 13 pages, 5 figures, 4 tables. Companion paper: "Lying Is Just a Phase: The Hidden Alignment Transition in Language Model Scaling." ( https://doi.org/10.48550/arXiv.2605.18838 ). Code: https://github.com/adilamin89/cape-scaling . Dashboard: https://zehenlabs.com/cape/
机构 * KAIST AI(韩国科学技术院人工智能研究所)
专题命中 偏好对齐 :DPO(abstract,comments);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Preprint, under review. 39 pages, 12 figures. Updates from v1: Added new theoretical results on DPO training dynamics and policy exploration, included experiments with Qwen3-4B, and refined the discussion of log-margin dynamics
机构 * The University of Tokyo(东京大学)
专题命中 偏好对齐 :RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG;alignment(comments)
Comments ICML 2025 Workshop on Models of Human Feedback for AI Alignment
AlpsBench: 一个面向真实对话记忆与偏好对齐的LLM个性化基准
机构 * University of Science and Technology of China(科学技术大学) ; National University of Singapore(新加坡国立大学)
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI
AI总结 AlpsBench通过真实人类与LLM对话数据构建,包含2500个长期交互序列及验证的记忆,评估个性化信息提取、更新、检索与利用等核心任务,揭示LLM在记忆管理中的局限性。
偏好学习用于AI对齐:一种因果视角
机构 * Department of Applied Mathematics and Theoretical Physics(应用数学与理论物理系)
专题命中 偏好对齐 :alignment(title);分类 cs.AI、cs.LG
AI总结 本文从因果视角出发,探讨了基于偏好数据的奖励建模在对齐大语言模型与人类价值观中的关键作用,提出因果工具箱以解决因果误识别、偏好异质性和用户特定因素的混淆问题,并提出未来研究的方向。
Journal ref Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025
链式放大:通过尺度自回归和偏好对齐实现极端超分辨率
机构 * KAIST AI(韩国科学技术院人工智能系)
专题命中 偏好对齐 :alignment(title);分类 cs.AI、cs.LG
AI总结 本文提出Chain-of-Zoom框架,通过自回归链和多尺度提示实现极端超分辨率,利用GRPO优化文本提示对齐人类偏好,实验显示标准4x扩散模型在CoZ下可实现256倍以上高质量放大。
Comments NeurIPS 2025 (Spotlight)
RLAIF-SPA: 结构化AI反馈用于语音合成中的语义-语调对齐
机构 * School of Computer Science and Engineering, Northeastern University, China(东北大学计算机科学与工程学院)
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI
AI总结 本文提出RLAIF-SPA框架,通过整合强化学习从AI反馈来优化语音合成中的情感表达和可懂度,实验显示其在多个数据集上均取得显著提升。
流畅对齐与不流畅评判:低资源语言的后训练方法
机构 * Language Technology Group, University of Oslo(奥斯陆大学语言技术组)
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI
AI总结 本文提出一种低资源语言的后训练方法,通过不流畅评判模型保持语言模型的流畅性,通过挪威语案例研究验证了基于策略的训练方法在无需额外数据时的有效性。
Journal ref The Fourteenth International Conference on Learning Representations (ICLR 2026)
机构 * Shanghai Jiao Tong University(上海交通大学) ; Alibaba Cloud Computing(阿里云计算)
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI
Comments emnlp 2025 main conference
机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) ; Independent Researcher(独立研究者)
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI
机构 * College of Computer Science and Technology, Dalian University of Technology(大连理工大学计算机科学与技术学院) ; Air Force Communications NCO Academy(空军通信NCO学院)
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI
Comments Accepted as a short paper at BlBM2025
机构 * Radha Gulhane(独立研究者) ; Sathish Reddy Indurthi(独立研究者)
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI
机构 * Imperial College London(伦敦帝国学院) ; Apta AI
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI
Comments 5 pages, 1 figure, 3 tables
机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) ; Shanghai AI Laboratory(上海人工智能实验室)
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI
Comments Accepted to ACL 2025, Oral & Panel Discussion
专题命中 偏好对齐 :RLHF(title);分类 cs.CL、cs.LG
机构 * First Author Affiliation(第一作者机构)
专题命中 偏好对齐 :DPO(title);分类 cs.CL、cs.LG
Comments v4: ACL'25 industry track camera ready; v3: minor modifications; v2: better writing & format for later submission; all release at https://github.com/Qihoo360/Light-R1
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI
专题命中 偏好对齐 :RLHF(title);分类 cs.CL、cs.LG
Comments Proceedings of the First Workshop on Theory of Mind in Communicating Agents at (TOM @ ICML 2023)
专题命中 偏好对齐 :RLHF(title);分类 cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.LG
专题命中 偏好对齐 :alignment(title,comments);分类 cs.CL
Comments Presented at AAAI 2025, special track on AI Alignment