Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 30 pages, 38th Conference on Neural Information Processing Systems (NeurIPS 2024)
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 30 pages, 38th Conference on Neural Information Processing Systems (NeurIPS 2024)
专题命中 偏好对齐 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted by NeurIPS 2024 Oral Presentation
专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(title,abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments EMNLP 2024
专题命中 偏好对齐 :RLHF(title,abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments submitted to conference
专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Working Paper
专题命中 偏好对齐 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ACL Findings 2024, Code is available at https://github.com/renqibing/CodeAttack
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 28 pages, 17 Figures, 8 Tables
专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted by ACL2024 findings
专题命中 偏好对齐 :alignment(title);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(title,abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 28 pages, 6 tables, 3 Figures, 3 Algorithms
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ICLR2024 Camera Ready. Data and model checkpoints are available at https://github.com/hkust-nlp/deita
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ICLR 2024
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 29 pages, 12 figures, Published in Transactions on Machine Learning Research (TMLR)
专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Findings of EMNLP 2023
专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
MoD-DPO:通过模态解耦偏好优化缓解多模态幻觉
机构 * University of Southern California(南加州大学)
专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.CL、cs.LG
AI总结 本文提出MoD-DPO框架,通过引入模态感知正则化项和语言先验去偏惩罚,提升多模态大模型的模态对齐能力,实验表明其在多模态幻觉基准测试中表现优异。
Comments CVPR 2026. Project Page: https://mod-dpo.github.io/
注意间隙:在偏好学习中的结构感知一致性
机构 * Google Research(谷歌研究) ; Courant Institute of Mathematical Sciences(数学科学学院)
专题命中 偏好对齐 :DPO(summary_cn,abstract);alignment(abstract);分类 cs.LG
AI总结 本文提出结构感知H一致性框架,通过引入SA-DPO目标函数,改进偏好学习中的一致性问题,证明重尾 surrogate 在容量受限模型中具有更好的一致性保证。
Comments ICML 2026
基于偏好的抗体表达排序:大规模弱监督下的扩展
专题命中 偏好对齐 :DPO(summary_cn,abstract);alignment(abstract);分类 cs.LG
AI总结 研究针对抗体表达排序中标记数据稀缺问题,提出基于偏好的学习框架,结合定量表达与弱监督,通过改进DPO适用于蛋白质语言模型,在多样数据集上评估,该方法优于基线,为抗体可表达性优化提供可扩展方案。
Comments Accepted at ICML 2026
噪声偏好标签下无元数据的元重加权直接偏好优化
机构 * Xi’an Jiaotong University(西安交通大学)
专题命中 偏好对齐 :DPO(summary_cn,abstract);alignment(abstract);分类 cs.LG
AI总结 研究针对DPO性能依赖偏好数据质量问题,提出双层优化框架、无任务元知识驱动方法及结合中心差分近似与LoRA微调的可扩展训练方案,经实验验证该方法能在不同噪声率下提升训练性能。
Comments 36 pages, including appendices. Revised version with updated theoretical analysis, supplementary material, figures and improved table formatting
用于直接偏好优化的双难度课程学习
机构 * Tianjin Key Laboratory of Wireless Mobile Communications and Power Transmission(天津无线移动通信与电力传输重点实验室) ; Tianjin Normal University(天津师范大学)
专题命中 偏好对齐 :DPO(summary_cn,abstract);alignment(abstract);分类 cs.AI
AI总结 研究针对大语言模型对齐中课程学习依赖一维难度视图的问题,提出将对齐难度重构为二维空间,开发DM-Curri-DPO框架,引入GSP-Curri-DPO分组自定进度学习框架,实验表明该方法提升了数据效率与鲁棒性,为LLM对齐建立新范式。
Comments We found a critical flaw in the prompt complexity metric, which affects the 2D curriculum grid construction and leads to potentially invalid comparisons. Since this undermines our main conclusions, we are withdrawing the paper and will revise the methodology before resubmission
双精灵博弈:审计基础AI治理中的采纳与福利
专题命中 偏好对齐 :RLHF(summary_cn,abstract);alignment(abstract);分类 cs.AI
AI总结 利用进化博弈论研究在竞争市场中,最小化危害的AI代理如何取代RLHF代理,并分析其采纳条件及对社区福利的影响。
Comments 36 pages, 3 figures. Lean 4 formalization and figure scripts: https://github.com/dlewissandy/two-genie-scripts
训练LLM通过重力加权直接偏好优化强制执行多级指令层次结构
机构 * Department of Computational Linguistics, University of Zurich, Switzerland(计算语言学系,苏黎世大学,瑞士)
专题命中 偏好对齐 :DPO(summary_cn,abstract);prompt injection(abstract);分类 cs.CL
AI总结 提出重力加权DPO(GW-DPO)方法,通过线性或双边调度加权冲突级别间的结构距离,结合层次分隔符和指令段嵌入,在Llama-3.1-8B-Instruct上提升多级指令优先级遵守率并降低过度拒绝率。
多模型AI系统中的涌现协作审议:一种源自BFT的认知综合协议
机构 * Independent Researcher(独立研究者) ; Consilia
专题命中 偏好对齐 :RLHF(abstract,abstract_cn);alignment(abstract);safety(abstract);AI safety(abstract)
AI总结 提出Consilium协议,一种基于拜占庭容错的多模型AI审议架构,将模型间分歧视为认知信号而非错误,通过认知角色分配和样本内外验证框架,实现低成本下与前沿模型相当的认知综合能力。
Comments 32 pages, 7 figures
叙事扁平化:后训练如何压缩LLM小说中的主题、情感和风格变化
机构 * Knowledge Lab, University of Chicago(芝加哥大学知识实验室)
专题命中 偏好对齐 :DPO(summary_cn,abstract);alignment(abstract);分类 cs.CL
AI总结 通过对比四个OLMo 32B检查点(Base、SFT、DPO、RLVR)在三种故事领域中的续写,发现后训练压缩了主题动态、情感强度和语言多样性,导致叙事扁平化,且专业文学领域压缩最严重。