arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3227 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3227 篇

2604.17299 2026-08-14 cs.CL cs.AI 版本更新 94%

Cat-DPO: Category-Adaptive Safety Alignment

Cat-DPO:基于类别的安全对齐

Tiankai Yang, Yi Nian, Xinyuan Li, Ruiyao Xu, Henry Peng Zou, Kaize Ding, Xiyang Hu, Yan Liu, Yue Zhao

机构 * University of Southern California(南加州大学) Northwestern University(西北大学)

专题命中 偏好对齐 :DPO(title,title_cn);alignment(title,abstract);safety(title,abstract);harmlessness(abstract)

AI总结 Cat-DPO通过将安全对齐转化为类别约束优化问题,为每个有害类别设置独立的适应性安全边际,提升整体帮助性和无害性,减少类别间的安全方差和最佳至最差差距。

Comments 23 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07678 2026-06-16 cs.LG cs.AI 新提交 94%

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment

DOG-DPO:几何中的动态优化用于安全对齐

Yi Nian, Tiankai Yang, Yudi Zhang, Qi Pan, Zelong Xu, Shenzhe Zhu, Qingqing Luan, Yue Huang, Xiangliang Zhang, Yue Zhao

机构 * University of Southern California(南加州大学) Iowa State University(爱荷华州立大学) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) UT Austin(德克萨斯大学奥斯汀分校) Independent Researcher(独立研究员) University of Notre Dame(圣母大学)

专题命中 偏好对齐 :DPO(title,title_cn);alignment(title,abstract);safety(title,abstract);分类 cs.AI、cs.LG

AI总结 提出DOG-DPO框架,将偏好对表示为模型表示空间中的方向,通过几何分解和多样性覆盖选择子集,仅用11%数据即可恢复大部分安全增益。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01523 2026-05-19 cs.LG stat.ML 93%

Beyond RLHF: A Unified Theoretical Framework of Alignment

超越RLHF:对齐的统一理论框架

Jihun Yun, Juno Kim, Jongho Park, Junhyuck Kim, Jongha Jon Ryu, Jaewoong Cho, Kwang-Sung Jun

机构 * KRAFTON UC Berkeley(加州大学伯克利分校) MIT(麻省理工学院) POSTECH

专题命中 偏好对齐 :RLHF(title,title_cn);alignment(title,abstract);DPO(abstract,abstract_cn);分类 cs.LG

AI总结 本文提出了一种统一的对齐理论框架,通过将对齐视为基于成对偏好的分布学习,推导出三种新的对齐目标,并证明了它们在非渐近情况下具有O(1/n)的收敛性,为RLHF提供了理论支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02018 2026-06-03 cs.CL 93%

Enhancing Paraphrase Type Generation: The Impact of DPO and RLHF Evaluated with Human-Ranked Data

增强释义类型生成:基于人工排序数据的DPO和RLHF评估影响

Christopher Lee Lübbers

机构 * University of Göttingen(哥廷根大学)

专题命中 偏好对齐 :DPO(title,title_cn);RLHF(title,title_cn);分类 cs.CL

AI总结 本研究利用人工排序的释义类型数据集,结合直接偏好优化(DPO)使模型输出与人类判断对齐,将释义类型生成准确率提升3个百分点,人类偏好评分提升7个百分点,并创建了新的标注数据集以支持更严格的评估。

Comments 21 pages, 11 figures. Master's thesis, University of Goettingen, December 2024. Code: https://github.com/cluebbers/dpo-rlhf-paraphrase-types. Models: https://huggingface.co/collections/cluebbers/enhancing-paraphrase-type-generation-673ca8d75dfe2ce962a48ac0

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26315 2026-05-27 cs.LG cs.AI 93%

Curriculum Learning for Safety Alignment

用于安全对齐的课程学习

Sandeep Kumar, Virginia Smith, Chhavi Yadav

机构 * Carnegie Mellon University(卡内基梅隆大学) Simons Institute, UC Berkeley(Simons研究所,伯克利大学)

专题命中 偏好对齐 :safety(title,abstract);DPO(summary_cn,abstract);alignment(title,abstract);jailbreak(abstract)

AI总结 提出基于课程学习的Staged-Competence框架,通过难度分级的偏好数据和渐进式参考模型更新,提升DPO安全对齐的鲁棒性,在三个模型族上平均降低16%的OOD有害响应率和20%的越狱攻击成功率。

Comments Accepted at the ICML 2026 GlobalSouthML Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20685 2026-04-23 cs.LG 93%

MGDA-Decoupled: Geometry-Aware Multi-Objective Optimisation for DPO-based LLM Alignment

MGDA-Decoupled:基于几何的多目标优化用于基于DPO的LLM对齐

Andor Vári-Kakas, Ji Won Park, Natasa Tagasovska

机构 * Prescient Design, CS CoE, Genentech | Roche(预见设计,计算机科学学院,基因泰克 | 罗氏)

专题命中 偏好对齐 :DPO(title,title_cn);alignment(title,abstract);harmlessness(abstract);分类 cs.LG

AI总结 本文提出MGDA-Decoupled算法,通过几何方法在DPO框架内实现更公平的多目标优化,实验显示其在UltraFeedback数据集上表现最优。

Comments Accepted to the Algorithmic Fairness Across Alignment Procedures and Agentic Systems Workshop at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30021 2026-06-04 cs.CL 93%

Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs

在不损失对齐的情况下恢复多样性:面向后训练大语言模型的DPO配方

Vinay Samuel, Yapei Chang, Mohit Iyyer

机构 * University of Maryland, College Park(马里兰大学 College Park 分校)

专题命中 偏好对齐 :DPO(title,title_cn);alignment(title,abstract);safety(abstract);分类 cs.CL

AI总结 提出REDIPO数据构建流程,通过离线DPO从基础模型生成中恢复多样性答案,同时保持指令模型的对齐性能。

Comments Under Review. 26 pages, 3 figures, 16 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03238 2026-07-10 cs.LG cs.AI 版本更新 93%

When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming

当RLHF失败时:奖励黑客、崩溃和评估者博弈的机制分类

Zelalem Abahana, David Evans, Satish Mahadevan Srinivasan, Matjaz Gams

机构 * First Citizens Bank(第一公民银行) Alma Mater Europaea University(欧洲大学)

专题命中 偏好对齐 :RLHF(title,title_cn);DPO(summary_cn,abstract);分类 cs.AI、cs.LG

AI总结 本文通过PPO、DPO等方法的对比实验,提出了一种基于奖励和评估者分数方向的机制分类法,将RLHF失败模式分类为可定位、可预测的训练动态。

Comments 20 pages, 8 figures; includes code, artifacts, and live demo

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23868 2026-05-15 cs.LG cs.CL 93%

GIFT: Group-Relative Implicit Fine-Tuning Integrates GRPO with DPO and UNA

GIFT: 组相对隐式微调整合GRPO与DPO和UNA

Zhichao Wang

机构 * Inflection AI

专题命中 偏好对齐 :DPO(title,title_cn);RLHF(summary_cn,abstract);分类 cs.CL、cs.LG

AI总结 GIFT结合GRPO组采样、DPO隐式奖励和UNA的隐式与显式优势MSE,通过z-score标准化消除DPO隐式奖励中的不可行分区函数Z(x)和RLHF/RLVR目标中的KL系数β,以组相对隐式微调解决相同参数策略族,用提示适应的β(x)替代外部调优的β。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21346 2026-02-26 cs.CL cs.AI 93%

Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment

基于对齐的DPO:一种原则性的推理方法以提高安全性对齐

Mengxuan Hu, Vivek V. Datla, Anoop Kumar, Zihan Guan, Sheng Li, Alfy Samuel, Daben Liu

机构 * University of Virginia(弗吉尼亚大学) Capital One

专题命中 偏好对齐 :alignment(title,abstract);DPO(title,abstract);safety(title,abstract);RLHF(abstract)

AI总结 本文提出基于对齐的DPO方法,通过引入推理意识的后训练和对齐加权机制,提升大语言模型的安全性对齐鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03520 2026-06-11 cs.LG cs.AI cs.SY eess.SY 版本更新 93%

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment

可认证安全RLHF:基于语义基础与固定惩罚约束优化的更安全大语言模型对齐

Kartik Pandit, Sourav Ganguly, Arnesh Banerjee, Shaahin Angizi, Arnob Ghosh

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) New Jersey Institute of Technology(新泽西理工学院) Department of Computer Engineering(计算机工程系) Heritage Institute of Technology(遗产理工学院)

专题命中 偏好对齐 :RLHF(title,title_cn);alignment(title);safety(abstract);分类 cs.AI、cs.LG

AI总结 针对现有RLHF方法依赖奖励/成本函数和双变量调优导致性能敏感且缺乏可证明安全保证的问题,提出CS-RLHF,通过语义基础成本模型和固定惩罚约束优化,实现可认证安全对齐,效率提升至少5倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21225 2026-05-21 cs.LG cs.AI 93%

PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment

PREFINE: 基于偏好的隐式奖励和成本微调以实现安全对齐

Richa Verma, Bavish Kulur, Sanjay Chawla, Balaraman Ravindran

机构 * TCS Research, \ of CSE, IIT Madras India Department of Computing Science, \ of Alberta Canada Qatar Computing Research Institute, \ Bin Khalifa University Qatar Department of Data Science \& AI, Wadhwani School of Data Science \& AI, IIT Madras India TCS Research, \ of CSE, IIT Madras Department of Computing Science, \ of Alberta Qatar Computing Research Institute, \ Bin Khalifa University Department of Data Science \& AI, Wadhwani School of Data Science \& AI, IIT Madras

专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);RLHF(abstract,abstract_cn);DPO(abstract,abstract_cn)

AI总结 该研究提出PREFINE方法,通过基于偏好的隐式奖励和成本微调,在连续控制环境中实现安全策略对齐,通过微调预训练强化学习策略以生成低成本行为同时保持高奖励。

Comments Accepted at AAMAS 2026 as a full paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04149 2026-05-19 cs.CL cs.AI cs.LG 93%

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap

基于难度的偏好数据选择:通过DPO隐式奖励差距

Xuan Qi, Rongwu Xu, Zhijing Jin

机构 * Paul G. Allen School of Computer Science & Engineering, University of Washington(华盛顿大学计算机科学与工程保罗·G·艾伦学校) Max Planck Institute for Intelligent Systems, Tübingen, Germany(德国图宾根马克斯·普朗克智能系统研究所) Jinesis Lab, University of Toronto & Vector Institute(多伦多大学Jinesis实验室及向量研究所)

专题命中 偏好对齐 :DPO(title,title_cn);RLHF(abstract,abstract_cn);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出基于难度的偏好数据选择方法,利用DPO隐式奖励机制选择奖励差距小的样本,提升数据效率和模型对齐性能,在多个数据集和对齐任务中优于五个基线方法。

Comments Our code and data are available at https://github.com/Difficulty-Based-Preference-Data-Select/Difficulty-Based-Preference-Data-Select

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09735 2026-06-09 cs.CL 新提交 92%

The Neutral Mask: How RLHF Provides Shallow Alignment while Leaving Partisan Structure Intact in a Large Language Model

中性面具:RLHF如何提供浅层对齐而保留大语言模型中的党派结构

Wendy K. Tam

机构 * Vanderbilt University(范德堡大学) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) National Center for Supercomputing Applications(国家超级计算应用中心)

专题命中 偏好对齐 :RLHF(title,title_cn);alignment(title,abstract);分类 cs.CL

AI总结 研究RLHF对Llama 3.1 8B党派倾向的影响,发现RLHF仅压缩党派信号方差以实现中性输出,而非移除党派结构,且特征级操控可绕过对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12843 2026-06-25 cs.LG cs.AI 版本更新 92%

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF

偏差拟合以缓解RLHF中奖励模型的长度偏差

Kangwen Zhao, Jianfeng Cai, Jinhua Zhu, Ruopei Sun, Dongyun Xue, Wengang Zhou, Li Li, Houqiang Li

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 偏好对齐 :RLHF(title,title_cn);DPO(abstract,abstract_cn);alignment(abstract);分类 cs.AI、cs.LG

AI总结 提出FiMi-RM框架,通过自学习长度与奖励的非线性关系并解耦,有效缓解RLHF中奖励模型因长度偏差导致的奖励破解问题。

Comments 16 pages, 12 figures. Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10892 2026-06-08 cs.LG 版本更新 92%

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models

多目标偏好优化:提升生成模型的人类对齐

Akhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng Wen

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学)

专题命中 偏好对齐 :RLHF(summary_cn,abstract);alignment(title,abstract);DPO(abstract,abstract_cn);safety(abstract)

AI总结 针对RLHF和偏好优化方法假设单一目标的问题,提出多目标偏好优化框架MOPO,通过约束KL散度最大化主要目标并保障次要目标下限,在合成基准和人类偏好数据上实现帕累托最优策略。

Comments arXiv admin note: text overlap with arXiv:2406.18853 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02193 2025-07-29 cs.AI 92%

More is Less: The Pitfalls of Multi-Model Synthetic Preference Data in DPO Safety Alignment

Yifan Wang, Runjin Chen, Bolian Li, David Cho, Yihe Deng, Ruqi Zhang, Tianlong Chen, Zhangyang Wang, Ananth Grama, Junyuan Hong

机构 * Purdue University(普渡大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校) The University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 偏好对齐 :alignment(title,abstract);DPO(title,abstract);safety(title,abstract);RLHF(abstract)

Comments This version includes updated results and expanded discussion

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10054 2026-05-26 cs.LG cs.AI cs.CL cs.CV 92%

Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs

Uni-DPO:大语言模型动态偏好优化的统一范式

Shangpin Peng, Weinong Wang, Zhuotao Tian, Senqiao Yang, Xing Wu, Haotian Xu, Chengquan Zhang, Takashi Isobe, Baotian Hu, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Xi’an Jiaotong University(西安交通大学) The Chinese University of Hong Kong(香港中文大学) University of Chinese Academy of Sciences(中国科学院大学) Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 偏好对齐 :DPO(title,title_cn);RLHF(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 针对现有DPO方法忽略数据质量和学习难度差异的问题,提出Uni-DPO统一框架,通过自适应重加权偏好对实现更有效的数据利用和更优性能。

Comments Accepted by ICLR 2026. Code & models: https://github.com/pspdada/Uni-DPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09055 2025-09-12 cs.CL cs.AI cs.LG 92%

Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M

Piyush Pant

机构 * Saarland University(萨尔兰大学)

专题命中 偏好对齐 :DPO(title,abstract);safety(title,abstract);alignment(abstract);RLHF(abstract)

Comments 17 pages, 3 figures. Code and dataset available at https://github.com/PiyushWithPant/Improving-LLM-Safety-and-Helpfulness-using-SFT-and-DPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11455 2025-02-18 cs.CR 92%

Adversary-Aware DPO: Enhancing Safety Alignment in Vision Language Models via Adversarial Training

Fenghua Weng, Jian Lou, Jun Feng, Minlie Huang, Wenjie Wang

专题命中 偏好对齐 :alignment(title,abstract);DPO(title,abstract);safety(title,abstract);jailbreak(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10998 2026-05-13 cs.CR cs.AI 92%

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs

少样本真正无害的DPO攻击用于对抗LLMs

Sangyeon Yoon, Wonje Jeung, Yoonjun Cho, Dongjae Jeon, Albert No

机构 * Yonsei University(延世大学)

专题命中 偏好对齐 :DPO(title,title_cn);alignment(abstract);safety(abstract);分类 cs.AI

AI总结 研究提出一种利用10对无害偏好对的真正无害DPO攻击,通过优化模型偏好来减少拒绝行为,从而在对抗LLMs时取得高成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00224 2026-05-04 cs.AI 92%

TUR-DPO: Topology- and Uncertainty-Aware Direct Preference Optimization

TUR-DPO:基于拓扑和不确定性的直接偏好优化

Abdulhady Abas Abdullah, Fatemeh Daneshfar, Seyedali Mirjalili, Mourad Oussalah

机构 * Artificial Intelligence and Innovation Centre, University of Kurdistan, Erbil, Iraq(人工智能与创新中心,乌尔米耶大学,伊拉克) Department of Computer Engineering, University of Kurdistan, Iran(计算机工程系,乌尔米耶大学,伊朗) Centre for Artificial Intelligence Research and Optimisation, Torrens University Australia, Brisbane, Australia(人工智能研究与优化中心,塔伦斯大学澳大利亚,布里斯班,澳大利亚) Research and Innovation Center, Obuda University, Budapest 1034, Hungary(研究与创新中心,奥布达大学,布达佩斯1034,匈牙利) Center for Machine Vision and Signal Analysis (CMVS), University of Oulu, Finland(机器视觉与信号分析中心(CMVS),奥卢大学,芬兰)

专题命中 偏好对齐 :DPO(title,title_cn);RLHF(abstract,abstract_cn);分类 cs.AI

AI总结 TUR-DPO通过引入轻量级推理拓扑和结合语义忠实度、效用和拓扑质量,提升偏好对齐的稳定性与鲁棒性,同时保持训练简洁性和无需在线回滚。

Comments Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19728 2026-04-15 cs.LG 92%

Hard Negative Sample-Augmented DPO Post-Training for Small Language Models

增强的DPO后训练用于小型语言模型

Haocheng Lu, Minjun Zhu, Henry Yu

机构 * Computer Science NYU Shanghai(纽约大学上海学院)

专题命中 偏好对齐 :DPO(title,title_cn);RLHF(abstract,abstract_cn);分类 cs.LG

AI总结 本文提出一种轻量级后训练方法,通过MathVerifier检测结构化错误,改进DPO以提升小型模型在数学推理中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00849 2024-03-11 cs.CL cs.CV 92%

RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback

Tianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu, Hai-Tao Zheng, Maosong Sun, Tat-Seng Chua

专题命中 偏好对齐 :alignment(title,abstract);RLHF(title,abstract);trustworthy(title,abstract);分类 cs.CL

Comments Accepted by CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10126 2026-08-12 cs.LG cs.AI cs.CL 新提交 91%

Procedural Fairness Failures in RLHF from Preference Averaging

来自偏好平均的RLHF中的程序公平性失败

M P V S Gopinadh, Karthik Kamuju, Kummari Avinash, John Joshua, Srinivasa Raju Rudraraju

机构 * Vishnu Institute of Technology(维什努理工学院)

专题命中 偏好对齐 :RLHF(title,title_cn);alignment(abstract,comments);分类 cs.CL、cs.AI、cs.LG

AI总结 该研究指出标准RLHF因偏好平均引发程序公平性失败,提出PA-RLHF分开优化不同偏好模式,提升了对齐准确率并缩小了群体公平差距,对大模型和智能体系统有重要意义。

Comments 4 pages, Accepted at the ICLR 2026 Workshop on Algorithmic Fairness Across Alignment Procedures and Agentic Systems (AFAA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04389 2026-06-23 cs.CL cs.AI 版本更新 91%

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment

安全并非普遍:大语言模型对齐中的选择性安全陷阱

Iago Alves Brito, Walcy Santos Rezende Rios, Julia Soares Dollis, Diogo Fernandes Costa Silva, Arlindo Rodrigues Galvão Filho

机构 * Advanced Knowledge Center for Immersive Technologies(沉浸式技术先进知识中心) Federal University of Goiás(戈亚斯联邦大学)

专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);DPO(summary_cn,abstract);分类 cs.CL、cs.AI

AI总结 研究揭示大语言模型在安全对齐中的选择性安全陷阱,通过MiJaBench基准测试发现安全防御存在人口统计学层级差异,通过DPO优化实现跨群体的安全泛化。

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21438 2026-05-11 cs.CL cs.LG 91%

UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function

UFT:通过通用隐式奖励函数统一SFT和RLHF/DPO/UNA的微调

Zhichao Wang, Bin Bi, Zixu Zhu, Xiangbo Mao, Jun Wang, Shiyu Wang, Cheng Wang, Dong Nie, Lingzi Hong

机构 * Salesforce RadixArk ChatAlpha AI University of North Texas(北卡罗来纳州立大学)

专题命中 偏好对齐 :RLHF(title,title_cn);DPO(title,title_cn);alignment(abstract);分类 cs.CL、cs.LG

AI总结 本文提出UFT框架,通过隐式奖励函数整合SFT与对齐过程,提升指令微调和事实性任务的性能,实验显示其优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25895 2026-04-29 cs.CY cs.AI cs.CL 91%

Three Models of RLHF Annotation: Extension, Evidence, and Authority

三种RLHF注释模型:扩展、证据与权威

Steve Coyne

机构 * University of Toronto(多伦多大学)

专题命中 偏好对齐 :RLHF(title,title_cn);alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文探讨RLHF注释的三种模型:扩展、证据与权威,分析其对注释流程设计的影响,并提出应根据不同维度选择适配的模型。

Comments 17 pages. Accepted to ACM FAccT '26, June 25-28, Montreal

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11206 2024-01-23 cs.CL 91%

InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance

Pengyu Wang, Dong Zhang, Linyang Li, Chenkun Tan, Xinghao Wang, Ke Ren, Botian Jiang, Xipeng Qiu

专题命中 偏好对齐 :alignment(title,abstract);harmlessness(title,abstract);RLHF(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17044 2026-05-26 cs.LG cs.AI cs.CV 91%

Do Understanding and Generation Fight? A Diagnostic Study of DPO for Unified Multimodal Models

理解与生成相冲突吗?统一多模态模型DPO的诊断研究

Abinav Rao, Sujan Rachuri

专题命中 偏好对齐 :DPO(title,title_cn);alignment(abstract);分类 cs.AI、cs.LG

AI总结 通过系统实验发现,在统一多模态模型上应用DPO时,生成质量难以对齐,主要原因是理解和生成梯度近乎正交且存在11-14倍的幅度不平衡,源于VQ token数量不对称。

Comments Experiments are inconclusive: The claim that architectures such as Chameleon or Emu would exhibit stronger gradient conflict is not supported by experiments or analysis, and all experiments are conducted on Janus-Pro without evaluation on other unified multimodal architectures

详情

展开后加载摘要…

URL PDF HTML 收藏