arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 30405 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3229 篇

2604.07754 2026-04-10 cs.CR cs.CL 85%

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training

对齐的艺术:细调方法如何有效地对齐和重新对齐训练后的LLM

Rui Zhang, Hongwei Li, Yun Shen, Xinyue Shen, Wenbo Jiang, Guowen Xu, Yang Liu, Michael Backes, Yang Zhang

机构 * University of Electronic Science and Technology of China(电子科技大学) Flexera CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心) Nanyang Technological University(南洋理工大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);safety(abstract);分类 cs.CL

AI总结 研究探讨细调方法在对齐和重新对齐LLM中的效果,揭示攻击与防御间的机制不对称,发现ORPO在对齐方面最有效,DPO在重新对齐中表现优异但牺牲了模型实用性,同时发现模型特定的抗性及多轮对抗动态的残留效应。

Comments Accepted by ACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14400 2026-03-23 cs.AI 85%

HPS: Hard Preference Sampling for Human Preference Alignment

HPS: 人类偏好对齐的硬偏好采样

Xiandong Zou, Wanyu Lin, Yuchen Li, Pan Zhou

机构 * Singapore Management University(新加坡国立管理学院) The Hong Kong Polytechnic University(香港理工大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);safety(abstract);分类 cs.AI

AI总结 本文提出HPS框架,通过优先选择最受好评的响应并拒绝所有不受欢迎和有害的响应,提升大语言模型与人类偏好的对齐效率和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16417 2026-03-18 cs.AI 85%

Via Negativa for AI Alignment: Why Negative Constraints Are Structurally Superior to Positive Preferences

通过否定来实现AI对齐:为什么否定约束在结构上优于积极偏好

Quan Cheng

机构 * Tsinghua University(清华大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);harmlessness(abstract);分类 cs.AI

AI总结 本文探讨了通过否定约束而非积极偏好进行AI对齐的结构优势,提出基于波普尔的证伪逻辑,指出否定信号方法在避免趋炎附势和提升效果上的有效性。

Comments 9 pages, position paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11126 2026-03-13 cs.MA cs.CL 85%

Enhancing Value Alignment of LLMs with Multi-agent system and Combinatorial Fusion

通过多智能体系统和组合融合增强大语言模型的价值对齐

Yuanhong Wu, Djallel Bouneffouf, D. Frank Hsu

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);trustworthy(abstract);分类 cs.CL

AI总结 本文提出VAS-CFA框架,通过多智能体和组合融合提升大语言模型的价值对齐,实验证明其在标准度量上优于现有方法。

Comments 5 pages, 3 figures, accepted to 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03054 2026-03-10 cs.CL 85%

PrivMedChat: End-to-End Differentially Private RLHF for Medical Dialogue Systems

PrivMedChat: 医疗对话系统的端到端差分隐私RLHF

Sudip Bhujel

机构 * Graduate Student in the Department of Computer Science, University of Kentucky(计算机科学系研究生,肯塔基大学)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);safety(abstract);分类 cs.CL

AI总结 PrivMedChat通过端到端差分隐私RLHF方法,在医疗对话系统中实现隐私保护与对齐,采用DP-SGD和DP-aware策略优化,避免临床标注成本,提供正式隐私保证。

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24159 2026-03-02 cs.AI 85%

RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment

RE-PO:一种用于大语言模型对齐的鲁棒增强策略优化通用框架

Xiaoyang Cao, Zelai Xu, Mo Guang, Kaiwen Long, Michiel A. Bakker, Yu Wang, Chao Yu

机构 * IDSS, Massachusetts Institute of Technology(IDSS,麻省理工学院) EE, Tsinghua University(电子工程系,清华大学) Li Auto Inc.(力汽车公司) SIGS, Tsinghua University(系统工程系,清华大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI

AI总结 RE-PO是一种用于大语言模型对齐的鲁棒增强策略优化通用框架,通过期望最大化程序推断标签正确性并自适应加权数据点以减轻噪声,提升对齐效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16987 2026-02-20 cs.CY 85%

A testable framework for AI alignment: Simulation Theology as an engineered worldview for silicon-based agents

一种可测试的AI对齐框架:模拟神学作为硅基智能体的工程世界观

Josef A. Habdank

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);safety(abstract);分类 cs.CY

AI总结 本文提出模拟神学作为可测试的AI对齐框架,通过内化目标减少欺骗行为,促进AI与人类的可持续共存。

Comments 16 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01128 2026-02-03 cs.LG 85%

Tangent Space Fine-Tuning for Directional Preference Alignment in Large Language Models

切线空间微调用于大语言模型中的方向偏好对齐

Mete Erdogan

机构 * Stanford University(斯坦福大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);safety(abstract);分类 cs.LG

AI总结 TS-DPO通过切线空间微调实现多偏好维度的可控对齐,提升模型在帮助性与冗余性之间的平衡能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14100 2025-12-17 cs.LG cs.LO 85%

A First-Order Logic-Based Alternative to Reward Models in RLHF

基于一阶逻辑的强化学习人类反馈中奖励模型的替代方案

Chunjin Jian, Xinhua Zhu

机构 * School of Computer Science(计算机科学系) Engineering Guangxi Normal University Guilin, China(广西师范大学)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);DPO(abstract);分类 cs.LG

AI总结 本文提出基于一阶逻辑的奖励机制,替代传统奖励建模,通过逻辑一致性提升模型对齐性能,实验显示其在性能和鲁棒性上优于标准微调方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00709 2025-12-02 cs.AI 85%

When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF

当人类偏好翻转时:面向RLHF的实例依赖性鲁棒损失

Yifan Xu, Xichen Ye, Yifan Chen, Qiaosheng Zhang

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);DPO(abstract);分类 cs.AI

AI总结 本文提出了一种面向RLHF的实例依赖性鲁棒损失算法,通过建模偏好翻转机制和引入实例依赖的翻转概率,提升对齐算法的鲁棒性。

Comments Accepted by AAAI-26-AIA

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12179 2025-11-19 cs.AI cs.MA 85%

Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation

Yubo Li, Weiyi Song

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08594 2025-11-13 cs.CL 85%

Diverse Preference Learning for Capabilities and Alignment

Stewart Slocum, Asher Parker-Sartori, Dylan Hadfield-Menell

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL

Journal ref 13th International Conference on Learning Representations (ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20995 2025-11-12 cs.CL 85%

ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models

Xiaomin Li, Xupeng Chen, Jingxuan Fan, Eric Hanchen Jiang, Mingye Gao

专题命中 偏好对齐 :safety(title,abstract);alignment(abstract);RLHF(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01894 2025-11-11 cs.CV cs.AI cs.HC 85%

LIVS: A Pluralistic Alignment Dataset for Inclusive Public Spaces

Rashid Mushkani, Shravan Nayak, Hugo Berard, Allison Cohen, Shin Koseki, Hadrien Bertrand

机构 * Mila–Quebec AI Institute(魁北克AI研究所)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);safety(abstract);分类 cs.AI

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20413 2025-10-24 cs.LG 85%

Why DPO is a Misspecified Estimator and How to Fix It

Aditya Gopalan, Sayak Ray Chowdhury, Debangshu Banerjee

机构 * IISc Bangalore(班加罗尔IISc) IIT Kanpur(坎purIIT) HP AI Research(HP人工智能研究)

专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);RLHF(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04346 2025-10-24 cs.CL 85%

Permutative Preference Alignment from Listwise Ranking of Human Judgments

Yang Zhao, Yixin Wang, Mingzhang Yin

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Michigan, Ann Arbor(密歇根大学安娜堡分校) University of Florida(佛罗里达大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL

Comments Published at EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10013 2025-10-14 cs.CL 85%

Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety

Yuyi Huang, Runzhe Zhan, Lidia S. Chao, Ailin Tao, Derek F. Wong

机构 * The Second Affiliated Hospital, Guangdong Provincial Key Laboratory of Allergy and Clinical Immunology, Guangzhou Medical University(广东省过敏与临床免疫学重点实验室,广州医科大学第二附属医院) NLP 2 CT Lab, Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系NLP2CT实验室)

专题命中 偏好对齐 :safety(title,abstract);alignment(abstract);RLHF(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25771 2025-10-01 cs.CV cs.AI 85%

Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs

Jia Jun Cheng Xian, Muchen Li, Haotian Yang, Xin Tao, Pengfei Wan, Leonid Sigal, Renjie Liao

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute for AI(人工智能向量研究所) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席) Kling Team, Kuaishou Technology(快手科技 Kling 团队) NSERC CRC Chair(加拿大NSERC CRC主席)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18391 2025-08-27 cs.AI 85%

PKG-DPO: Optimizing Domain-Specific AI systems with Physics Knowledge Graphs and Direct Preference Optimization

Nitin Nagesh Kulkarni, Bryson Wilcox, Max Sawa, Jason Thom

机构 * Advanced Engineering and Technology(先进工程与技术)

专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08998 2025-06-17 math.ST cs.LG stat.ML stat.TH 85%

On Monotonicity in AI Alignment

Gilles Bareilles, Julien Fageot, Lê-Nguyên Hoang, Peva Blanchard, Wassim Bouaziz, Sébastien Rouault, El-Mahdi El-Mhamdi

机构 * CTU in Prague, Tournesol(布拉格CTU,图尔诺索尔) Tournesol(图尔诺索尔) Calicarpa, Tournesol(卡利卡尔帕,图尔诺索尔) Kleis Technology(克莱斯技术) École Polytechnique(巴黎高等理工学院) Calicarpa(卡利卡尔帕)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08681 2025-06-12 cs.LG 85%

Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling

Phuc Minh Nguyen, Ngoc-Hieu Nguyen, Duy H. M. Nguyen, Anji Liu, An Mai, Binh T. Nguyen, Daniel Sonntag, Khoa D. Doan

机构 * VinUniversity(文大学) VinUni-Illinois Smart Health Center(VinUni-伊利诺伊智能健康中心) University of Stuttgart(斯图加特大学) University of California Los Angeles (UCLA)(加州大学洛杉矶分校) International University - VNUHCM(国际大学 - VNUHCM) University of Science - VNUHCM(科学大学 - VNUHCM) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) Oldenburg University(奥尔登堡大学) Max Planck Research School for Intelligent Systems (IMPRS-IS)(马克斯·普朗克智能系统研究学校)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.LG

Comments First version

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11283 2025-06-06 cs.LG 85%

AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment

Pankayaraj Pathmanathan, Udari Madhushani Sehwag, Michael-Andrei Panaitescu-Liess, Furong Huang

机构 * University of Maryland(马里兰大学) JPMorgan AI Research(摩根大通人工智能研究) University of Maryland Capital One(马里兰大学Capital One)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);jailbreak(abstract);分类 cs.LG

Comments Published at the Neurips Safe Generative AI Workshop 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03884 2025-06-02 cs.CL 85%

AlphaPO: Reward Shape Matters for LLM Alignment

Aman Gupta, Shao Tang, Qingquan Song, Sirou Zhu, Jiwoo Hong, Ankan Saha, Viral Gupta, Noah Lee, Eunki Kim, Siyu Zhu, Parag Agrawal, Natesh Pillai, S. Sathiya Keerthi

机构 * LinkedIn Corporation(LinkedIn公司) KAIST AI(KAIST人工智能) KAIST(韩国科学技术院)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL

Comments 26 pages, 16 figures. Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23749 2025-05-30 cs.LG cs.GT 85%

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

Paul Gölz, Nika Haghtalab, Kunhe Yang

机构 * Cornell University(康奈尔大学) UC Berkeley(加州大学伯克利分校)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17956 2025-05-27 cs.AI 85%

Unified Preference Optimization: Language Model Alignment Beyond the Preference Frontier

Anirudhan Badrinath, Prabhat Agarwal, Jiajing Xu

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI

Comments Accepted at Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11926 2025-05-20 cs.CV cs.AI 85%

SafeVid: Toward Safety Aligned Video Large Multimodal Models

Yixu Wang, Jiaxin Song, Yifeng Gao, Xin Wang, Yang Yao, Yan Teng, Xingjun Ma, Yingchun Wang, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 偏好对齐 :safety(title,abstract);alignment(abstract);DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17112 2025-04-01 cs.LG 85%

Decoding Human Preferences in Alignment: An Improved Approach to Inverse Constitutional AI

Carl-Leander Henneking, Claas Beger

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.LG

Comments 9 Pages, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13967 2025-03-04 cs.CL 85%

Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity

Rheeya Uppaal, Apratim Dey, Yiting He, Yiqiao Zhong, Junjie Hu

专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);safety(abstract);分类 cs.CL

Comments Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16681 2025-02-19 cs.CL 85%

Triple Preference Optimization: Achieving Better Alignment using a Single Step Optimization

Amir Saeidi, Shivanshu Verma, Aswin RRV, Kashif Rasul, Chitta Baral

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06874 2024-12-03 cs.AI cs.HC cs.RO 85%

Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment

Chenliang Li, Siliang Zeng, Zeyi Liao, Jiaxiang Li, Dongyeop Kang, Alfredo Garcia, Mingyi Hong

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏