arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 30405 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3229 篇

2602.02824 2026-02-04 cs.CL 83%

CATNIP: LLM Unlearning via Calibrated and Tokenized Negative Preference Alignment

CATNIP:通过校准和令牌化的负偏好对齐实现LLM去学习

Zhengbang Yang, Yisheng Zhong, Junyuan Hong, Zhuangdi Zhu

机构 * George Mason University(乔治·马歇尔大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL

AI总结 CATNIP通过校准和令牌化的负偏好对齐方法,实现有效LLM去学习,无需保留数据或对比对,提升知识遗忘与保留的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10416 2026-01-16 cs.AI 83%

LLMdoctor: Token-Level Flow-Guided Preference Optimization for Efficient Test-Time Alignment of Large Language Models

LLMdoctor: 基于令牌级流引导的偏好优化用于大语言模型高效测试时间对齐

Tiesunlong Shen, Rui Mao, Jin Wang, Heming Sun, Jian Zhang, Xuejie Zhang, Erik Cambria

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.AI

AI总结 LLMdoctor通过令牌级流引导偏好优化,实现高效的大语言模型测试时间对齐,提升对齐精度并保留生成多样性。

Comments Accepted by AAAI26

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06596 2026-01-13 cs.CR cs.AI 83%

Are LLMs Vulnerable to Preference-Undermining Attacks (PUA)? A Factorial Analysis Methodology for Diagnosing the Trade-off between Preference Alignment and Real-World Validity

大型语言模型是否易受偏好削弱攻击(PUA)攻击?一种诊断偏好对齐与现实有效性之间权衡的因子分析方法

Hongjun An, Yiliang Song, Jiangan Chen, Jiawei Shao, Chi Zhang, Xuelong Li

机构 * School of Artificial Intelligence, OPtics and ElectroNics, Northwestern Polytechnical University(人工智能学院、光学与电子学院、西北工业大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) School of Economics and Management, Guangxi Normal University(经济管理学院,广西师范大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI

AI总结 本文提出了一种因子分析方法,用于诊断LLM在偏好对齐与现实有效性之间的权衡,揭示了高级模型可能更易受操纵性提示攻击,并强调了定制防御的重要性。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05075 2026-01-09 cs.CL 83%

SemPA: Improving Sentence Embeddings of Large Language Models through Semantic Preference Alignment

SemPA:通过语义偏好对齐提升大语言模型的句子嵌入

Ziyang Chen, Zhenxuan Huang, Yile Wang, Weiqin Wang, Lu Yin, Hui Huang

机构 * College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) School of Computer Science and Electronic Engineering, University of Surrey(Surrey大学计算机科学与电子工程学院)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

AI总结 SemPA通过语义偏好对齐提升大语言模型的句子嵌入能力,同时保持其生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03589 2026-01-08 cs.CL 83%

OLA: Output Language Alignment in Code-Switched LLM Interactions

OLA:代码切换交互中的输出语言对齐

Juhyun Oh, Haneul Yoo, Faiz Ghifari Haznitrama, Alice Oh

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

AI总结 OLA研究了代码切换交互中LLM输出语言对齐问题,发现现有模型存在隐含语言期望识别偏差,通过Code-Switching Aware DPO方法显著提升对齐效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00647 2026-01-05 cs.CL cs.CE q-bio.QM 83%

Physio-DPO: Aligning Large Language Models with the Protein Energy Landscape to Eliminate Structural Hallucinations

Physio-DPO:将大型语言模型与蛋白质能量景观对齐以消除结构幻觉

QiWei Meng

机构 * Xi’an Jiaotong University(西安交通大学)

专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.CL

AI总结 Physio-DPO通过引入物理感知的目标,将蛋白质语言模型与热力学稳定性对齐,有效减少结构幻觉并提升折叠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11518 2026-01-05 cs.CL 83%

W2S-AlignTree: Weak-to-Strong Inference-Time Alignment for Large Language Models via Monte Carlo Tree Search

W2S-AlignTree: 通过蒙特卡洛树搜索实现大语言模型的弱到强推理时对齐

Zhenyu Ding, Yuhao Wang, Tengyue Xiao, Haoying Wang, Caigui Jiang, Ning Ding

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

AI总结 W2S-AlignTree通过结合MCTS与弱到强泛化范式,实现大语言模型在推理时的高效对齐,提升摘要任务性能15.9%。

Comments AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15843 2026-01-01 cs.CL 83%

Pre-DPO: Improving Data Utilization in Direct Preference Optimization Using a Guiding Reference Model

Pre-DPO:通过引导参考模型提升直接偏好优化的数据利用

Junshu Pan, Wei Shen, Shulin Huang, Qiji Zhou, Yue Zhang

专题命中 偏好对齐 :DPO(title,abstract);RLHF(abstract);分类 cs.CL

AI总结 Pre-DPO通过引入引导参考模型提升直接偏好优化的数据利用效率,有效提高大语言模型的性能表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04453 2025-12-23 cs.LG 83%

ESSA: Evolutionary Strategies for Scalable Alignment

ESSA:可扩展对齐的进化策略

Daria Korotyshova, Boris Shaposhnikov, Alexey Malakhov, Alexey Khokhulin, Nikita Surnachev, Kirill Ovcharenko, George Bredis, Alexey Gorbatovski, Viacheslav Sinii, Daniil Gavrilov

机构 * T-Tech

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.LG

AI总结 ESSA通过进化策略实现大规模LLM对齐,无需梯度优化,提升模型准确率并减少计算开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13837 2025-12-17 cs.LG 83%

Explainable reinforcement learning from human feedback to improve alignment

通过人类反馈改进强化学习的可解释性以提升对齐

Shicheng Liu, Siyuan Xu, Wenjie Qiu, Hangfan Zhang, Minghui Zhu

机构 * Department of Electrical Engineering, Pennsylvania State University(宾夕法尼亚州立大学电气工程系) Department of Computer Science, Rutgers University(罗格斯大学计算机科学系) College of Information Sciences and Technology, Pennsylvania State University(宾夕法尼亚州立大学信息科学与技术学院)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.LG

AI总结 本文提出通过纠正RLHF中不满意响应的原因来提升语言模型对齐性的方法,结合事后解释与卸载技术改进模型响应质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11998 2025-12-16 cs.CL 83%

Direct Confidence Alignment: Aligning Verbalized Confidence with Internal Confidence In Large Language Models

直接置信对齐:将 verbalized 置信度与内部置信度对齐于大语言模型

Glenn Zhang, Treasure Mayowa, Jason Fan, Yicheng Fu, Aaron Sandoval, Sean O'Brien, Kevin Zhu

机构 * Algoverse AI Research(Algoverse AI研究院)

专题命中 偏好对齐 :alignment(title,abstract);trustworthy(abstract);分类 cs.CL

AI总结 本文提出直接置信对齐方法,通过优化使LLM的 verbalized 置信度与内部置信度对齐,提升模型透明度和可靠性。

Comments Accepted at ACL 2025 SRW, 5 pages body, 14 pages total

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06515 2025-12-09 cs.CL 83%

ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models

ProSocialAlign:语言模型中的偏好条件测试时间对齐

Somnath Banerjee, Sayan Layek, Sayantan Adak, Mykola Pechenizkiy, Animesh Mukherjee, Rima Hazra

机构 * Cisco Research(思科研究) Indian Institute of Technology Kharagpur(印度理工学院卡里格普尔分校) Eindhoven University of Technology, Netherlands (TU/e)(埃因霍温理工大学(荷兰))

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL

AI总结 ProSocialAlign通过测试时间参数高效方法,在不重新训练模型的情况下,引导生成安全、富有同理心且价值观一致的响应,提升语言模型的安全性和对齐性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16929 2025-12-09 cs.CV cs.AI 83%

TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment

TEMPLE:通过渐进式预SFT对齐激励视频大语言模型的时序理解

Shicheng Li, Lei Li, Kun Ouyang, Shuhuai Ren, Yuanxin Liu, Yuanxing Zhang, Fuzheng Zhang, Lingpeng Kong, Qi Liu, Xu Sun

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.AI

AI总结 TEMPLE通过渐进式预SFT对齐策略提升视频大语言模型的时序理解能力,利用直接偏好优化和课程学习增强模型对时间信息的感知。

Comments Accepted to AAAI 2026. Code available at https://github.com/lscpku/TEMPLE

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23316 2025-12-04 cs.CL 83%

Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPO

近端化偏好优化用于多样化反馈类型:对DPO的分解视角

Kaiyang Guo, Yinchuan Li, Zhitang Chen

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.CL

AI总结 本文提出PRO方法,通过分解DPO损失并恢复完整正则化项,解决似然不足确定性问题,提升对多样化反馈类型的适应能力。

Comments NeurIPS'2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19387 2025-11-27 cs.LG cs.SY eess.SY math.OC 83%

Alignment of large language models with constrained learning

受限学习下的大语言模型对齐

Botong Zhang, Shuo Li, Ignacio Hounie, Osbert Bastani, Dongsheng Ding, Alejandro Ribeiro

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.LG

AI总结 本文提出了一种基于拉格朗日对偶性的迭代对偶对齐方法,用于在满足约束条件下优化大语言模型策略,通过实验验证了其在受限对齐问题中的有效性。

Comments 51 pages, 5 figures, 11 tables; Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02709 2025-11-19 cs.LG 83%

Preference Robustness for DPO with Applications to Public Health

Cheol Woo Kim, Shresth Verma, Mauricio Tec, Milind Tambe

专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09385 2025-11-18 cs.CL 83%

AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment

Ruibo Deng, Duanyu Feng, Wenqiang Lei

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

Comments AAAI 2026 AIA oral, our code is available at https://github.com/Shiroha-Offical/AMaPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02095 2025-11-04 cs.CV cs.LG 83%

Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences

Hyojin Bahng, Caroline Chan, Fredo Durand, Phillip Isola

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21794 2025-10-28 cs.CV cs.AI 83%

Token-Level Inference-Time Alignment for Vision-Language Models

Kejia Chen, Jiawen Zhang, Jiacong Hu, Kewei Gao, Jian Lou, Zunlei Feng, Mingli Song

机构 * Zhejiang University(浙江大学) Sun Yat-sen University(中山大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07520 2025-10-24 cs.SD cs.AI eess.AS 83%

LeVo: High-Quality Song Generation with Multi-Preference Alignment

Shun Lei, Yaoxun Xu, Zhiwei Lin, Huaicheng Zhang, Wei Tan, Hangting Chen, Jianwei Yu, Yixuan Zhang, Chenyu Yang, Haina Zhu, Shuai Wang, Zhiyong Wu, Dong Yu

机构 * Shenzhen International Graduate School, Tsinghua University, Shenzhen(清华大学深圳国际研究生院) Tencent AI Lab(腾讯AI实验室) Wuhan University(武汉大学) The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen)(香港中文大学(深圳)) X-LANCE Lab, Shanghai Jiao Tong University, Shanghai(上海交通大学X-LANCE实验室) School of Intelligence Science and Technology, Nanjing University, Suzhou, China(南京大学智能科学与技术学院)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01203 2025-10-21 cs.LG stat.ML 83%

KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity

Gholamali Aminian, Amir R. Asadi, Idan Shenfeld, Youssef Mroueh

机构 * The Alan Turing Institute(艾伦·图灵研究所) Statistical Laboratory(统计实验室) University of Cambridge(剑桥大学) Massachusetts Institute of Technology(麻省理工学院) IBM Research USA(IBM美国研究)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments Extra experiments are added in new version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13694 2025-10-16 cs.LG 83%

Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking

Yuchun Miao, Liang Ding, Sen Zhang, Rong Bao, Lefei Zhang, Dacheng Tao

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments 46 pages, 36 figures, submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12195 2025-10-15 cs.CL 83%

DPO-Tuned Large Language Models for Segmentation in Simultaneous Speech Translation

Zeyu Yang, Satoshi Nakamura

专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02850 2025-10-06 cs.AI 83%

Reward Model Routing in Alignment

Xinle Wu, Yao Lu

机构 * National University of Singapore(新加坡国立大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08541 2025-09-16 cs.CL 83%

CM-Align: Consistency-based Multilingual Alignment for Large Language Models

Xue Zhang, Yunlong Liang, Fandong Meng, Songming Zhang, Yufeng Chen, Jinan Xu, Jie Zhou

机构 * Key Laboratory of Big Data & Artificial Intelligence in Transportation, Beijing Jiaotong University, Ministry of Education(大数据与人工智能交通运输 key laboratory,北京交通大学,教育部) School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China(计算机科学与技术学院,北京交通大学,北京,中国) Pattern Recognition Center, WeChat AI, Tencent Inc, China(模式识别中心,微信AI,腾讯公司,中国)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20655 2025-09-12 cs.CV cs.CL 83%

Improving Alignment in LVLMs with Debiased Self-Judgment

Sihan Yang, Chenhang Cui, Zihao Zhao, Yiyang Zhou, Weilong Yan, Ying Wei, Huaxiu Yao

机构 * Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学) UNC-Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02949 2025-09-09 cs.CL 83%

RADIANT: Retrieval AugmenteD entIty-context AligNmenT -- Introducing RAG-ability and Entity-Context Divergence

Vipula Rawte, Rajarshi Roy, Gurpreet Singh, Danush Khanna, Yaswanth Narsupalli, Basab Ghosh, Abhay Gupta, Argha Kamal Samanta, Aditya Shingote, Aadi Krishna Vikram, Vinija Jain, Aman Chadha, Amit Sheth, Amitava Das

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00309 2025-09-03 cs.CL 83%

Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models

Chen Zheng, Yiyuan Ma, Yuan Yang, Deyi Liu, Jing Liu, Zuquan Song, Yuxin Song, Cheng Ren, Hang Zhu, Xin Liu, Yiyuan Ma, Siyuan Qiao, Xun Zhou, Liang Xiang, Yonghui Wu

机构 * Peking University(北京大学)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16982 2025-08-26 cs.CL 83%

Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens

Ilias Chalkidis

机构 * Department of Computer Science, University of Copenhagen(哥本哈根大学计算机科学系)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

Comments This is a working paper and will be updated with new information or corrections based on community feedback

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16455 2025-08-26 stat.ML cs.LG stat.ME 83%

On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization

Jiancong Xiao, Ziniu Li, Xingyu Xie, Emily Getzen, Cong Fang, Qi Long, Weijie J. Su

机构 * University of Pennsylvania(宾夕法尼亚大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) National University of Singapore(新加坡国立大学) Peking University(北京大学) Joint corresponding authors(联合通讯作者)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments Accepted for publication in the Journal of the American Statistical Association

详情

展开后加载摘要…

URL PDF HTML 收藏