arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 30405 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3229 篇

2404.05530 2024-08-07 cs.CL cs.AI cs.CR cs.LG 82%

Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data

Tim Baumgärtner, Yang Gao, Dana Alon, Donald Metzler

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18676 2024-07-19 cs.CL cs.AI cs.LG 82%

Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation

Guanting Dong, Yutao Zhu, Chenghao Zhang, Zechen Wang, Zhicheng Dou, Ji-Rong Wen

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13173 2024-07-17 cs.CV cs.AI cs.CL cs.LG 82%

Biomedical Visual Instruction Tuning with Clinician Preference Alignment

Hejie Cui, Lingjun Mao, Xin Liang, Jieyu Zhang, Hui Ren, Quanzheng Li, Xiang Li, Carl Yang

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13213 2024-07-08 cs.LG cs.CL cs.CY 82%

From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards

Khaoula Chehbouni, Megha Roshan, Emmanuel Ma, Futian Andrew Wei, Afaf Taik, Jackie CK Cheung, Golnoosh Farnadi

专题命中 偏好对齐 :safety(title,abstract);分类 cs.CL、cs.CY、cs.LG

Comments 9 pages, 4 figures. Accepted to Findings of the Association for Computational Linguistics: ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02552 2024-07-04 cs.CL cs.AI cs.LG 82%

RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs

John Dang, Arash Ahmadian, Kelly Marchisio, Julia Kreutzer, Ahmet Üstün, Sara Hooker

专题命中 偏好对齐 :RLHF(title);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07971 2024-06-14 cs.CL cs.AI cs.LG 82%

It Takes Two: On the Seamlessness between Reward and Policy Model in RLHF

Taiming Lu, Lingfeng Shen, Xinyu Yang, Weiting Tan, Beidi Chen, Huaxiu Yao

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10958 2024-05-29 cs.CL cs.AI cs.LG 82%

Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts

Yueqin Yin, Zhendong Wang, Yi Gu, Hai Huang, Weizhu Chen, Mingyuan Zhou

专题命中 偏好对齐 :alignment(title);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.06639 2024-05-13 cs.LG cs.AI cs.CL 82%

Value Augmented Sampling for Language Model Alignment and Personalization

Seungwook Han, Idan Shenfeld, Akash Srivastava, Yoon Kim, Pulkit Agrawal

专题命中 偏好对齐 :alignment(title);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Website: https://sites.google.com/view/llm-vas

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08555 2024-04-17 cs.LG cs.AI cs.CL 82%

RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs

Shreyas Chaudhari, Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande, Bruno Castro da Silva

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06326 2024-03-12 cs.CL cs.AI cs.LG 82%

From Instructions to Constraints: Language Model Alignment with Automatic Constraint Verification

Fei Wang, Chao Shang, Sarthak Jain, Shuai Wang, Qiang Ning, Bonan Min, Vittorio Castelli, Yassine Benajiba, Dan Roth

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10529 2024-03-01 cs.CL cs.AI cs.CY 82%

Evaluating Psychological Safety of Large Language Models

Xingxuan Li, Yutong Li, Lin Qiu, Shafiq Joty, Lidong Bing

专题命中 偏好对齐 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments Preprint. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06452 2024-02-20 cs.LG cs.AI cs.CL 82%

Understanding the Effects of RLHF on LLM Generalisation and Diversity

Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, Roberta Raileanu

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Code available here: https://github.com/facebookresearch/rlfh-gen-div

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07319 2024-02-13 cs.LG cs.AI cs.CL 82%

ODIN: Disentangled Reward Mitigates Hacking in RLHF

Lichang Chen, Chen Zhu, Davit Soselia, Jiuhai Chen, Tianyi Zhou, Tom Goldstein, Heng Huang, Mohammad Shoeybi, Bryan Catanzaro

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16335 2024-01-30 cs.LG cs.AI cs.CL stat.ML 82%

Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

Banghua Zhu, Michael I. Jordan, Jiantao Jiao

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01320 2023-08-04 cs.LG cs.AI cs.CL 82%

DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales

Zhewei Yao, Reza Yazdani Aminabadi, Olatunji Ruwase, Samyam Rajbhandari, Xiaoxia Wu, Ammar Ahmad Awan, Jeff Rasley, Minjia Zhang, Conglong Li, Connor Holmes, Zhongzhu Zhou, Michael Wyatt, Molly Smith, Lev Kurilenko, Heyang Qin, Masahiro Tanaka, Shuai Che, Shuaiwen Leon Song, Yuxiong He

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.06859 2023-01-18 cs.HC cs.AI cs.CL cs.LG 82%

Methodological reflections for AI alignment research using human feedback

Thilo Hagendorff, Sarah Fabi

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12205 2026-01-23 cs.LG cs.AI cs.IT math.IT math.ST stat.ML stat.TH 82%

On the Exponential Convergence for Offline RLHF with Pairwise Comparisons

关于通过成对比较进行离线RLHF的指数收敛性

Zhirui Chen, Vincent Y. F. Tan

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG;alignment(comments)

AI总结 本文提出了一种在离线RLHF中通过成对比较实现指数收敛的算法RL-LOW,并推导了实例依赖的下界。

Comments Accepted as an oral presentation at AAAI 2026 (AI Alignment Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23630 2024-11-01 cs.LG cs.AI 82%

Adaptive Alignment: Dynamic Preference Adjustments via Multi-Objective Reinforcement Learning for Pluralistic AI

Hadassah Harland, Richard Dazeley, Peter Vamplew, Hashini Senaratne, Bahareh Nakisa, Francisco Cruz

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments Accepted for the Pluralistic Alignment workshop at NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00861 2021-12-13 cs.CL cs.LG 82%

A General Language Assistant as a Laboratory for Alignment

Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Jared Kaplan

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments 26+19 pages; v2 typos fixed, refs added, figure scale / colors fixed; v3 correct very non-standard TruthfulQA formatting and metric, alignment implications slightly improved

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08491 2026-08-11 cs.AI 新提交 81%

TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models

TrustRoboReward:面向多范式机器人奖励模型的偏好有序保序分数编辑方法

Yidong Wang, Yan Zhan, Ziteng Feng, Zhenyu Cui, Ziyi Zhou, Renzhao Liang, Jiaxuan Zhu, Zilei Yang, Yiran Zhao, Zhongkuan Mao, Bo Jia, Hanchu Ni, Chenggang Xie, Biao Liu, Yi Zhang, Yong Dai, Xiaozhu Ju, Wei Ye, Shikun Zhang

机构 * Peking University(北京大学) Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) University of Science and Technology of China(中国科学技术大学) Southeast University(东南大学) Southern University of Science and Technology(南方科技大学) Beijing University of Aeronautics and Astronautics(北京航空航天大学) Beijing Language and Culture University(北京语言大学) Sichuan University(四川大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 偏好对齐 :RLHF(abstract,abstract_cn);DPO(abstract,abstract_cn);分类 cs.AI

AI总结 针对现有机器人奖励模型的跨范式偏好与分数不一致问题,本文提出 TrustRoboReward 框架,通过 POISE 方法解决反转冲突,训练的 Qwen3-VL-4B 性能接近 GPT-5-mini,优于 RoboReward 基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06526 2026-08-10 cs.CL 新提交 81%

GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy Optimization

GRASP:用组相对策略优化增强语言模型匿名化器

Sajjad Ghiasvand, Nader Sehatbakhsh

机构 * UC Santa Barbara(加州大学圣巴巴拉分校) UC Los Angeles(加州大学洛杉矶分校)

专题命中 偏好对齐 :DPO(summary_cn,abstract);分类 cs.CL

AI总结 该研究提出GRASP方法,用组相对策略优化增强设备端小型语言模型匿名化器,在隐私-效用权衡、对抗匿名化防御及运行成本上均优于DPO蒸馏基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05341 2026-08-07 cs.CV cs.LG 新提交 81%

Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation

面向胸部X射线报告生成的正例-无标签偏好优化

Yuta Kobayashi, Pradyun Ramesh, Muhammad Ahmed Chaudhry, Vincent Jeanselme, Judy Wawira Gichoya, Sanmi Koyejo, Kathleen Capaccione, Shalmali Joshi

机构 * Columbia University(哥伦比亚大学) Stanford University(斯坦福大学) Emory University(埃默里大学)

专题命中 偏好对齐 :DPO(summary_cn,abstract);分类 cs.LG

AI总结 该研究针对放射报告生成VLM的遗漏噪声问题,提出PU-DPO框架,将未提及项视为无标签,通过对比对优化,提升病理检测率与隐藏正例恢复能力,增强对遗漏噪声的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05385 2026-07-24 cs.IR cs.AI 版本更新 81%

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework

TeaRAG:一种令牌高效的智能检索增强生成框架

Chao Zhang, Yuhao Wang, Derong Xu, Haoxin Zhang, Yuanjie Lyu, Yuhao Chen, Shuochen Liu, Tong Xu, Xiangyu Zhao, Yan Gao, Yao Hu, Enhong Chen

机构 * University of Science and Technology of China(中国科学技术大学) City University of Hong Kong(香港城市大学) Xiaohongshu Inc.(小红书公司)

专题命中 偏好对齐 :DPO(summary_cn,abstract);分类 cs.AI

AI总结 研究提出TeaRAG框架解决智能RAG令牌开销大问题,通过基于简洁三元组的图检索压缩检索内容,用IP-DPO减少推理步骤,在多个数据集上实验,提高了平均精确匹配,减少了输出令牌。

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18091 2026-07-21 cs.CV cs.GR cs.LG 新提交 81%

SciForma: Structure-Faithful Generation of Scientific Diagrams

SciForma:科学图表的结构忠实生成

Yuxuan Luo, Peng Zhang, Xinjie Zhang, Xun Guo, Zhouhui Lian, Yan Lu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机技术研究所) State Key Lab of CAD & CG, Zhejiang University(浙江大学CAD&CG国家重点实验室) Microsoft Research Asia(微软亚洲研究院)

专题命中 偏好对齐 :DPO(summary_cn,abstract);分类 cs.LG

AI总结 研究针对科学图表结构保真度问题,提出SciForma框架,分解图表质量为三个结构轴,用M-DPO优化,策划训练和评估数据,实现迭代编辑,使SciForma-9B超越开源基线和GPT-Image-1.5,提升科学图表生成的结构保真度。

Comments 30 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04136 2026-07-07 cs.AI 版本更新 81%

A Technical Survey of Reinforcement Learning Techniques for Large Language Models

大语言模型强化学习技术的技术综述

Saksham Sahai Srivastava, Vaneet Aggarwal

机构 * University of Georgia(佐治亚大学) Purdue University(普渡大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);safety(abstract)

AI总结 综述强化学习与语言模型整合,介绍如近端策略优化等算法,分析其在各领域应用,对失败模式算法分析,给出分类法,评估趋势并探讨挑战与新兴方向,为研究提供路线图。

Comments Accepted to ACM Transactions on Intelligent Systems and Technology, June 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06625 2026-07-01 cs.CL 版本更新 81%

FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge

FairJudge: 一种自适应、去偏且一致的LLM作为评判者

Bo Yang, Lanfei Feng, Yunkui Chen, Yu Zhang, Xiao Xu, Shijian Li

机构 * School of Software Technology, Zhejiang University(浙江大学软件学院)

专题命中 偏好对齐 :DPO(summary_cn,abstract);分类 cs.CL

AI总结 针对LLM评判系统在适应性、偏见和一致性上的局限,提出FairJudge,通过可学习正则化策略和课程式SFT-DPO-GRPO训练范式,实现自适应、去偏和一致评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28898 2026-06-30 cs.CL 81%

PASTA: A Paraphrasing And Self-Training Approach for Knowledge Updating in LLMs

PASTA: 一种用于大语言模型知识更新的释义与自训练方法

Takayuki Yamamoto, Daisuke Kawahara

机构 * Waseda University(早稻田大学)

专题命中 偏好对齐 :DPO(summary_cn,abstract);分类 cs.CL

AI总结 提出PASTA框架,通过数据增强、问答生成和自学习DPO过程,将新闻事实知识集成到LLM中,实现知识覆盖与幻觉抑制,准确率从0.02提升至0.82。

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24993 2026-06-25 cs.LG 新提交 81%

The Geometry of Sequential Learning: Lie-Bracket Prediction of Transfer Order

序列学习的几何:李括号预测迁移顺序

John Sweeney

机构 * Sideplane AI

专题命中 偏好对齐 :DPO(summary_cn,abstract);分类 cs.LG

AI总结 提出李括号对偶性分数预测序列学习中的源域顺序,通过李括号锦标赛实现O(N log N)排序,在指令微调和DPO中达到高准确率。

Comments Accepted to ICML 2026. 20 pages, including appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19607 2026-06-19 cs.AI stat.AP 新提交 81%

Which Pairs to Compare for LLM Post-Training?

LLM后训练中应比较哪些对?

Jiangze Han, Vineet Goyal, Will Ma

机构 * Columbia University(哥伦比亚大学)

专题命中 偏好对齐 :DPO(summary_cn,abstract);分类 cs.AI

AI总结 研究偏好后训练中如何选择最具信息量的比较对,提出基于采样设计的比较策展方法,通过DPO训练的理论分析给出优化准则,实验证明能提升样本效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18961 2026-06-18 cs.LG 新提交 81%

Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization

做自己的老师:通过无监督奖励优化引导蛋白质语言模型

Lanqing Li, Shentong Mo, Yang Yu, Pheng-Ann Heng

机构 * The Chinese University of Hong Kong(香港中文大学) MBZUAI Hong Kong University of Science and Technology(香港科学理工大学)

专题命中 偏好对齐 :RLHF(abstract,abstract_cn);DPO(abstract,abstract_cn);分类 cs.LG

AI总结 提出无监督奖励优化框架,结合模型不确定性和语义一致性作为代理奖励,通过SRO和BRO算法优化PLMs,在无标签数据下实现可控蛋白质生成,性能接近有监督方法。

Comments 24 pages, 2 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏