arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3235 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3235 篇

2503.21878 2025-04-09 cs.AI cs.LG stat.ML 81%

Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment

Audrey Huang, Adam Block, Qinghua Liu, Nan Jiang, Akshay Krishnamurthy, Dylan J. Foster

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19201 2025-03-26 cs.LG cs.AI 81%

A Shared Low-Rank Adaptation Approach to Personalized RLHF

Renpu Liu, Peng Wang, Donghao Li, Cong Shen, Jing Yang

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

Comments Published as a conference paper at AISTATS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18130 2025-03-25 cs.LG cs.AI 81%

Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization

Juntao Dai, Taiye Chen, Yaodong Yang, Qian Zheng, Gang Pan

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

Comments Published as a conference paper at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17965 2025-03-25 cs.CL cs.AI 81%

Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts

Beining Xu, Arkaitz Zubiaga

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI

Comments 14 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15880 2025-03-21 cs.LG cs.CL 81%

InCo-DPO: Balancing Distribution Shift and Data Quality for Enhanced Preference Optimization

Yunan Wang, Jijie Li, Bo-Wen Zhang, Liangdong Wang, Guang Liu

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04070 2025-03-14 cs.CL cs.AI 81%

PAD: Personalized Alignment of LLMs at Decoding-Time

Ruizhe Chen, Xiaotian Zhang, Meng Luo, Wenhao Chai, Zuozhu Liu

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14509 2025-03-05 cs.CL cs.CY cs.HC 81%

Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits

Tuhin Chakrabarty, Philippe Laban, Chien-Sheng Wu

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.CY

Comments ACM CHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05760 2025-02-28 cs.CV cs.AI cs.LG math.OC stat.ML 81%

Training-free Diffusion Model Alignment with Sampling Demons

Po-Hung Yeh, Kuang-Huei Lee, Jun-Cheng Chen

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 35 pages

Journal ref Proceedings of the Thirteenth International Conference on Learning Representations (ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15145 2025-02-25 cs.LG cs.AI 81%

Projection Optimization: A General Framework for Multi-Objective and Multi-Group RLHF

Nuoya Xiong, Aarti Singh

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14204 2025-02-21 cs.CL cs.AI 81%

On-the-fly Preference Alignment via Principle-Guided Decoding

Mingye Zhu, Yi Liu, Lei Zhang, Junbo Guo, Zhendong Mao

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00883 2025-02-21 cs.LG cs.CL 81%

SimPER: A Minimalist Approach to Preference Alignment without Hyperparameters

Teng Xiao, Yige Yuan, Zhengyu Chen, Mingxiao Li, Shangsong Liang, Zhaochun Ren, Vasant G Honavar

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19320 2025-02-20 cs.LG cs.AI stat.ML 81%

Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Shicong Cen, Jincheng Mei, Katayoon Goshvadi, Hanjun Dai, Tong Yang, Sherry Yang, Dale Schuurmans, Yuejie Chi, Bo Dai

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10545 2025-02-19 cs.LG cs.CL 81%

Efficient Alignment of Large Language Models via Data Sampling

Amrit Khera, Rajat Ghosh, Debojyoti Dutta

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments Original work accepted at NeurIPS Efficient Natural Language and Speech Processing Workshop. PMLR, 2024. Experiments with a larger model from a different family, Llama-30B have been added to the appendix for generalizability

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02628 2025-02-06 cs.LG cs.AI 81%

e-SimFT: Alignment of Generative Models with Simulation Feedback for Pareto-Front Design Exploration

Hyunmin Cheong, Mohammadmehdi Ataei, Amir Hosein Khasahmadi, Pradeep Kumar Jayaraman

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17117 2025-01-29 cs.CL cs.AI 81%

Histoires Morales: A French Dataset for Assessing Moral Alignment

Thibaud Leteno, Irina Proskurina, Antoine Gourru, Julien Velcin, Charlotte Laclau, Guillaume Metzler, Christophe Gravier

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02790 2025-01-07 cs.CL cs.AI 81%

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Yueqin Yin, Shentao Yang, Yujia Xie, Ziyi Yang, Yuting Sun, Hany Awadalla, Weizhu Chen, Mingyuan Zhou

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20834 2024-12-31 cs.CL cs.AI 81%

Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment

Jianfei Zhang, Jun Bai, Bei Li, Yanmeng Wang, Rumei Li, Chenghua Lin, Wenge Rong

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Coling 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13998 2024-12-19 cs.LG cs.AI 81%

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes

Katarzyna Kobalczyk, Claudio Fanconi, Hao Sun, Mihaela van der Schaar

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06000 2024-12-10 cs.CL cs.LG 81%

Does RLHF Scale? Exploring the Impacts From Data, Model, and Method

Zhenyu Hou, Pengfan Du, Yilin Niu, Zhengxiao Du, Aohan Zeng, Xiao Liu, Minlie Huang, Hongning Wang, Jie Tang, Yuxiao Dong

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04305 2024-12-06 cs.CL cs.LG 81%

ALMA: Alignment with Minimal Annotation

Michihiro Yasunaga, Leonid Shamis, Chunting Zhou, Andrew Cohen, Jason Weston, Luke Zettlemoyer, Marjan Ghazvininejad

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09223 2024-11-22 cs.CL cs.AI 81%

Word Alignment as Preference for Machine Translation

Qiyu Wu, Masaaki Nagata, Zhongtao Miao, Yoshimasa Tsuruoka

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2024 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11937 2024-11-20 cs.LG cs.AI 81%

Value Imprint: A Technique for Auditing the Human Values Embedded in RLHF Datasets

Ike Obi, Rohan Pant, Srishti Shekhar Agrawal, Maham Ghazanfar, Aaron Basiletti

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14758 2024-11-08 cs.GT cs.AI cs.LG 81%

Axioms for AI Alignment from Human Feedback

Luise Ge, Daniel Halpern, Evi Micha, Ariel D. Procaccia, Itai Shapira, Yevgeniy Vorobeychik, Junlin Wu

专题命中 偏好对齐 :alignment(title);RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09345 2024-11-04 cs.LG cs.AI 81%

InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling

Yuchun Miao, Sen Zhang, Liang Ding, Rong Bao, Lefei Zhang, Dacheng Tao

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

Comments The paper has been accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09385 2024-10-31 cs.CL cs.AI 81%

Reward Difference Optimization For Sample Reweighting In Offline RLHF

Shiqi Wang, Zhengze Zhang, Rui Zhao, Fei Tan, Cam Tu Nguyen

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2024 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22304 2024-10-30 cs.CL cs.LG 81%

Flow-DPO: Improving LLM Mathematical Reasoning through Online Multi-Agent Learning

Yihe Deng, Paul Mineiro

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL、cs.LG

Comments 5 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19206 2024-10-28 cs.LG cs.CL 81%

Inference time LLM alignment in single and multidomain preference spectrum

Sadat Shahriar, Zheng Qi, Nikolaos Pappas, Srikanth Doss, Monica Sunkara, Kishaloy Halder, Manuel Mager, Yassine Benajiba

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18533 2024-10-25 cs.CL cs.AI 81%

LOGO -- Long cOntext aliGnment via efficient preference Optimization

Zecheng Tang, Zechen Sun, Juntao Li, Qiaoming Zhu, Min Zhang

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09528 2024-10-25 cs.LG cs.AI 81%

Boosting Deductive Reasoning with Step Signals In RLHF

Jialian Li, Yipin Zhang, Wei Shen, Yuzi Yan, Jian Xie, Dong Yan

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00508 2024-10-15 cs.CL cs.AI 81%

FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization

Mingye Zhu, Yi Liu, Quan Wang, Junbo Guo, Zhendong Mao

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2024 Main track

详情

展开后加载摘要…

URL PDF HTML 收藏