arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 30405 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3229 篇

2409.14509 2025-03-05 cs.CL cs.CY cs.HC 81%

Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits

Tuhin Chakrabarty, Philippe Laban, Chien-Sheng Wu

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.CY

Comments ACM CHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05760 2025-02-28 cs.CV cs.AI cs.LG math.OC stat.ML 81%

Training-free Diffusion Model Alignment with Sampling Demons

Po-Hung Yeh, Kuang-Huei Lee, Jun-Cheng Chen

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 35 pages

Journal ref Proceedings of the Thirteenth International Conference on Learning Representations (ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15145 2025-02-25 cs.LG cs.AI 81%

Projection Optimization: A General Framework for Multi-Objective and Multi-Group RLHF

Nuoya Xiong, Aarti Singh

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14204 2025-02-21 cs.CL cs.AI 81%

On-the-fly Preference Alignment via Principle-Guided Decoding

Mingye Zhu, Yi Liu, Lei Zhang, Junbo Guo, Zhendong Mao

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00883 2025-02-21 cs.LG cs.CL 81%

SimPER: A Minimalist Approach to Preference Alignment without Hyperparameters

Teng Xiao, Yige Yuan, Zhengyu Chen, Mingxiao Li, Shangsong Liang, Zhaochun Ren, Vasant G Honavar

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19320 2025-02-20 cs.LG cs.AI stat.ML 81%

Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Shicong Cen, Jincheng Mei, Katayoon Goshvadi, Hanjun Dai, Tong Yang, Sherry Yang, Dale Schuurmans, Yuejie Chi, Bo Dai

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10545 2025-02-19 cs.LG cs.CL 81%

Efficient Alignment of Large Language Models via Data Sampling

Amrit Khera, Rajat Ghosh, Debojyoti Dutta

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments Original work accepted at NeurIPS Efficient Natural Language and Speech Processing Workshop. PMLR, 2024. Experiments with a larger model from a different family, Llama-30B have been added to the appendix for generalizability

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02628 2025-02-06 cs.LG cs.AI 81%

e-SimFT: Alignment of Generative Models with Simulation Feedback for Pareto-Front Design Exploration

Hyunmin Cheong, Mohammadmehdi Ataei, Amir Hosein Khasahmadi, Pradeep Kumar Jayaraman

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17117 2025-01-29 cs.CL cs.AI 81%

Histoires Morales: A French Dataset for Assessing Moral Alignment

Thibaud Leteno, Irina Proskurina, Antoine Gourru, Julien Velcin, Charlotte Laclau, Guillaume Metzler, Christophe Gravier

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02790 2025-01-07 cs.CL cs.AI 81%

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Yueqin Yin, Shentao Yang, Yujia Xie, Ziyi Yang, Yuting Sun, Hany Awadalla, Weizhu Chen, Mingyuan Zhou

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20834 2024-12-31 cs.CL cs.AI 81%

Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment

Jianfei Zhang, Jun Bai, Bei Li, Yanmeng Wang, Rumei Li, Chenghua Lin, Wenge Rong

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Coling 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13998 2024-12-19 cs.LG cs.AI 81%

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes

Katarzyna Kobalczyk, Claudio Fanconi, Hao Sun, Mihaela van der Schaar

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06000 2024-12-10 cs.CL cs.LG 81%

Does RLHF Scale? Exploring the Impacts From Data, Model, and Method

Zhenyu Hou, Pengfan Du, Yilin Niu, Zhengxiao Du, Aohan Zeng, Xiao Liu, Minlie Huang, Hongning Wang, Jie Tang, Yuxiao Dong

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04305 2024-12-06 cs.CL cs.LG 81%

ALMA: Alignment with Minimal Annotation

Michihiro Yasunaga, Leonid Shamis, Chunting Zhou, Andrew Cohen, Jason Weston, Luke Zettlemoyer, Marjan Ghazvininejad

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09223 2024-11-22 cs.CL cs.AI 81%

Word Alignment as Preference for Machine Translation

Qiyu Wu, Masaaki Nagata, Zhongtao Miao, Yoshimasa Tsuruoka

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2024 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11937 2024-11-20 cs.LG cs.AI 81%

Value Imprint: A Technique for Auditing the Human Values Embedded in RLHF Datasets

Ike Obi, Rohan Pant, Srishti Shekhar Agrawal, Maham Ghazanfar, Aaron Basiletti

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14758 2024-11-08 cs.GT cs.AI cs.LG 81%

Axioms for AI Alignment from Human Feedback

Luise Ge, Daniel Halpern, Evi Micha, Ariel D. Procaccia, Itai Shapira, Yevgeniy Vorobeychik, Junlin Wu

专题命中 偏好对齐 :alignment(title);RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09345 2024-11-04 cs.LG cs.AI 81%

InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling

Yuchun Miao, Sen Zhang, Liang Ding, Rong Bao, Lefei Zhang, Dacheng Tao

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

Comments The paper has been accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09385 2024-10-31 cs.CL cs.AI 81%

Reward Difference Optimization For Sample Reweighting In Offline RLHF

Shiqi Wang, Zhengze Zhang, Rui Zhao, Fei Tan, Cam Tu Nguyen

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2024 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22304 2024-10-30 cs.CL cs.LG 81%

Flow-DPO: Improving LLM Mathematical Reasoning through Online Multi-Agent Learning

Yihe Deng, Paul Mineiro

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL、cs.LG

Comments 5 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19206 2024-10-28 cs.LG cs.CL 81%

Inference time LLM alignment in single and multidomain preference spectrum

Sadat Shahriar, Zheng Qi, Nikolaos Pappas, Srikanth Doss, Monica Sunkara, Kishaloy Halder, Manuel Mager, Yassine Benajiba

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18533 2024-10-25 cs.CL cs.AI 81%

LOGO -- Long cOntext aliGnment via efficient preference Optimization

Zecheng Tang, Zechen Sun, Juntao Li, Qiaoming Zhu, Min Zhang

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09528 2024-10-25 cs.LG cs.AI 81%

Boosting Deductive Reasoning with Step Signals In RLHF

Jialian Li, Yipin Zhang, Wei Shen, Yuzi Yan, Jian Xie, Dong Yan

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00508 2024-10-15 cs.CL cs.AI 81%

FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization

Mingye Zhu, Yi Liu, Quan Wang, Junbo Guo, Zhendong Mao

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2024 Main track

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08639 2024-10-15 cs.AI cs.LG 81%

$β$-DPO: Direct Preference Optimization with Dynamic $β$

Junkang Wu, Yuexiang Xie, Zhengyi Yang, Jiancan Wu, Jinyang Gao, Bolin Ding, Xiang Wang, Xiangnan He

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08464 2024-10-08 cs.CL cs.AI 81%

Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Zhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng, Radha Poovendran, Yejin Choi, Bill Yuchen Lin

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Link: https://magpie-align.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01789 2024-10-03 cs.LG cs.AI 81%

Investigating on RLHF methodology

Alexey Kutalev, Sergei Markoff

专题命中 偏好对齐 :RLHF(title);alignment(abstract);分类 cs.AI、cs.LG

Comments 23 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18995 2024-10-01 cs.CL cs.AI 81%

Systematic Characterization of the Effectiveness of Alignment in Large Language Models for Categorical Decisions

Isaac Kohane

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments 19 pages (without Appendix) Appendix 7 pages. 7 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10392 2024-08-21 cs.CL cs.LG 81%

Value Alignment from Unstructured Text

Inkit Padhi, Karthikeyan Natesan Ramamurthy, Prasanna Sattigeri, Manish Nagireddy, Pierre Dognin, Kush R. Varshney

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04655 2024-08-13 cs.CL cs.AI 81%

Strong and weak alignment of large language models with human values

Mehdi Khamassi, Marceau Nahon, Raja Chatila

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted for publication in Scientific Reports, special issue on AI aligment

详情

展开后加载摘要…

URL PDF HTML 收藏