arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3229 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3229 篇

2508.08466 2025-08-13 cs.CL 83%

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints

Daren Yao, Jinsong Yuan, Ruike Chen

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18802 2025-07-28 cs.HC cs.AI 83%

DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition

Danqing Shi, Furui Cheng, Tino Weinkauf, Antti Oulasvirta, Mennatallah El-Assady

机构 * Aalto University(阿alto大学) ETH Zürich(苏黎世联邦理工学院) KTH Royal Institute of Technology(皇家理工学院)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04427 2025-07-09 cs.CL 83%

One fish, two fish, but not the whole sea: Alignment reduces language models' conceptual diversity

Sonia K. Murthy, Tomer Ullman, Jennifer Hu

机构 * School of Engineering and Applied Sciences, Harvard University(哈佛大学工程与应用科学学院) Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(哈佛大学自然与人工智能研究学院) Department of Psychology, Harvard University(哈佛大学心理学系)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

Comments 17 pages, 10 figures; updated with publishing information

Journal ref Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02255 2025-07-04 cs.IR cs.LG 83%

Listwise Preference Alignment Optimization for Tail Item Recommendation

Zihao Li, Chao Yang, Tong Zhang, Yakun Chen, Xianzhi Wang, Guandong Xu, Daoyi Dong

机构 * Australian Artificial Intelligence Institute (AAII) and School of Compute Sicence, Faculty of Engineering and Information Technology, University of Technology Sydney(澳大利亚人工智能研究所(AAII)和计算机科学学院,工程与信息技术学院,新南威尔士大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01368 2025-07-03 cs.CV cs.LG 83%

Activation Reward Models for Few-Shot Model Alignment

Tianning Chai, Chancharik Mitra, Brandon Huang, Gautam Rajendrakumar Gare, Zhiqiu Lin, Assaf Arbelle, Leonid Karlinsky, Rogerio Feris, Trevor Darrell, Deva Ramanan, Roei Herzig

机构 * University of California, Berkeley(加州大学伯克利分校) Carnegie Mellon University(卡内基梅隆大学) IBM Research(IBM研究院) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12999 2025-06-17 cs.CL 83%

POROver: Improving Safety and Reducing Overrefusal in Large Language Models with Overgeneration and Preference Optimization

Batuhan K. Karaman, Ishmam Zabir, Alon Benhaim, Vishrav Chaudhary, Mert R. Sabuncu, Xia Song

机构 * Cornell University(康奈尔大学) Microsoft(微软公司) Meta

专题命中 偏好对齐 :safety(title,abstract);alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09329 2025-06-12 cs.CL 83%

Towards Efficient and Effective Alignment of Large Language Models

Yuxin Jiang

机构 * Division of Emerging Interdisciplinary Areas(新兴跨学科领域 division)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12882 2025-06-09 cs.CR cs.CL cs.SE 83%

ProSec: Fortifying Code LLMs with Proactive Security Alignment

Xiangzhe Xu, Zian Su, Jinyao Guo, Kaiyuan Zhang, Zhenting Wang, Xiangyu Zhang

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL

Comments The first two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01704 2025-06-06 cs.CL 83%

An Exploration of Self-Supervised Mutual Information Alignment for Multi-Task Settings

Soham V. Govande

机构 * Stanford University(斯坦福大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04063 2025-06-05 cs.HC cs.DC cs.LG 83%

Crowd-SFT: Crowdsourcing for LLM Alignment

Alex Sotiropoulos, Sulyab Thottungal Valapu, Linus Lei, Jared Coleman, Bhaskar Krishnamachari

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20417 2025-05-28 cs.AI 83%

SCAR: Shapley Credit Assignment for More Efficient RLHF

Meng Cao, Shuyuan Zhang, Xiao-Wen Chang, Doina Precup

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19428 2025-05-27 cs.CL 83%

Frictional Agent Alignment Framework: Slow Down and Don't Break Things

Abhijnan Nath, Carine Graff, Andrei Bachinin, Nikhil Krishnaswamy

机构 * Department of Computer Science, Colorado State University(计算机科学系,科罗拉多州立大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

Comments 48 pages (main paper: 10 pages incl. Limitations and Acknowledgments; references: 6 pages; appendix: 32 pages), 9 figures, 12 tables, appearing in Proceedings of ACL 2025, Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18531 2025-05-27 cs.AI cs.CV 83%

Generative RLHF-V: Learning Principles from Multi-modal Human Preference

Jiayi Zhou, Jiaming Ji, Boyuan Chen, Jiapeng Sun, Wenqi Chen, Donghai Hong, Sirui Han, Yike Guo, Yaodong Yang

机构 * Peking University(北京大学) Hong Kong University of Science and Technology(香港科技大学) University College London(伦敦大学学院)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.AI

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04240 2025-05-27 cs.CL 83%

DiffPO: Diffusion-styled Preference Optimization for Efficient Inference-Time Alignment of Large Language Models

Ruizhe Chen, Wenhao Chai, Zhifei Yang, Xiaotian Zhang, Joey Tianyi Zhou, Tony Quek, Soujanya Poria, Zuozhu Liu

机构 * Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence(浙江医学影像人工智能重点实验室) Zhejiang University(浙江大学) SUTD(新加坡科技设计大学) Princeton University(普林斯顿大学) Peking University(北京大学) Nanyang Technological University(南洋理工大学) A*STAR Centre for Frontier AI Research(A*STAR前沿人工智能研究中心)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15610 2025-05-27 cs.LG 83%

On The Global Convergence Of Online RLHF With Neural Parametrization

Mudit Gaur, Amrit Singh Bedi, Raghu Pasupathy, Vaneet Aggarwal

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments The updated version of this paper is arXiv:2503.17644

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16714 2025-04-22 cs.CL 83%

Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment

Mingzhi Wang, Chengdong Ma, Qizhi Chen, Linjian Meng, Yang Han, Jiancong Xiao, Zhaowei Zhang, Jing Huo, Weijie J. Su, Yaodong Yang

机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) China Telecom(中国电信) University of Pennsylvania(宾夕法尼亚大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11337 2025-04-16 cs.CL 83%

REWARD CONSISTENCY: Improving Multi-Objective Alignment from a Data-Centric Perspective

Zhihao Xu, Yongqi Tong, Xin Zhang, Jun Zhou, Xiting Wang

专题命中 偏好对齐 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04950 2025-04-08 cs.LG 83%

A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Wenyuan Xu, Xiaochen Zuo, Chao Xin, Yu Yue, Lin Yan, Yonghui Wu

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments 11oages,2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02708 2025-04-04 cs.CL 83%

The Hidden Space of Safety: Understanding Preference-Tuned LLMs in Multilingual context

Nikhil Verma, Manasa Bharadwaj

专题命中 偏好对齐 :safety(title,abstract);alignment(abstract);分类 cs.CL

Comments 14 pages, 11 Figures, 2 Tables, currently under review at ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17696 2025-03-28 cs.CL 83%

Understanding the Logic of Direct Preference Alignment through Logic

Kyle Richardson, Vivek Srikumar, Ashish Sabharwal

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18454 2025-03-25 cs.CV cs.LG 83%

InPO: Inversion Preference Optimization with Reparametrized DDIM for Efficient Diffusion Model Alignment

Yunhong Lu, Qichao Wang, Hengyuan Cao, Xierui Wang, Xiaoyin Xu, Min Zhang

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.LG

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11779 2025-03-11 cs.CL 83%

Personality Alignment of Large Language Models

Minjun Zhu, Yixuan Weng, Linyi Yang, Yue Zhang

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

Comments Acecpt in ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02451 2025-03-11 cs.AI 83%

Strong Preferences Affect the Robustness of Preference Models and Value Alignment

Ziwei Xu, Mohan Kankanhalli

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.AI

Comments 21 Pages. Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05079 2025-03-10 cs.LG 83%

On a Connection Between Imitation Learning and RLHF

Teng Xiao, Yige Yuan, Mingxiao Li, Zhengyu Chen, Vasant G Honavar

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09724 2025-03-04 cs.CL 83%

Taming Overconfidence in LLMs: Reward Calibration in RLHF

Jixuan Leng, Chengsong Huang, Banghua Zhu, Jiaxin Huang

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10914 2025-02-21 cs.CL 83%

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment

Sizhe Wang, Yongqi Tong, Hengyuan Zhang, Dawei Li, Xin Zhang, Tianlong Chen

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

Comments The 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL 2025)- Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12735 2025-02-10 cs.LG 83%

Online Preference Alignment for Language Models via Count-based Exploration

Chenjia Bai, Yang Zhang, Shuang Qiu, Qiaosheng Zhang, Kang Xu, Xuelong Li

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.LG

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15453 2025-01-29 cs.CL 83%

Data-adaptive Safety Rules for Training Reward Models

Xiaomin Li, Mingye Gao, Zhiwei Zhang, Jingxuan Fan, Weiyu Li

专题命中 偏好对齐 :safety(title,abstract);RLHF(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12895 2025-01-23 cs.CL 83%

Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback

Yafu Li, Xuyang Hu, Xiaoye Qu, Linjie Li, Yu Cheng

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL

Comments 43 pages; work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15623 2024-12-23 cs.CR cs.AI 83%

JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs

Hongyi Li, Jiawei Ye, Jie Wu, Tianjie Yan, Chu Wang, Zhixin Li

专题命中 偏好对齐 :jailbreak(title,abstract);alignment(abstract);分类 cs.AI

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏