Are LLMs Vulnerable to Preference-Undermining Attacks (PUA)? A Factorial Analysis Methodology for Diagnosing the Trade-off between Preference Alignment and Real-World Validity
大型语言模型是否易受偏好削弱攻击(PUA)攻击?一种诊断偏好对齐与现实有效性之间权衡的因子分析方法
Hongjun An, Yiliang Song, Jiangan Chen, Jiawei Shao, Chi Zhang, Xuelong Li
机构
*
School of Artificial Intelligence, OPtics and ElectroNics, Northwestern Polytechnical University(人工智能学院、光学与电子学院、西北工业大学)
;
Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信)
;
School of Economics and Management, Guangxi Normal University(经济管理学院,广西师范大学)
机构
*
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
;
School of Computer Science and Electronic Engineering, University of Surrey(Surrey大学计算机科学与电子工程学院)
机构
*
Department of Electrical Engineering, Pennsylvania State University(宾夕法尼亚州立大学电气工程系)
;
Department of Computer Science, Rutgers University(罗格斯大学计算机科学系)
;
College of Information Sciences and Technology, Pennsylvania State University(宾夕法尼亚州立大学信息科学与技术学院)
机构
*
Cisco Research(思科研究)
;
Indian Institute of Technology Kharagpur(印度理工学院卡里格普尔分校)
;
Eindhoven University of Technology, Netherlands (TU/e)(埃因霍温理工大学(荷兰))
机构
*
Shenzhen International Graduate School, Tsinghua University, Shenzhen(清华大学深圳国际研究生院)
;
Tencent AI Lab(腾讯AI实验室)
;
Wuhan University(武汉大学)
;
The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen)(香港中文大学(深圳))
;
X-LANCE Lab, Shanghai Jiao Tong University, Shanghai(上海交通大学X-LANCE实验室)
;
School of Intelligence Science and Technology, Nanjing University, Suzhou, China(南京大学智能科学与技术学院)
KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity
Gholamali Aminian, Amir R. Asadi, Idan Shenfeld, Youssef Mroueh
机构
*
The Alan Turing Institute(艾伦·图灵研究所)
;
Statistical Laboratory(统计实验室)
;
University of Cambridge(剑桥大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
IBM Research USA(IBM美国研究)
机构
*
Key Laboratory of Big Data & Artificial Intelligence in Transportation, Beijing Jiaotong University, Ministry of Education(大数据与人工智能交通运输 key laboratory,北京交通大学,教育部)
;
School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China(计算机科学与技术学院,北京交通大学,北京,中国)
;
Pattern Recognition Center, WeChat AI, Tencent Inc, China(模式识别中心,微信AI,腾讯公司,中国)
On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization
Jiancong Xiao, Ziniu Li, Xingyu Xie, Emily Getzen, Cong Fang, Qi Long, Weijie J. Su
机构
*
University of Pennsylvania(宾夕法尼亚大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
National University of Singapore(新加坡国立大学)
;
Peking University(北京大学)
;
Joint corresponding authors(联合通讯作者)