机构
*
University of Copenhagen(哥本哈根大学)
;
IIIT Ranchi(印度信息技术学院兰契分校)
;
ISI Kolkata(印度统计学院加尔各答分校)
;
NIT Andhra Pradesh(印度国家理工学院安得拉邦分校)
;
IGDTUW(英迪拉·甘地德里技术女子大学)
;
IIT Kharagpur(印度理工学院卡哈拉格普尔分校)
;
Google DeepMind(谷歌DeepMind)
;
Google(谷歌)
;
AI Institute, University of South Carolina(南卡罗来纳大学人工智能研究所)
Aligning Large Language Models with Searcher Preferences
将大型语言模型对齐于搜索者偏好
Wei Wu, Peilun Zhou, Liyi Chen, Qimeng Wang, Chengqiang Lu, Yan Gao, Yi Wu, Yao Hu, Hui Xiong
机构
*
School of Artificial Intelligence and Data Science, University of Science and Technology of China(中国科学技术大学人工智能与数据科学学院)
;
Xiaohongshu Inc.(小红书公司)
;
Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能研究所)
;
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
评分标准作为攻击面:LLM裁判中的隐蔽偏好漂移
Ruomeng Ding, Yifei Pang, He Sun, Yizhong Wang, Zhiwei Steven Wu, Zhun Deng
机构
*
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Yale University(耶鲁大学)
;
The University of Texas at Austin(德克萨斯大学奥斯汀分校)