A Descriptive and Normative Theory of Human Beliefs in RLHF
基于人类反馈强化学习中的人类信念描述性与规范性理论
Sylee Dandekar, Shripad Deshmukh, Frank Chiu, W. Bradley Knox, Scott Niekum
机构
*
College of Information and Computer Sciences University of Massachusetts Amherst(信息与计算机科学学院 马萨诸塞大学阿姆赫斯特分校)
;
Department of Computer Science The University of Texas at Austin(计算机科学系 德州大学奥斯汀分校)
Staleness-Learning Rate Scaling Laws for Asynchronous RLHF
异步RLHF的陈旧度-学习率缩放定律
Jingwei Song, Haofeng Xu, Jie Xiao, Chengke Bao, Jingwei Shi, Pengbin Feng, Weixun Wang, Yuhang Han, Chuan Wu, Linfeng Zhang, Bill Shi
机构
*
The University of Hong Kong(香港大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Gradient
;
University of Southern California(南加州大学)
;
The Hong Kong Polytechnic University(香港理工大学)
机构
*
University of Arizona, USA(亚利桑那大学)
;
Arizona State University, USA(亚利桑那州立大学)
;
Now at Google LLC, work done at Rice University(现就职于谷歌公司,曾就职于里士大学)
;
Clemson University, USA(克莱姆森大学)
;
Washington University in St. Louis, USA(圣路易斯华盛顿大学)
;
Halmstad University, Sweden(哈姆斯塔德大学)
;
Guangdong Institute of Intelligence Science and Technology, China(广东智能科学与技术研究院)
专题命中
后训练与偏好优化
:preference optimization(title,abstract);LLM(abstract_cn);large language model(abstract);language model(abstract)
机构
*
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(人工智能与数字经济广东实验室)
;
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院)
;
Baidu Inc.(百度公司)
专题命中
后训练与偏好优化
:LLM(title,abstract);large language model(abstract);language model(abstract);preference optimization(abstract)
Beyond Pairwise: Empowering LLM Alignment With Ranked Choice Modeling
超越成对:通过排名选择建模增强大语言模型对齐
Yuxuan Tang, Yifan Feng
机构
*
Institute of Operations Research and Analytics(运营研究与分析研究所)
;
National University of Singapore(新加坡国立大学)
;
Department of Analytics and Operations(分析与运营系)
;
NUS Business School(新加坡国立大学商学院)
专题命中
后训练与偏好优化
:LLM(title,abstract);large language model(abstract);language model(abstract);preference optimization(abstract)
Joint Continual Learning of Local Language Models and Cloud Offloading Decisions with Budget Constraints
带有预算约束的本地语言模型与云卸载决策的联合持续学习
Evan Chen, Wenzhi Fang, Shiqiang Wang, Christopher Brinton
机构
*
Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN(电子与计算机工程学院,普渡大学)
;
Department of Computer Science, University of Exeter, UK(计算机科学系,埃克塞特大学)
专题命中
后训练与偏好优化
:language model(title,abstract);large language model(abstract);small language model(abstract);post-training(abstract)