Doubly Robust Alignment for Large Language Models
机构 * Department of Statistics(统计系) ; LSE London, UK(伦敦大学学院) ; Department of Mathematics(数学系) ; Tsinghua University(清华大学) ; School of Design LCC, UAL London, UK(伦敦艺术大学设计学院) ; Department of Engineering Science(工程科学系) ; University of Oxford(牛津大学)
专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);RLHF(abstract);preference optimization(abstract)
Comments Accepted to NeurIPS 2025