Actor-Curator: Co-adaptive Curriculum Learning via Policy-Improvement Bandits for RL Post-Training
Actor-Curator: 通过策略改进带状机为RL后训练实现联合适应课程学习
机构 * University of Illinois Chicago(伊利诺伊大学香槟分校) ; Caltech(加州理工学院) ; RPI(罗切斯特理工学院) ; MBZUAI(澳门大学人工智能研究院) ; University of Chicago(芝加哥大学) ; NEC Laboratories America(NEC美国实验室)
专题命中 后训练与偏好优化 :post-training(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)
AI总结 ACTOR-CURATOR通过策略改进带状机实现RL后训练的联合适应课程学习,有效提升训练稳定性和效率。
Comments 37 pages, 8 figures, 1 table. Preprint under review. Equal contribution by first two authors