机构
*
University of Maryland(马里兰大学)
;
Brown University(布朗大学)
;
Washington University in St. Louis(圣路易斯华盛顿大学)
;
Adobe
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Southern California(南加州大学)
;
NVIDIA
专题命中
后训练与偏好优化
:language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.LG
Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
逐步引导策略优化:在GRPO中着色你的错误推理
Peter Chen, Xiaopeng Li, Ziniu Li, Xi Chen, Tianyi Lin
机构
*
Department of Industrial Engineering and Operations Research(工业工程与运营管理系)
;
Department of Mathematics(数学系)
;
Columbia University(哥伦比亚大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Stern School of Business, New York University(纽约大学斯特恩商学院)
专题命中
后训练与偏好优化
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
机构
*
School of Computer Science, Peking University(北京大学计算机科学系)
;
BioGeometry
;
Mila - Québec AI Institute(魁北克人工智能研究所)
;
Université de Montréal(蒙特利尔大学)
;
HEC Montréal(蒙特利尔HEC商学院)
;
CIFAR AI Research Chair(CIFAR人工智能研究主席)
ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning
ADHint: 基于难度先验的自适应提示用于强化学习
Feng Zhang, Zezhong Tan, Xinhong Ma, Ziqiang Dong, Xi Leng, Jianfei Zhao, Xin Sun, Yang Yang
机构
*
Alibaba Group(阿里巴巴集团)
;
Beijing Institute of Technology(北京理工大学)
;
Peking University(北京大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Zhongguancun Academy(中关村学院)
Journal refIn Proceedings of the 3rd International Workshop on AI for Quantum and Quantum for AI (AIQxQIA 2025), co-located with ECAI 2025, CEUR Workshop Proceedings, Vol. 4153, Paper 6, pp. 1-12, 2025