机构
*
School of Computer Science, Peking University(北京大学计算机科学系)
;
Tongyi Lab, Alibaba Group(阿里集团通义实验室)
;
Department of Computing Science, University of Alberta(阿尔伯塔大学计算机科学系)
;
Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
数据混合代理:学习重新加权领域以实现持续预训练
Kailai Yang, Xiao Liu, Lei Ji, Hao Li, Xiao Liang, Zhiwei Liu, Yeyun Gong, Peng Cheng, Mao Yang
机构
*
The University of Manchester(曼彻斯特大学)
;
Microsoft Research(微软研究院)
;
Imperial College London(伦敦帝国学院)
;
University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
Principled RL for Diffusion LLMs Emerges from a Sequence-Level Perspective
从序列层面视角出发的扩散大语言模型原理化强化学习
Jingyang Ou, Jiaqi Han, Minkai Xu, Shaoxuan Xu, Jianwen Xie, Stefano Ermon, Yi Wu, Chongxuan Li
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院)
;
Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大模型与智能治理研究重点实验室)
;
Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程技术研究中心)
;
Stanford University(斯坦福大学)
;
Lambda, Inc(Lambda公司)
;
Tsinghua University(清华大学)
More Than One Teacher: Adaptive Multi-Guidance Policy Optimization for Diverse Exploration
Xiaoyang Yuan, Yujuan Ding, Yi Bin, Wenqi Shao, Jinyu Cai, Jingkuan Song, Yang Yang, Heng Tao Shen
机构
*
Tongji University(同济大学)
;
Hong Kong Polytechnic University(香港理工大学)
;
Shanghai AI Lab(上海人工智能实验室)
;
National University of Singapore(新加坡国立大学)
;
University of Electronic Science and Technology of China(电子科技大学)
Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration From Cognitive Psychology
Wei Xie, Shuoyoucheng Ma, Zhenhua Wang, Enze Wang, Kai Chen, Xiaobing Sun, Baosheng Wang
机构
*
College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机科学与技术学院)
;
Institute of High Performance Computing, Agency for Science, Technology and Research (A*STAR)(科学、技术和研究局高性能计算研究所)
;
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
CommentsThank you for your attention. This paper was accepted by the CogSci 2025 conference in April and published in August. The location in the proceedings is: https://escholarship.org/uc/item/24x9t7s1
CommentsPreliminary version, v3, added the missing name of x-axis in the left part of Fig.1 and corrected a wrong number in Fig.3. Project page: https://anitaleungxx.github.io/ReMix
RLPR: Extrapolating RLVR to General Domains without Verifiers
Tianyu Yu, Bo Ji, Shouli Wang, Shu Yao, Zefan Wang, Ganqu Cui, Lifan Yuan, Ning Ding, Yuan Yao, Zhiyuan Liu, Maosong Sun, Tat-Seng Chua
机构
*
Tsinghua University(清华大学)
;
National University of Singapore(新加坡国立大学)
;
Shanghai Qi Zhi Institute(上海启智研究院)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)