From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space
从P(y|x)到P(y):在预训练空间中研究强化学习
机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; National University of Singapore(新加坡国立大学) ; Tencent AI Lab(腾讯AI实验室)
AI总结 本文提出PreRL和DSRL方法,通过优化预训练空间中的边际分布P(y),提升LLM推理能力,并通过NSR机制增强推理效果,实验表明DSRL在推理任务中表现优异。
Comments Preprint. Our code is available at https://github.com/Trae1ounG/Pretrain_Space_RLVR