VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training
VRPO: 重新思考价值建模以实现大语言模型后训练中噪声监督下的鲁棒强化学习
机构 * College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) ; Honor Device Co., Ltd(荣耀设备有限公司) ; Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究院) ; Shanghai Key Laboratory of Multimodal Embodied AI(上海多模态具身人工智能重点实验室)
专题命中 数学推理 :reasoning(abstract);math reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 针对强化学习中噪声监督导致策略不稳定问题,提出VRPO框架,通过熵和困惑度辅助损失及变分信息瓶颈增强价值模型,使其能过滤噪声并生成可靠优势估计,在多项任务上优于PPO和GRPO。