Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
Nemotron-Cascade:基于级联强化学习的通用推理模型规模化
机构 * NVIDIA(英伟达)
专题命中 测试时计算 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出级联域强化学习(Cascade RL)方法,开发出Nemotron-Cascade模型,可在指令和深度思考模式间切换,性能不逊于纯思考模型,同时提升推理能力并保持跨领域基准性能。
Comments We publicly release the Nemotron-Cascade models and the full collection of training data at: https://huggingface.co/collections/nvidia/nemotron-cascade