Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization
策略分裂:通过双模式熵正则化激励大语言模型强化学习中的双模式探索
机构 * Beijing Institute of Technology(北京理工大学) ; Tsinghua University(清华大学) ; Beihang University(北航)
AI总结 提出Policy Split方法,将策略分裂为正常和高熵两种模式,通过协作双模式熵正则化在保持准确性的同时促进多样化探索,实验表明在通用和创造性任务上优于现有基线。
Comments preprint