Targeted Exploration via Unified Entropy Control for Reinforcement Learning
通过统一熵控制实现定向探索的强化学习
机构 * College of Software, Nankai University(南开大学软件学院) ; Zhongguancun Academy(中关村学院) ; Shanghai Jiao Tong University(上海交通大学) ; Chinese Academy of Sciences(中国科学院) ; Tsinghua University(清华大学) ; Sun Yat-sen Univeristy(中山大学)
AI总结 本文提出UEC-RL框架,通过定向探索机制和稳定器解决强化学习中的熵崩溃问题,提升大模型推理性能。
Comments Accepted for publication in Findings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)