MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models
MLREF:基于大语言模型的强化学习奖励设计中高效模块复用框架
机构 * Institute for Interdisciplinary Information Sciences Tsinghua University(清华大学交叉信息研究院)
专题命中 仓库级理解 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 该研究针对强化学习奖励设计瓶颈,提出MLREF框架,通过模块池复用奖励组件,结合三种优化机制,在17个任务上实现比强基线更优且更稳定的性能。
Comments 22 pages, 5 figures, 4 tables