Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization
Active-GRPO:用于分子优化的自适应模仿与自我改进推理
机构 * School of Medicine, Stanford University(斯坦福大学医学院) ; Data Science Institute, University of Chicago(芝加哥大学数据科学研究所) ; Pritzker School of Molecular Engineering, University of Chicago(芝加哥大学普利兹克分子工程学院) ; Department of Computer Science, University of Chicago(芝加哥大学计算机科学系) ; Argonne National Laboratory(阿贡国家实验室)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI、cs.LG
AI总结 提出Active-GRPO方法,通过主动模仿-强化和主动参考机制,在分子优化中自适应平衡模仿与自我改进,显著提升性能。