ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement
ThinkTwice:联合优化大型语言模型以进行推理和自我完善
机构 * Department of Computer Science, University of Toronto(多伦多大学计算机科学系) ; Coolwei AI Lab(Coolwei AI实验室)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.AI
AI总结 ThinkTwice通过两阶段框架联合优化LLM,提升推理和自我完善性能,基于GRPO方法,在多个数学推理基准上优于现有基线。
Comments 27 pages,7 figures,5 tables