HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs
HiPO:用于大语言模型自适应推理的分层偏好优化
机构 * Vellore Institute of Technology(维洛雷理工学院) ; University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校) ; Northwestern University(西北大学) ; Yale University(耶鲁大学) ; Algoverse AI Research(Algoverse AI研究)
AI总结 HiPO通过分层优化提升大语言模型在复杂推理任务中的表现,结合偏好优化与结构化推理的优势,实现更高效的训练和更一致的输出。
Comments 12 pages, 4 figures, 6 tables. Includes ablation study across Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct on 5 math reasoning benchmarks (GSM8K, MATH500, Minerva, AIME24, Gaokao2023). GPT-4.1 used for structured evaluation of reasoning quality