Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
通过贡献加权群体相对策略优化增强基于LLM的搜索代理
机构 * Fudan University(复旦大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Shanghai Key Laboratory of Intelligent Information Processing(上海智能信息处理关键实验室)
AI总结 本文提出CW-GRPO框架,通过整合过程监督与群体相对策略优化,提升搜索代理的信用分配效率,实验表明其在多个知识密集型基准上表现优异。
Comments Accepted to the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), Main Conference