Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning
推理与工具使用在智能体强化学习中的竞争:从量化干扰到解耦调优
机构 * School of Information, Renmin University of China(中国人民大学信息学院) ; Bytedance Inc.(字节跳动公司)
专题命中 工具调用 :tool-use(title,abstract);agentic(title,abstract);agent(abstract);tool use(abstract)
AI总结 本文通过引入能力效应归因(CEA)量化推理与工具使用行为之间的干扰,并提出解耦动作-推理调优(DART)框架,通过分离参数更新来提升智能体强化学习的性能。