Prompting Policies for Multi-step Reasoning and Tool-Use in Black-box LLMs with Iterative Distillation of Experience
多步推理和工具使用中黑盒LLM的提示策略:基于经验迭代蒸馏的强化学习框架
机构 * Google Research(谷歌研究)
专题命中 效率与部署 :LLM(title_cn,summary_cn);prompting(title,abstract);large language model(abstract);language model(abstract)
AI总结 本文提出基于经验迭代蒸馏的强化学习框架,用于训练提示策略以提升黑盒LLM的多步推理和工具使用能力,实验显示在逻辑推理和工具使用任务中性能显著提升。
Comments 10 pages and reference, appendix