Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study
推理努力,而非工具访问,决定了智能体代码生成的首试可靠性:一项观察性研究
机构 * TrendAI
专题命中 代码生成 :code generation(title);分类 cs.SE、cs.AI
AI总结 本研究通过90次独立智能体运行构建同一应用,发现推理努力(从高到极高)将首试完美运行率从28%提升至89%,而测试工具虽增加成本却未改善功能得分或可靠性。
Comments 22 pages, 5 figures, 10 tables. Dataset and evaluation artifacts: https://doi.org/10.5281/zenodo.21134406