Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering
可重复、可解释和有效的代理AI在软件工程中的评估
机构 * Norwegian University of Science and Technology(挪威科技大学)
专题命中 Agent评测 :agentic(title,abstract);autonomous agent(abstract);分类 cs.AI、cs.SE
AI总结 本文分析了18篇关于代理AI在软件工程中评估的论文,提出指南以提升评估的可重复性、可解释性和有效性,通过公开TAR轨迹和LLM交互数据促进不同方法的系统比较。
Comments 7 pages, 5 figures, accepted to the 2nd International Workshop on Responsible Software Engineering (ResponsibleSE 2026), co-located with FSE