SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
SlopCodeBench:评估编码代理在长周期迭代任务中性能退化的基准测试
机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) ; Washington State University(华盛顿州立大学) ; MIT(麻省理工学院)
专题命中 软件智能体 :coding agent(title,abstract);分类 cs.SE、cs.CL、cs.AI
AI总结 本文提出SlopCodeBench,通过36个问题和196个检查点评估编码代理在长周期迭代任务中的性能退化,发现代理代码在结构上逐渐退化并产生冗余代码,人类代码退化更慢。
Comments Code and Leaderboards are located at https://www.scbench.ai