RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications
RuBench:一个具有原生编写的俄语任务规范的仓库级智能编码基准测试
机构 * Independent Researcher(独立研究者)
专题命中 代码评测 :repository(title);coding agent(abstract,comments);分类 cs.SE、cs.CL、cs.AI
AI总结 RuBench 1.0是含25个俄语任务的仓库级智能编码基准测试,任务源于开源仓库修复提交,按客户请求风格编写。评估多种产品配置,报告运行结果,发现最佳配置解决78.7%任务,还发现产品替换模型问题,为智能编码评估提供新基准。
Comments 20 pages, 1 figure, 9 tables. v2 adds Round 2: Russian-market coding agents (SourceCraft CLI, Koda CLI), Antigravity with Gemini 3.1 Pro / 3.5 Flash, and Codex CLI with GPT-5.6 on the same frozen task set, plus a tool-call contamination re-audit (network + disk layers). Data, full trajectories and harness: https://github.com/eugeneshilow/rubench