Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate
Mage:多轴评估LLM生成的可执行游戏场景超越编译通过率
机构 * Chalmers University of Technology and University of Gothenburg(楚尔姆斯理工大学和哥德堡大学)
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.AI、cs.LG
AI总结 本文提出Mage多轴评估框架,通过编译通过率、运行通过率、结构保真度和机制符合度四个维度评估LLM生成的游戏场景,发现编译通过率与功能正确性反相关,证明多轴评估对检测领域差异的必要性。
Comments Main Content: 10 pages, 1 figure. In total 22 pages