Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play
评估大型语言模型作为实时战略代理:提供商性能、混合分解及时间风险游戏中的操作差距
专题命中 规划决策 :agent(abstract);planning(abstract);分类 cs.AI
AI总结 本文研究了大型语言模型在实时策略环境中的表现,发现其性能受目标跟踪、执行转换、成本和运行时可靠性等因素影响,支持将LLM作为受限制工作流中的组件进行评估,而非孤立的基准测试对象。
Comments 13 pages, 7 figures. Code and tracked notes: https://github.com/hcekne/risk-game . Public runtime artifact index: https://github.com/hcekne/risk-game/blob/main/docs/article-plans/public_experiment_artifacts.md