Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
开放式多智能体协作在语言智能体中的基准测试
机构 * University of Edinburgh(爱丁堡大学) ; University of Oxford(牛津大学) ; University College London(伦敦大学学院)
专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);autonomous agent(abstract);分类 cs.AI、cs.LG
AI总结 提出基于JAX的开放式多智能体协作基准Alem,评估13种现代LLM在长时生存世界中的零样本协作能力,发现协调能力是前沿LLM智能体的独立瓶颈。
Comments 42 pages, preprint