CausalGame: Benchmarking Causal Thinking of LLM Agents in Games
因果游戏:在游戏中对大语言模型智能体的因果思维进行基准测试
机构 * MBZUAI(穆罕默德·本·扎耶德人工智能大学) ; Carnegie Mellon University(卡内基梅隆大学) ; Hong Kong Baptist University(香港浸会大学) ; University of Oxford(牛津大学) ; New York University, Abu Dhabi(纽约大学阿布扎比分校)
AI总结 研究旨在通过因果游戏基准测试大语言模型智能体的因果思维能力,让智能体设计实验、收集数据并给出解决方案。设计含多种挑战的场景,发现当前智能体因果思维能力欠佳,为评估提供了测试平台。
Comments Zhenhao, Yongqiang, and Chenxi contributed equally to the project. A short version is accepted at the Forty-Third International Conference on Machine Learning (ICML) 2026 as an Oral presentation. Project website https://causalgame.github.io/