Evaluating the Architectural Reasoning Capabilities of LLM Provers via the Obfuscated Natural Number Game
通过混淆的自然数游戏评估大语言模型证明者的架构推理能力
机构 * Lixing Li(李立星)
AI总结 本文通过混淆的自然数游戏评估大语言模型的架构推理能力,发现推理模型在无语义线索下仍保持准确率,为数学推理能力提供了量化评估标准。
Comments 4 pages. Accepted as a short paper to the AAAI 2026 Spring Symposium on Machine Learning and Knowledge Engineering for Knowledge-Grounded Semantic Agents (MAKE 2026)