Can LLMs Write Correct TLA+ Specifications? Evaluating Natural-Language-to-TLA+ Generation
LLM 能写出正确的 TLA+ 规范吗?自然语言到 TLA+ 生成的评估
机构 * Department of Computer Science, Loyola University Chicago(洛约奈大学芝加哥分校计算机科学系)
专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG
AI总结 本文首次系统评估基于 LLM 从自然语言合成 TLA+ 规范的能力,发现模型在语义正确性上仅达 8.6%,且成功依赖于渐进式提示,揭示了模型大小与质量无关、代码专用模型表现不佳等关键发现。
Comments 12 pages, 11 tables. Accepted at the 21st International Conference on Software Technologies (ICSOFT 2026); Recommended as Best Paper Award Candidate
Journal ref ICSOFT 2026, pp. 39-50, 2026