When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models
当经过验证的世界模型仍然失败时:大语言模型合成代码世界模型中的游戏充分性与预测准确性
机构 * AGILabs(AGILabs实验室)
专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)
AI总结 研究大语言模型合成的代码世界模型,发现其虽预测准确率高,但在游戏中仍失败,存在验证与正确的差距,揭示危害规律,指出更多数据无法修复,相同机制在信念推理函数上重现,表明规划世界模型充分性应基于搜索分布或游戏衡量。
Comments 41 pages, 4 figures. Code and reproduction log: https://github.com/JaviMaligno/code-world-models