The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World
重放差距:大语言模型智能体中模型切换的静态评估评估的是错误的世界
机构 * Carnegie Mellon University(卡内基梅隆大学)
AI总结 该研究发现大语言模型智能体的模型切换静态评估基于错误假设,通过分支展开实验证明模型交换会导致轨迹分歧,重放评估无法准确预测结果,发布了相关工具与轨迹。
Comments 8 pages, 3 figures. Accepted at the Conference on Language Modeling 2026. Code: https://github.com/AshrithaG/replay-gap Data: https://huggingface.co/datasets/ashritha0907/replay-gap-trajectories