Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities
数学文本延续的似然评分:一个自监督基准及快捷漏洞测试
机构 * Department of Physics, California Institute of Technology(加州理工学院物理系)
AI总结 本文提出一个自动基准,用于预测技术论文中的隐藏文本。通过比较模型生成的辅助预测字符串与评分器对延续的预测,评估信息传递效果。实验显示,GPT-5.5等模型在方程后缀预测任务中优于基线,支持似然评分作为静态基准和快捷漏洞测试的工具。
Comments 13 pages + appendices, 4 figures; v2: expanded related work