Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings
大语言模型工具中的部分相关验证级联:凹对数优势、多项式可靠性和盲点上限
机构 * Independent researcher(独立研究者)
专题命中 测试时计算 :verifier(title,abstract);分类 cs.AI、cs.LG
AI总结 研究大语言模型工具中部分相关验证级联,将生成器错误接受率建模为潜在变量,给出理论。包括凹对数优势、多项式可靠性等特性,可测量,通过合成测试对比不同方法效果,指出实际应通过去相关而非加门提升可靠性。
Comments 14 pages, 2 figures. Code and synthetic-recovery experiments: https://github.com/jianganghan/harness-verifier-cascades