When benchmark inferences do not compose: Projectibility in AI evaluation
当基准推理无法组合:AI评估中的可投射性
机构 * Humber Polytechnic(汉伯理工学院) ; University of Toronto(多伦多大学)
AI总结 本文针对AI评估中基准推理无法组合的问题,提出非组合原则,结合古德曼的竞争延伸问题与基于论证的有效性框架,通过案例和模拟开发可投射性审计以诊断基准到应用论证的衔接缺陷。
Comments 34 pages, 2 figures, 5 tables. v2 substantially revises Secs. 5-8 and the conclusion, adds a measured instance of factor-structure instability, and corrects a claim in Sec. 3.3 that endpoint alignment suffices for composition. Supersedes the withdrawn arXiv:2510.15236. Code: https://github.com/BrettRey/benchmark-inference-composition