Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning
评估评估者:面向大语言模型推理的验证自主等级(L0-L5)
专题命中 测试时计算 :reasoning(title);verifier(abstract);分类 cs.CL
AI总结 该研究提出验证自主等级(VAL)元标准,解决验证领域的等级概念混淆,明确不同验证方案的完备性边界,相关代码与评估材料已公开。
Comments Code and data: this https URL (https://github.com/1549080929-debug/math_agent) Keywords: LLM verification; verification autonomy; completeness; ground truth; trustworthy AI Writing and implementation assisted by an AI language model; all experiments, data, and research decisions are the author's own