On Cost-Effective LLM-as-a-Judge Improvement Techniques
关于成本效益的LLM作为评判者的改进技术
专题命中 评测与基准 :LLM(title,title_cn);RLHF(abstract,abstract_cn);language model(abstract);prompting(abstract)
AI总结 研究通过集成评分、任务特定标准注入等四种技术提高LLM评判准确性,在RewardBench 2上达到85.8%准确率,成本效益显著。
Comments Accepted at the ICML 2026 workshops "Statistical Frameworks for Uncertainty in Agentic Systems" and "Combining Theory and Benchmarks: Towards a Virtuous Cycle to Understand and Guarantee Foundation Model Performance". 13 pages, 9 figures