PRECISE: Reducing the Bias of LLM Evaluations Using Prediction-Powered Ranking Estimation
PRECISE: 使用预测驱动的排名估计减少LLM评估的偏差
机构 * Primary contributor and corresponding author(主要贡献者及通讯作者)
专题命中 预训练与数据 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 提出PRECISE框架,通过结合少量人工标注与LLM判断,利用预测驱动推断(PPI)方法,在低资源下可靠估计搜索、排序和RAG系统的指标,并校正LLM偏差。
Comments Accepted at AAAI 2026 - Innovative Applications of AI (IAAI-26)