机构
*
Department of Computer Science(计算机科学系)
;
Stony Brook University(石英布鲁克大学)
;
Department of Applied Mathematics and Statistics(应用数学与统计学系)
;
Yale University(耶鲁大学)
;
Department of Data Science(数据科学系)
;
New Jersey Institute of Technology(新泽西理工学院)
;
Department of Biomedical Informatics(生物医学信息学系)
CommentsRevised version following peer review. Expanded methodological details, practitioner-survey validation, statistical analyses, and discussion; main conclusions unchanged
机构
*
Nanjing University(南京大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias
Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standard
基于精确贝叶斯最优标准测量语言模型的上下文内算法推理能力
Hector Zenil, Luan Ozelim
机构
*
Oxford Immune Algorithmics(牛津免疫算法学公司)
;
Oxford University Innovation(牛津大学创新公司)
;
London Institute for Healthcare Engineering(伦敦医疗工程研究所)
;
King’s College London(伦敦国王学院)
CommentsThis paper has been withdrawn by the authors due to a significant bug discovered in our data processing pipeline. This bug affects the validity of the experimental results, and we can no longer stand by the conclusions presented