Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring
对多模态大语言模型评分者的审计:临床顺序评分中的中间倾向偏差
Jiaqing Zhang, Sandeep Elluri, Bhanu Cherukuvada, Yonah Joffe, Jessica Sena, Miguel Contreras, Scott Siegel, Subhash Nerella, Catherine Price, Parisa Rashidi
机构
*
Department of Electrical & Computer Engineering(电气与计算机工程系)
;
Department of Computer and Information Science and Engineering(计算机与信息科学与工程系)
;
Department of Clinical and Health Psychology(临床与健康心理学系)
;
Department of Biomedical Engineering(生物医学工程系)
专题命中
评测与基准
:LLM(title,summary_cn);large language model(abstract);language model(abstract)
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
在没有基准的情况下:在无标签的情况下验证比较LLM安全性评分
Sushant Gautam, Finn Schwall, Annika Willoch Olstad, Fernando Vallecillos Ruiz, Birk Torpmann-Hagen, Sunniva Maria Stordal Bjørklund, Leon Moonen, Klas Pettersen, Michael A. Riegler
机构
*
Simula Metropolitan Center for Digital Engineering(Simula 数字工程中心)
;
Oslo Metropolitan University(奥斯陆 Metropolitan 大学)
;
University of Oslo(奥斯陆大学)
;
Simula Research Laboratory(Simula 研究实验室)
;
Norwegian Directorate of Health(挪威健康 Directorate)
GazeMind: A Gaze-Guided LLM Agent for Personalized Cognitive Load Assessment
GazeMind:一种基于注视的LLM代理用于个性化认知负荷评估
Bin Wang, Yue Liu, Benjamin Newman, Ajoy S. Fernandes, Zhiyuan Wang, Robert Cavin, Michele A. Cox, Vijay Rajanna, Takumi Bolte, Melissa Hunfalvay, Ulas Bagci, Michael J. Proulx
Comments14 pages of main content, 3 figures, 4 tables, 9 appendices. This paper has been submitted to the Becker Friedman Institute 2026 AI in Social Sciences conference for peer review
CommentsRevised version: Corrected citation errors in the Learning Check subsection (Section RQ2), updated the RQ2 summary to more accurately reflect the mixed results in the data, and added clarifying notes on implementation structure as a moderating variable for learning outcomes. 16 pages, 8 tables, 2 figures