LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers
LLM-as-a-Reviewer: 基准测试它们作为论文审稿人的能力、分歧和提示注入抵抗性
机构 * University of South Florida(佛罗里达南大学) ; Missouri University of Science and Technology(密苏里科技大学) ; University of Alabama(阿拉巴马大学) ; Florida International University(佛罗里达国际大学) ; University of Cincinnati(辛辛那提大学) ; George Mason University(乔治·梅森大学)
专题命中 提示注入 :prompt injection(title,abstract);alignment(abstract);分类 cs.CL、cs.CY
AI总结 本研究通过一个系统基准测试,评估了12个大型语言模型在论文评审中的表现,包括评分校准、与人类审稿人的分歧以及对不可见字体映射攻击的抵抗性,发现LLMs存在系统性高估弱论文、与人类关注点不同以及易受提示注入攻击等问题。