Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
通过隐式奖励模型实现LLM生成文本的零样本检测
机构 * School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学)
专题命中 后训练与偏好优化 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
AI总结 本文提出IRM方法,利用隐式奖励模型实现LLM生成文本的零样本检测,无需偏好收集或额外训练,在DetectRL基准上表现优于现有方法。
Comments NeurIPS 2025