Prompt Embedding Probes (PEP): Hallucination Detection in LLMs from Hidden States
提示嵌入探针(PEP):基于隐藏状态的大语言模型幻觉检测
专题命中 知识编辑与模型理解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
AI总结 本研究提出提示嵌入探针(PEP),一种白盒幻觉检测方法,通过少量可学习提示嵌入扩充标准线性探针,在TriviaQA等数据集的Qwen3模型上验证其在生成前、跨模型等场景的有效性,仅需少量参数即可提升检测性能。
Comments 10 pages, 7 figures. Code available at https://github.com/zazamrykh/internal_probing