HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models
HalluWorld: 一个用于通过参考世界模型控制幻觉的基准
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Patronus AI ; Independent Researcher(独立研究者) ; Stanford University(斯坦福大学) ; The Ohio State University(俄亥俄州立大学) ; DegenAI Labs(DegenAI实验室)
专题命中 Agent评测 :agentic(abstract);分类 cs.AI、cs.CL、cs.LG
AI总结 本文提出HalluWorld基准,通过显式参考世界模型研究语言模型的幻觉问题,发现不同任务中幻觉表现不一致,表明幻觉源于多种失败模式而非单一能力。
Comments HalluWorld benchmark (code and data) at github.com/DegenAI-Labs/HalluWorld