Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
检测和纠正商业大语言模型和深度研究代理中的参考幻觉
机构 * University of Pennsylvania(宾夕法尼亚大学)
AI总结 研究通过评估10种模型和代理在DRBench和ExpertQA上的引用URL可靠性,发现3-13%的URL存在幻觉,5-18%无法解析。提出urlhealth工具可有效减少非解析URL,验证引用URL的有效性与可纠正性。
Comments 25 pages