A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
针对特定语料库的临床检索增强生成(RAG)系统在HealthBench上的表现与较新的前沿大语言模型相当或更优
专题命中 RAG评测 :RAG(title,title_cn);retrieval-augmented generation(abstract);knowledge retrieval(abstract);分类 cs.IR、cs.CL、cs.AI
AI总结 本文评估专为中低收入环境构建的临床RAG系统VITA,其在HealthBench基准测试中与前沿LLM表现相当或更优,语料库特异性可提升模型落地性但会降低沟通流畅度。
Comments 2 tables