MedMT-Bench: Can LLMs Memorize and Understand Long Multi-Turn Conversations in Medical Scenarios?
MedMT-Bench: LLMs能否在医学场景中记忆并理解长多轮对话?
机构 * ByteDance(字节跳动)
专题命中 医学数据与评测 :medical AI(abstract);diagnosis(abstract)
AI总结 MedMT-Bench通过模拟诊断治疗流程,测试LLM在长上下文记忆、干扰鲁棒性和安全防御方面的表现,发现17种前沿模型均无法达到60%以上的准确率。