MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding
MLLM-HWSI: 一种用于分层全滑动图像理解的多模态大语言模型
机构 * Department of Computer Science, Khalifa University of Science and Technology(卡利法科技大学计算机科学系) ; Information Technology University(信息技术大学) ; KAU(卡乌大学) ; University of the Western Australia(西澳大学)
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);MLLM(title,abstract);grounding(abstract);分类 cs.CV
AI总结 本文提出MLLM-HWSI,一种分层全滑动图像级多模态大语言模型,通过四级尺度对齐视觉特征与病理语言,提升解释性证据接地推理能力,在六个CPath任务上取得新SOTA结果。