Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis
将2D多模态大语言模型适应于3DCT图像分析
机构 * Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(香港科技大学电子及计算机工程学系) ; Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程学系) ; Institute of Medical Intelligence and XR, The Chinese University of Hong Kong(香港中文大学医学智能与扩展现实研究所) ; Center for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(中国科学院香港创新研究院人工智能与机器人创新中心)
专题命中 多模态生成 :multi-modal(title);multimodal(abstract);MLLM(abstract);cross-modal(abstract)
AI总结 本文提出将2D多模态大语言模型适应于3DCT图像分析,设计Text-Guided Hierarchical MoE框架,采用两阶段训练策略提升医学报告生成和视觉问答任务性能。