From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM
从通用到专用:基于冻结视觉语言模型(VLM)的内窥镜息肉报告上下文融合框架
机构 * Zhejiang University(浙江大学) ; Shanghai Institute for Advanced Study, Zhejiang University(浙江大学上海高等研究院) ; Shanghai Key Laboratory of MICCAI(上海市MICCAI重点实验室) ; Digital Medical Research Center, School of Basic Medical Sciences, Fudan University(复旦大学基础医学院数字医学研究中心) ; Endoscopy Center and Endoscopy Research Institute, Zhongshan Hospital, Fudan University(复旦大学附属中山医院内镜中心及内镜研究所) ; Shanghai Collaborative Innovation Center of Endoscopy(上海市内镜协同创新中心) ; Alliance Manchester Business School, The University of Manchester(曼彻斯特大学联盟曼彻斯特商学院)
专题命中 效率与部署 :language model(abstract);分类 cs.AI
AI总结 本研究针对内窥镜息肉报告需求,提出一种不修改预训练权重的上下文融合框架,通过隐式指令与显式转换上下文对冻结通用VLM专用化,在2056张标注图像实验中实现最优性能且参数量仅为冻结VLM的0.006%。