Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
迈向统一的手术场景理解:通过多模态大语言模型弥合推理与 grounding
机构 * Southern University of Science and Technology(南方科技大学) ; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) ; Northwestern University(西北大学) ; University of Alberta(阿尔伯塔大学) ; Yale University(耶鲁大学) ; Nanfang Hospital(南华医院) ; Shenzhen University of Advanced Technology(深圳大学先进技术研究院)
专题命中 视觉定位与Grounding :grounding(title,title_cn);MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 本文提出SurgMLLM框架,通过统一推理与视觉 grounding 实现手术场景理解,提升三元组识别和分割精度,实验表明其在三元组识别指标AP_IVT上提升显著。