SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding
SurgNarrator:面向手术视频理解的生成式检索框架
机构 * University of Liverpool(利物浦大学) ; City University of Hong Kong(香港城市大学) ; University of Oxford(牛津大学) ; University of Strasbourg(斯特拉斯堡大学) ; IHU Strasbourg(斯特拉斯堡大学医院研究所) ; Imperial College London(帝国理工学院)
专题命中 视频理解 :video understanding(title,abstract);video-language(abstract);分类 cs.CV
AI总结 SurgNarrator是专为手术视频理解设计的生成式检索框架,通过构建手术词汇表、适配Qwen3-VL-Embedding-8B并采用分层检索策略,在零样本设置下于12个基准上实现性能提升且延迟大幅降低。
Comments This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible