Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology
面向放射学的空间定位2D视觉-语言模型的可扩展训练
机构 * Computer Vision Group, University of Freiburg, Germany(德国弗莱堡大学计算机视觉组) ; Department of Radiology, Medical Center -- University of Freiburg, Germany(德国弗莱堡大学医学中心放射科) ; CRIION-AI Lab, Freiburg, Germany(德国弗莱堡CRIION-AI实验室)
专题命中 医疗多模态 :radiology(title,abstract);CT(abstract,abstract_cn);分类 cs.CV、cs.LG
AI总结 提出RefRad2D大规模双语数据集,通过LLM和自动分割生成空间定位数据,训练RadGrounder模型联合完成报告生成、VQA和空间定位,在外部基准上取得竞争性结果。
Comments Accepted for MICCAI 2026. First two authors: equal contribution. Last two authors: equal supervision