From Classification to Cross-Modal Understanding: Leveraging Vision-Language Models for Fine-Grained Renal Pathology
Zhenhao Guo, Rachit Saluja, Tianyuan Yao, Quan Liu, Junchao Zhu, Haibo Wang, Daniel Reisenbüchler, Yuankai Huo, Benjamin Liechty, David J. Pisapia, Kenji Ikemura, Steven Salvatoree, Surya Seshane, Mert R. Sabuncu, Yihe Yang, Ruining Deng
机构
*
New York University(纽约大学)
;
Cornell Tech(康奈尔科技)
;
Vanderbilt University(范德比大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
University of Regensburg(莱茵河畔大学)
;
Weill Cornell Medicine(韦尔·科恩医学中心)
;
Northwell Health(北well健康)
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
University of California, Santa Barbara(加州大学圣巴巴拉分校)
;
Peking Union Medical College Hospital(北京佑安医学大学医院)
RadVLM: A Multitask Conversational Vision-Language Model for Radiology
Nicolas Deperrois, Hidetoshi Matsuo, Samuel Ruipérez-Campillo, Moritz Vandenhirtz, Sonia Laguna, Alain Ryser, Koji Fujimoto, Mizuho Nishio, Thomas M. Sutter, Julia E. Vogt, Jonas Kluckert, Thomas Frauenfelder, Christian Blüthgen, Farhad Nooralahzadeh, Michael Krauthammer
机构
*
Department of Radiology, Kobe University(金泽大学放射科)
;
Department of Computer Science, ETH Zurich(苏黎世联邦理工学院计算机科学系)
;
Department of Advanced Imaging in Medical Magnetic Resonance, Kyoto University(京都大学医学磁共振高级成像部门)
;
Department of Quantitative Biomedicine, University of Zurich(苏黎世大学定量生物医学系)
;
Diagnostic and Interventional Radiology, University Hospital Zurich(苏黎世大学医院诊断与介入放射科)
No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
Min Woo Sun, Alejandro Lozano, Javier Gamazo Tejero, Vishwesh Nath, Xiao Xiao Sun, James Burgess, Yuhui Zhang, Kun Yuan, Robert Tibshirani, Sean Huver, Serena Yeung-Levy
DentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice
Zijie Meng, Jin Hao, Xiwei Dai, Yang Feng, Jiaxiang Liu, Bin Feng, Huikai Wu, Xiaotang Gai, Hengchuan Zhu, Tianxiang Hu, Yangyang Wu, Hongxia Xu, Jin Li, Jun Xiao, Xiaoqiang Liu, Joey Tianyi Zhou, Fudong Zhu, Zhihe Zhao, Lunguo Xia, Bing Fang, Jimeng Sun, Jian Wu, Zuozhu Liu
机构
*
Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine, Zhejiang University, Hangzhou(牙科医院,口腔医学院,浙江大学医学院,浙江大学,杭州)
;
College of Computer Science and Technology, Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University, Hangzhou(计算机科学与技术学院,浙江大学-伊利诺伊大学 Urbana-Champaign 院,浙江大学,杭州)
;
Department of Orthodontics, Shanghai Ninth People’s Hospital, College of Stomatology, Shanghai Jiao Tong University School of Medicine(正畸科,上海第九人民医院,口腔医学院,上海交通大学医学院)
;
Angelalign Technology Inc.(Angelalign 技术公司)
机构
*
East China Normal University(华东师范大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Northwest A&F University(西北农林科技大学)
;
Tencent Youtu Lab(腾讯优图实验室)
;
Xiamen University(厦门大学)
Simulating Clinical AI Assistance using Multimodal LLMs: A Case Study in Diabetic Retinopathy
Nadim Barakat, William Lotter
机构
*
Dana-Farber Cancer Institute & Tufts University School of Medicine(达纳-法伯癌症研究所及塔夫茨大学医学院)
;
Dana-Farber Cancer Institute Brigham and Women’s Hospital & Harvard Medical School(达纳-法伯癌症研究所布里特妇女医院及哈佛医学院)
On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI
David Restrepo, Ira Ktena, Maria Vakalopoulou, Stergios Christodoulidis, Enzo Ferrante
机构
*
MICS, CentraleSupélec - Université Paris-Saclay, France(MICS,中央圣艾尔布兰大学-巴黎萨克雷大学,法国)
;
Google DeepMind, London, UK(谷歌DeepMind,伦敦,英国)
;
CONICET, Universidad de Buenos Aires, Argentina(CONICET,布宜诺斯艾利斯大学,阿根廷)
On the Importance of Text Preprocessing for Multimodal Representation Learning and Pathology Report Generation
Ruben T. Lucassen, Tijn van de Luijtgaarden, Sander P. J. Moonemans, Gerben E. Breimer, Willeke A. M. Blokx, Mitko Veta
机构
*
Dept. of Pathology, University Medical Center Utrecht(病理学系,乌得勒支大学医学中心)
;
Dept. of Biomedical Engineering, Eindhoven University of Technology(生物医学工程系,埃因霍温理工大学)
;
Dept. of Mathematics and Computer Science, Eindhoven University of Technology(数学与计算机科学系,埃因霍温理工大学)
MAIRA-1: A specialised large multimodal model for radiology report generation
Stephanie L. Hyland, Shruthi Bannur, Kenza Bouzid, Daniel C. Castro, Mercy Ranjit, Anton Schwaighofer, Fernando Pérez-García, Valentina Salvatelli, Shaury Srivastav, Anja Thieme, Noel Codella, Matthew P. Lungren, Maria Teodora Wetscherek, Ozan Oktay, Javier Alvarez-Valle
专题命中
医疗多模态
:radiology(title,abstract);分类 cs.CV
Comments18 pages, 9 tables, 5 figures. v2 adds test IDs and image encoder citation. v3 fixes error in NPV/specificity
Learning to Exploit Temporal Structure for Biomedical Vision-Language Processing
Shruthi Bannur, Stephanie Hyland, Qianchu Liu, Fernando Pérez-García, Maximilian Ilse, Daniel C. Castro, Benedikt Boecking, Harshita Sharma, Kenza Bouzid, Anja Thieme, Anton Schwaighofer, Maria Wetscherek, Matthew P. Lungren, Aditya Nori, Javier Alvarez-Valle, Ozan Oktay
A Foundational Multimodal Vision Language AI Assistant for Human Pathology
Ming Y. Lu, Bowen Chen, Drew F. K. Williamson, Richard J. Chen, Kenji Ikamura, Georg Gerber, Ivy Liang, Long Phi Le, Tong Ding, Anil V Parwani, Faisal Mahmood
BioFact-MoE: Biologically Factorized Mixture of Experts for Vision-Language Prognostic Modeling in Hepatocellular Carcinoma
BioFact-MoE:基于生物学因子分解的混合专家模型用于肝细胞癌的视觉-语言预后建模
Junlin Yang, Tian Yu, Nicha C. Dvornek, Yuexi Du, Peiyu Duan, Annabella Shewarega, Lawrence H. Staib, James S. Duncan, Julius Chapiro
机构
*
Department of Radiology \& Biomedical Imaging, Department of Biomedical Engineering, Department of Electrical Engineering, Department of Statistics \& Data Science Yale University, New Haven, CT, 06510, USA
MeD-3D: A Multimodal Deep Learning Framework for Precise Recurrence Prediction in Clear Cell Renal Cell Carcinoma (ccRCC)
Hasaan Maqsood, Saif Ur Rehman Khan
机构
*
Skolkovo Institute of Science and Technology (Skoltech)(斯克罗夫诺科学与技术学院(Skoltech))
;
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI))