机构
*
School of Biomedical Engineering, Tsinghua University(清华大学生物医学工程学院)
;
School of Biomedical Engineering, Shanghai Jiao Tong University(上海交通大学生物医学工程学院)
;
DAMO Academy, Alibaba Group(阿里云达摩院)
;
Hupan Laboratory(壶辰实验室)
;
Department of Biomedical Engineering, National University of Singapore(新加坡国立大学生物医学工程系)
;
Department of Radiology, Guizhou Provincial People’s Hospital(贵州省级人民医院放射科)
;
Department of Radiology, The First Affiliated Hospital, Zhejiang University School of Medicine(浙江大学医学院附属第一医院放射科)
;
Department of Radiology, Shanghai Sixth People’s Hospital Affiliated to Shanghai Jiao Tong University School of Medicine(上海交通大学医学院附属第六人民医院放射科)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
机构
*
Chinese University of Hong Kong(香港中文大学)
;
Westlake University(西湖大学)
;
Southern Medical University(南方医科大学)
;
Jiangnan University(江南大学)
;
Dalian University of Technology(大连理工大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
University of Copenhagen(哥本哈根大学)
机构
*
The Hong Kong Polytechnic University(香港理工大学)
;
Sun Yat-sen University(中山大学)
;
InfiX.ai
;
PolyU-Daya Bay Technology and Innovation Research Institute(PolyU-大亚湾技术与创新研究院)
Med-R2: Perception and Reflection-driven Complex Reasoning for Medical Report Generation
Med-R2:面向医学报告生成的感知与反思驱动复杂推理
Hao Wang, Shuchang Ye, Jinghao Lin, Usman Naseem, Jinman Kim
机构
*
The School of Computer Science, The University of Sydney(悉尼大学计算机科学学院)
;
The School of Computing, Macquarie University(麦考瑞大学计算机学院)
;
Doubao Medical Group, ByteDance(字节跳动 doubao 医疗集团)
6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models
6根手指,1个肾脏:自然对抗性医学图像揭示视觉语言模型的关键弱点
Leon Mayer, Piotr Kalinowski, Caroline Ebersbach, Marcel Knopp, Tim Rädsch, Evangelia Christodoulou, Annika Reinke, Fiona R. Kolbinger, Lena Maier-Hein
机构
*
German Cancer Research Center (DKFZ) Heidelberg, Division of Intelligent Medical Systems(德国癌症研究中心(DKFZ)海德堡,智能医学系统部门)
;
Medical Faculty, Heidelberg University(海德堡大学医学院)
;
Faculty of Mathematics and Computer Science, Heidelberg University(海德堡大学数学与计算机科学学院)
;
HIDSS4Health - Helmholtz Information and Data Science School for Health, Karlsruhe/Heidelberg(HIDSS4Health - 哈勃-马克斯信息与数据科学健康学院,卡尔斯鲁厄/海德堡)
;
Helmholtz Imaging, German Cancer Research Center (DKFZ)(哈勃-马克斯成像,德国癌症研究中心(DKFZ))
;
Engineering Faculty, Heidelberg University(海德堡大学工程学院)
;
School of Computation, Information and Technology, TUM(技术大学(TUM)计算、信息与技术学院)
;
Weldon School of Biomedical Engineering, Purdue University(普渡大学韦尔登生物医学工程学院)
;
Department of Visceral, Thoracic and Vascular Surgery, University Hospital and Faculty of Medicine Carl Gustav Carus, TUD Dresden University of Technology(visceral、胸腔和血管外科部门,技术大学(TUD)德累斯顿大学医院和医学院)
;
National Center for Tumor Diseases (NCT), NCT Heidelberg, a partnership between DKFZ and University Hospital Heidelberg(肿瘤疾病国家中心(NCT),海德堡NCT,DKFZ与海德堡大学医院之间的合作)
;
Heidelberg University Hospital, Surgical Clinic, Surgical AI Research Group(海德堡大学医院,外科诊所,外科人工智能研究组)
;
Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, UAE(Mohamed Bin Zayed人工智能大学(MBZUAI),阿布扎赫,阿拉伯联合酋长国)
RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model
RetiBridge:用知识引导的多模态大语言模型连接定量视网膜生物标志物与定性诊断
Zhuangzhi Gao, Hongyi Qin, He Zhao, Qinkai Yu, Feixiang Zhou, Fu Wang, Jinru Ding, Eduard Shantsila, Uazman Alam, Alena Shantsila, Wahbi El-Bouri, Gregory Y. H. Lip, Yalin Zheng
机构
*
University of Liverpool(利物浦大学)
;
Institute of Life Course & Medical Sciences(生命课程与医学科学研究院)
;
Department of Eye and Vision Sciences(眼科与视觉科学系)
;
Computer Science Department(计算机科学系)
;
Cardiovascular & Metabolic Medicine(心血管与代谢医学)
;
Liverpool Centre for Cardiovascular Science(利物浦心血管科学中心)
Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models
对前沿视觉-语言模型进行审计以实现可信的医学视觉问答:定位失败、格式崩溃和领域适应
Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li
机构
*
New York University, New York, USA(纽约大学)
;
Tsinghua University, Beijing, China(清华大学)
;
Columbia University, New York, USA(哥伦比亚大学)
;
University of Michigan, Ann Arbor, USA(密歇根大学)
机构
*
National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University(国家多媒体软件工程技术研究中心,武汉大学计算机学院)
;
College of Computing and Data Science, Nanyang Technological University(computing and Data Science学院,南洋理工大学)
机构
*
Nepal Applied Mathematics and Informatics Institute for Research(尼泊尔应用数学与信息技术研究所)
;
GastroIntestinal Department, Dhulikhel Hospital(杜尔基hel医院消化内科)
;
Univesity of Lausanne(洛桑大学)
;
University of West Virginia(西弗吉尼亚大学)
;
University of Utah(犹他大学)
;
University of Aberdeen(阿伯丁大学)
Enhancing Pathological VLMs with Cross-scale Reasoning
增强病理视觉语言模型的跨尺度推理能力
Chi Phan, Tianyi Zhang, Qiaochu Xue, Yufeng Wu, Dan Hu, Zeyu Liu, Sudong Wang, Yueming Jin
机构
*
Department of Electrical and Computer Engineering, National University of Singapore(新加坡国立大学电气与计算机工程系)
;
PuzzleLogic Pte Ltd(PuzzleLogic私人有限公司)
;
Department of Pathology, Fujian Medical University Cancer Hospital & Fujian Cancer Hospital(福建医科大学附属肿瘤医院病理科暨福建省肿瘤医院)
EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models
EasyLens: 一种无需训练的即插即用型微病变表示放大器,用于医学视觉语言模型
Qiwei Zeng, Hao Wang, Jinghao Lin, Shuchang Ye, Yuezhe Yang, Yige Peng, Haoyuan Che, Jinman Kim, Lei Bi
机构
*
Jilin University(吉林大学)
;
School of Computer Science, The University of Sydney(悉尼大学计算机科学学院)
;
ByteDance(字节跳动)
;
Institute of Translational Medicine, Shanghai Jiao Tong University(上海交通大学转化医学研究院)
机构
*
MoE Key Lab of Artificial Intelligence(人工智能MOE实验室)
;
AI Institute(人工智能研究院)
;
School of Computer Science(计算机科学学院)
;
Shanghai Jiao Tong University(上海交通大学)
;
Department of Radiology(放射科)
;
The First Affiliated Hospital(第一附属医院)
;
School of Medicine(医学院)
;
Zhejiang University(浙江大学)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
State Key Laboratory of Infrared Physics(红外物理国家重点实验室)
;
Shanghai Institute of Technical Physics(上海技术物理研究所)
;
Chinese Academy of Science(中国科学院)
机构
*
Department of Computer Science, Southern Methodist University(计算机科学系,南方 Methodist 大学)
;
Department of Ophthalmology, UT Southwestern Medical Center(眼科学系,UT 南方医学中心)
;
High Performance Computing Program, National Aeronautics and Space Administration(高性能计算计划,国家航空航天局)
机构
*
School of Computing and Informatics, University of Louisiana at Lafayette(路易斯安那州立大学拉法叶分校计算机与信息学学院)
;
Louisiana Center for Health Innovation and College of Nursing & Health Sciences, University of Louisiana at Lafayette(路易斯安那州立大学拉法叶分校健康创新中心及护理与健康科学学院)
;
Massachusetts Eye and Ear, Harvard Medical School(哈佛医学院马萨诸塞眼耳医院)
;
Department of Computer Science, University of Central Florida(佛罗里达州立大学计算机科学系)
;
Department of Electrical Engineering and Computer Science, Florida Atlantic University(佛罗里达Atlantic大学电子工程与计算机科学系)
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
ClinFusion:用于整体医学理解的以视觉为中心的多模态大语言模型系统
Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
机构
*
DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)
;
Hupan Laboratory(湖畔实验室)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Department of Radiology, The Affiliated Yangming Hospital of Ningbo University(宁波大学附属阳明医院放射科)
;
Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University(浙江大学伊利诺伊大学厄巴纳香槟校区联合学院,浙江大学)
;
Hepato-Pancreato-Biliary Center, Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua Medicine, Tsinghua University(清华长庚医院肝胆胰中心,清华大学临床医学院,清华医学,清华大学)
;
School of Software, Tsinghua University(清华大学软件学院)
;
Beijing National Research Center for Information Science and Technology, Tsinghua University(清华大学北京信息科学与技术国家研究中心)
MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
MEDIC-AD:迈向医疗视觉-语言模型的临床智能
Woohyeon Park, Jaeik Kim, Sunghwan Steve Cho, Pa Hong, Wookyoung Jeong, Yoojin Nam, Namjoon Kim, Ginny Y. Wong, Ka Chun Cheung, Jaeyoung Do
机构
*
AIDAS Laboratory, Seoul National University(首尔大学AIDAS实验室)
;
Samsung Changwon Hospital(三星昌原医院)
;
Samsung Medical Center(三星医疗中心)
;
NVIDIA, Santa Clara, USA(英伟达(美国圣克拉拉))
机构
*
School of iOPEN, Northwestern Polytechnical University(iOPEN学院,西北工业大学)
;
The First Affiliated Hospital of Sun Yat-Sen University(中山大学附属第一医院)
;
College of Mechanical Engineering, Tongji University(同济大学机械工程学院)
Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey
视觉与视觉-语言应用中的模态感知特征匹配:全面综述
Weide Liu, Wei Zhou, Jun Liu, Ping Hu, Jun Cheng, Jungong Han, Weisi Lin
机构
*
School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics(江西财经大学计算机与人工智能学院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院)
;
School of Computer Science and Informatics, Cardiff University(卡迪夫大学计算机科学与信息学院)
;
School of Computing and Communications, Lancaster University(兰卡斯特大学计算机与通讯学院)
;
School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)
;
Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR)(新加坡资讯研究院,科技研究局(A*STAR))
;
Department of Automation, Tsinghua University(清华大学自动化系)