arXivDaily arXiv每日学术速递 周一至周五更新

科学与医疗

医学 AI

医学智能、临床 AI、医学影像、病理、诊断和医疗健康大模型。

共收录 966 信号源:cs.CV, cs.LG, q-bio, eess.IV, eess.SP

1. 医疗多模态 966 篇

2601.22696 2026-02-02 cs.CV cs.LG 62%

Bi-MCQ: Reformulating Vision-Language Alignment for Negation Understanding

Bi-MCQ:重新表述视觉-语言对齐以理解否定

Tae Hun Kim, Hyun Gyu Lee

机构 * Department of Electrical and Computer Engineering, Inha University, Republic of Korea(电气与计算机工程系,印哈大学,大韩民国) College of Medicine, Inha University, Republic of Korea(医学学院,印哈大学,大韩民国)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV、cs.LG

AI总结 Bi-MCQ通过重新表述视觉-语言对齐为条件语义比较,提升医学VLM对否定理解的性能。

Comments 15 pages, 4 figures, Submitted to ICPR 2026 (under review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02443 2026-01-07 cs.CV cs.AI eess.IV 62%

Evaluating the Diagnostic Classification Ability of Multimodal Large Language Models: Insights from the Osteoarthritis Initiative

评估多模态大语言模型的诊断分类能力:来自骨关节炎倡议的见解

Li Wang, Xi Chen, XiangWen Deng, HuaHui Yi, ZeKun Jiang, Kang Li, Jian Li

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV、eess.IV

AI总结 研究发现,多模态大语言模型在医学图像分类中表现不佳,建议优先优化视觉编码器和数据集质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20100 2026-01-07 cs.LG cs.AI cs.CL cs.CV 62%

MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations

MIRAGE:农业专家引导对话中多模态信息检索与推理的基准

Vardhan Dongre, Chi Gui, Shubham Garg, Hooshang Nayyeri, Gokhan Tur, Dilek Hakkani-Tür, Vikram S. Adve

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV、cs.LG

AI总结 MIRAGE是一个用于农业专家引导对话中多模态信息检索与推理的基准,通过真实用户-专家交互数据,提供高保真的多模态推理评估平台。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23185 2025-12-30 eess.IV cs.AI cs.CV 62%

EIR: Enhanced Image Representations for Medical Report Generation

EIR: 增强的医学报告生成图像表示

Qiang Sun, Zongcheng Ji, Yinlong Xiao, Peng Chang, Jun Yu

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV、eess.IV

AI总结 EIR通过跨模态Transformer融合元数据与图像表示,结合医学领域预训练模型,有效解决信息不对称和领域差距问题,提升胸部X光报告生成的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02677 2025-12-17 eess.IV cs.CV 62%

Multimodal Deep Learning for Stroke Prediction and Detection using Retinal Imaging and Clinical Data

基于视网膜成像和临床数据的多模态深度学习用于中风预测与检测

Saeed Shurrab, Aadim Nepal, Terrence J. Lee-St. John, Nicola G. Ghazi, Bartlomiej Piechowski-Jozwiak, Farah E. Shamout

机构 * Division of Engineering, New York University Abu Dhabi(纽约大学阿布扎赫尔分校工程系) Institute for Healthier Living Abu Dhabi(阿布扎赫尔健康生活研究所) Eye Institute at Cleveland Clinic Abu Dhabi(阿布扎赫尔克利夫兰医学中心眼科研究所) Canberra Hospital(堪培拉医院)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV、eess.IV

AI总结 本研究提出一种多模态深度学习方法,利用视网膜成像和临床数据预测中风风险,实验结果显示在中风检测和风险预测方面优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27680 2025-12-02 cs.CV cs.AI cs.LG 62%

PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated Reporting

PETAR:基于掩码感知的视觉-语言建模的局部发现生成用于PET自动报告

Danyal Maqbool, Changhee Lee, Zachary Huemann, Samuel D. Church, Matthew E. Larson, Scott B. Perlman, Tomas A. Romero, Joshua D. Warner, Meghan Lubner, Xin Tie, Jameson Merkow, Junjie Hu, Steve Y. Cho, Tyler J. Bradshaw

机构 * University of Wisconsin–Madison Department of Computer Sciences(威斯康星大学麦迪逊分校计算机科学系) University of Wisconsin–Madison Department Radiology(威斯康星大学麦迪逊分校放射学系) Microsoft(微软公司)

专题命中 医疗多模态 :CT(abstract);分类 cs.CV、cs.LG

AI总结 PETAR通过引入PETARSeg-11K数据集和PETAR-4B模型,实现基于掩码感知的3D PET自动报告生成,提升医学影像分析的精度与实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17635 2025-11-25 cs.CV cs.LG 62%

Upstream Probabilistic Meta-Imputation for Multimodal Pediatric Pancreatitis Classification

上游概率元填补用于多模态儿童胰腺炎分类

Max A. Nelson, Elif Keles, Eminenur Sen Tasci, Merve Yazol, Halil Ertugrul Aktas, Ziliang Hong, Andrea Mia Bejar, Gorkem Durak, Oznur Leman Boyunaga, Ulas Bagci

专题命中 医疗多模态 :MRI(abstract);分类 cs.CV、cs.LG

AI总结 本文提出UPMI方法,通过元特征空间中的概率元填补提升多模态儿童胰腺炎分类性能,实现AUC提升5%。

Comments 5 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03328 2025-11-06 cs.CL cs.AI cs.CV cs.LG 62%

Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks

Jindong Hong, Tianjie Chen, Lingjie Luo, Chuanyang Zheng, Ting Xu, Haibao Yu, Jianing Qiu, Qianzhong Chen, Suning Huang, Yan Xu, Yong Gui, Yijun He, Jiankai Sun

机构 * Bytedance(字节跳动) Peking University(北京大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Mohamed bin Zayed University of Artificial Intelligence(马尔代夫人工智能大学) Stanford University(斯坦福大学) University of Michigan(密歇根大学)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22728 2025-10-28 cs.LG cs.CV 62%

S-Chain: Structured Visual Chain-of-Thought For Medicine

Khai Le-Duc, Duy M. H. Nguyen, Phuong T. H. Trinh, Tien-Phat Nguyen, Nghiem T. Diep, An Ngo, Tung Vu, Trinh Vuong, Anh-Tien Nguyen, Mau Nguyen, Van Trung Hoang, Khai-Nguyen Nguyen, Hy Nguyen, Chris Ngo, Anji Liu, Nhat Ho, Anne-Christin Hauschild, Khanh Xuan Nguyen, Thanh Nguyen-Tang, Pengtao Xie, Daniel Sonntag, James Zou, Mathias Niepert, Anh Totti Nguyen

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV、cs.LG

Comments First version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11666 2025-10-14 eess.SP cs.LG 62%

Explainable Deep Neural Network for Multimodal ECG Signals: Intermediate vs Late Fusion

Timothy Oladunni, Ehimen Aneni

专题命中 医疗多模态 :medical AI(abstract);分类 cs.LG、eess.SP

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09230 2025-10-13 cs.CV cs.AI cs.CL cs.LG 62%

Diagnosing Shoulder Disorders Using Multimodal Large Language Models and Consumer-Grade Cameras

Jindong Hong, Wencheng Zhang, Shiqin Qiao, Jianhai Chen, Jianing Qiu, Chuanyang Zheng, Qian Xu, Yun Ji, Qianyue Wen, Weiwei Sun, Hao Li, Huizhen Li, Huichao Wang, Kai Wu, Meng Li, Yijun He, Lingjie Luo, Jiankai Sun

机构 * Bytedance(字节跳动) Peking University(北京大学) Peking University People’s Hospital(北京大学人民医院) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01558 2025-10-03 cs.CE cs.LG eess.SP 62%

CardioRAG: A Retrieval-Augmented Generation Framework for Multimodal Chagas Disease Detection

Zhengyang Shen, Xuehao Zhai, Hua Tu, Mayue Shi

机构 * Department of Electrical and Electronic Engineering Imperial College London(帝国理工学院电子与电气工程系) Department of Civil and Environmental Engineering Imperial College London(帝国理工学院土木与环境工程系) Institute of Biomedical Engineering Department of Engineering Science University of Oxford(牛津大学生物医学工程研究所)

专题命中 医疗多模态 :medical AI(abstract);分类 cs.LG、eess.SP

Comments 4 pages, 2 figures. Accepted for oral presentation at the 52nd international Computing in Cardiology Conference (CinC2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20188 2025-08-29 cs.CV cs.LG 62%

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study

Max Torop, Masih Eskandar, Nicholas Kurtansky, Jinyang Liu, Jochen Weber, Octavia Camps, Veronica Rotemberg, Jennifer Dy, Kivanc Kose

机构 * Northeastern University(东北大学) Memorial Sloan Kettering Cancer Center(纪念斯隆凯特琳癌症中心)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19862 2025-08-28 cs.CV cs.LG 62%

Multimodal Conditional MeshGAN for Personalized Aneurysm Growth Prediction

Long Chen, Ashiv Patel, Mengyun Qiao, Mohammad Yousuf Salmasi, Salah A. Hammouche, Vasilis Stavrinides, Jasleen Nagi, Soodeh Kalaie, Xiao Yun Xu, Wenjia Bai, Declan P. O'Regan

机构 * MRC Laboratory of Medical Sciences Imperial College London(医学科学实验室 Imperial College London) Imperial College Healthcare NHS Trust(帝国理工医疗 NHS Trust) Department of Mechanical Engineering University College London(机械工程系 University College London) London Postgraduate School of Surgery NHS England(伦敦外科研究生学校 NHS England) Faculty of Medicine Imperial College London(医学系 Imperial College London) Department of Chemical Engineering Imperial College London(化学工程系 Imperial College London) Department of Brain Sciences&Computing Imperial College London(脑科学与计算系 Imperial College London)

专题命中 医疗多模态 :CT(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19319 2025-08-28 eess.IV cs.AI cs.CV 62%

MedVQA-TREE: A Multimodal Reasoning and Retrieval Framework for Sarcopenia Prediction

Pardis Moradbeiki, Nasser Ghadiri, Sayed Jalal Zahabi, Uffe Kock Wiil, Kristoffer Kittelmann Brockhattingen, Ali Ebrahimi

机构 * Department of Electrical and Computer Engineering, Isfahan University of Technology(电气与计算机工程系,伊斯法罕技术大学) SDU Health Informatics and Technology, The Maersk Mc-Kinney Moller Institute, University of Southern Denmark(南部丹麦大学健康信息学与技术,马士基麦金尼莫勒研究所) Geriatric Research Unit, Department of Clinical Research, University of Southern Denmark(老年医学研究单元,临床研究系,南部丹麦大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV、eess.IV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16882 2025-08-26 eess.IV cs.CV 62%

Multimodal Medical Endoscopic Image Analysis via Progressive Disentangle-aware Contrastive Learning

Junhao Wu, Yun Li, Junhao Li, Jingliang Bian, Xiaomao Fan, Wenbin Lei, Ruxin Wang

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) First Affiliated Hospital, Sun Yat-sen University(中山大学第一附属医院) College of Big Data and Internet, Shenzhen Technology University(深圳技术大学大数据与互联网学院)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV、eess.IV

Comments 12 pages,6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14706 2025-08-21 cs.CL cs.AI cs.CV cs.LG cs.MM 62%

ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine

Junying Chen, Zhenyang Cai, Zhiheng Liu, Yunjin Yang, Rongsheng Wang, Qingying Xiao, Xiangyi Feng, Zhan Su, Jing Guo, Xiang Wan, Guangjun Yu, Haizhou Li, Benyou Wang

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09717 2025-08-14 cs.CV cs.LG 62%

Multimodal Sheaf-based Network for Glioblastoma Molecular Subtype Prediction

Shekhnaz Idrissova, Islem Rekik

机构 * BASIRA Lab, Imperial-X(BASIRA实验室、Imperial-X) Department of Computing, Imperial College London, United Kingdom(计算系、帝国理工学院伦敦分校,英国)

专题命中 医疗多模态 :MRI(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02865 2025-08-13 eess.IV cs.AI cs.CL cs.CV 62%

VisionUnite: A Vision-Language Foundation Model for Ophthalmology Enhanced with Clinical Knowledge

Zihan Li, Diping Song, Zefeng Yang, Deming Wang, Fei Li, Xiulan Zhang, Paul E. Kinahan, Yu Qiao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Washington(华盛顿大学) Shenzhen Institutes of Advanced Technology(深圳先进技术研究所) Chinese Academy of Sciences(中国科学院) State Key Laboratory of Ophthalmology(眼科学国家重点实验室) Zhongshan Ophthalmic Center(中山眼科中心) Sun Yat-sen University(中山大学) Guangdong Provincial Key Laboratory of Ophthalmology and Visual Science(广东省眼科学与视觉科学重点实验室) Guangdong Provincial Clinical Research Center for Ocular Diseases(广东省眼科临床研究中心) Department of Bioengineering(生物工程系) Department of Radiology(放射科)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV、eess.IV

Comments Accepted by IEEE TPAMI, 14 pages, 15 tables, 4 figures with Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06701 2025-08-12 cs.CV cs.AI cs.CL cs.LG cs.SD eess.AS 62%

MMFformer: Multimodal Fusion Transformer Network for Depression Detection

Md Rezwanul Haque, Md. Milon Islam, S M Taslim Uddin Raju, Hamdi Altaheri, Lobna Nassar, Fakhri Karray

机构 * Centre for Pattern Analysis and Machine Intelligence, Department of Electrical and Computer Engineering, University of Waterloo(模式分析与机器智能中心,电气与计算机工程系,滑铁卢大学) School of Engineering and Computing, Department of Computer Science and Engineering, American University of Ras Al Khaimah(工程与计算学院,计算机科学与工程系,阿联酋拉线哈姆斯美国大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV、cs.LG

Comments Accepted for the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03734 2025-08-07 eess.IV cs.AI cs.CV 62%

A Survey of Multimodal Ophthalmic Diagnostics: From Task-Specific Approaches to Foundational Models

Xiaoling Luo, Ruli Zheng, Qiaojian Zheng, Zibo Du, Shuo Yang, Meidan Ding, Qihao Xu, Chengliang Liu, Linlin Shen

机构 * College of Computer Science and Software Engineering, Shenzhen University, Shenzhen, China(深圳大学计算机科学与软件工程学院) Shenzhen Key Laboratory of Visual Object Detection and Recognition, Harbin Institute of Technology, Shenzhen, 518055, China(视觉对象检测与识别深圳重点实验室) Laboratory for Artificial Intelligence in Design, Hong Kong(人工智能设计实验室) School of Artificial Intelligence, Shenzhen University, Shenzhen, China(深圳大学人工智能学院)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV、eess.IV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03008 2025-08-06 eess.IV cs.AI cs.CV 62%

ClinicalFMamba: Advancing Clinical Assessment using Mamba-based Multimodal Neuroimaging Fusion

Meng Zhou, Farzad Khalvati

机构 * TD Bank Group, Toronto, Canada(TD银行集团,加拿大) Department of Computer Science, University of Toronto, Toronto, Canada(多伦多大学计算机科学系,加拿大) Department of Medical Imaging, University of Toronto, Toronto, Canada(多伦多大学医学影像系,加拿大)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV、eess.IV

Comments Accepted at MICCAI MLMI 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23402 2025-08-01 cs.CV cs.AI cs.LG 62%

AGA: An adaptive group alignment framework for structured medical cross-modal representation learning

Wei Li, Xun Gong, Jiao Li, Xiaobin Sun

机构 * School of Computing and Artificial Intelligence(计算机与人工智能学院) Southwest Jiaotong University(西南交通大学) Department of Gastroenterology(消化内科部) The Third People’s Hospital of Chengdu(成都第三人民医院)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07001 2025-06-24 cs.CV cs.LG 62%

Hallucination-Aware Multimodal Benchmark for Gastrointestinal Image Analysis with Large Vision-Language Models

Bidur Khanal, Sandesh Pokhrel, Sanjay Bhandari, Ramesh Rana, Nikesh Shrestha, Ram Bahadur Gurung, Cristian Linte, Angus Watson, Yash Raj Shrestha, Binod Bhattarai

机构 * Rochester Institute of Technology, Rochester, NY, USA(罗切斯特技术学院) Nepal Applied Mathematics and Informatics Institute for Research (NAAMII)(尼泊尔应用数学与信息技术研究所) Kathmandu University(加德满都大学) University of Lausanne(洛桑大学) University of Aberdeen(阿伯丁大学)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV、cs.LG

Comments Accepted at MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13964 2025-06-18 eess.IV cs.LG 62%

Comparison of ConvNeXt and Vision-Language Models for Breast Density Assessment in Screening Mammography

Yusdivia Molina-Román, David Gómez-Ortiz, Ernestina Menasalvas-Ruiz, José Gerardo Tamez-Peña, Alejandro Santos-Díaz

机构 * School of Engineering and Sciences(工程与科学学院) School of Medical and Health Sciences(医学与健康科学学院)

专题命中 医疗多模态 :radiology(abstract);分类 cs.LG、eess.IV

Comments 6 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11178 2025-06-16 cs.CV cs.LG cs.NE 62%

BrainMAP: Multimodal Graph Learning For Efficient Brain Disease Localization

Nguyen Linh Dan Le, Jing Ren, Ciyuan Peng, Chengyao Xie, Bowen Li, Feng Xia

机构 * School of Computing Technologies, RMIT University(计算技术学院,拉筹纳斯大学) Institute of Innovation, Science and Sustainability, Federation University Australia(创新、科学与可持续性研究所,联邦大学澳大利亚)

专题命中 医疗多模态 :pathology(abstract);分类 cs.CV、cs.LG

Comments 6 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06141 2025-06-05 cs.CV cs.AI cs.CL cs.LG 62%

MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization

Kangyu Zhu, Peng Xia, Yun Li, Hongtu Zhu, Sheng Wang, Huaxiu Yao

机构 * UNC Chapel-Hill(UNC夏洛特-希尔分校) Brown University(布朗大学) University of Washington(华盛顿大学)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV、cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02229 2025-06-04 cs.CV cs.AI cs.CL cs.LG 62%

VLCD: Vision-Language Contrastive Distillation for Accurate and Efficient Automatic Placenta Analysis

Manas Mehta, Yimu Pan, Kelly Gallagher, Alison D. Gernand, Jeffery A. Goldstein, Delia Mwinyelle, Leena Mithal, James Z. Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Northwestern University(西北大学) University of Chicago(芝加哥大学) Lurie Children’s Hospital(Lurie儿童医院)

专题命中 医疗多模态 :pathology(abstract);分类 cs.CV、cs.LG

Comments Proceedings of the 9th International Workshop on Health Intelligence, in conjunction with the Annual AAAI Conference on Artificial Intelligence, Philadelphia, Pennsylvania, March 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24238 2025-06-03 cs.CV cs.LG 62%

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM

Bowen Dong, Minheng Ni, Zitong Huang, Guanglei Yang, Wangmeng Zuo, Lei Zhang

机构 * Harbin Institute of Technology(哈尔滨工业大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01377 2025-06-03 cs.CL cs.AI cs.CV cs.LG 62%

Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Yucheng Zhou, Lingran Song, Jianbing Shen

机构 * SKL-IOTSC, CIS, University of Macau(SKL-IOTSC、CIS、澳门大学)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV、cs.LG

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏