arXivDaily arXiv每日学术速递 周一至周五更新

科学与医疗

医学 AI

医学智能、临床 AI、医学影像、病理、诊断和医疗健康大模型。

共收录 966 信号源:cs.CV, cs.LG, q-bio, eess.IV, eess.SP

1. 医疗多模态 966 篇

2512.13747 2025-12-17 cs.CV cs.AI 61%

Why Text Prevails: Vision May Undermine Multimodal Medical Decision Making

为何文本占上风:视觉可能损害多模态医疗决策制定

Siyuan Dai, Lunxiao Li, Kun Zhao, Eardi Lila, Paul K. Crane, Heng Huang, Dongkuan Xu, Haoteng Tang, Liang Zhan

机构 * University of Texas Rio Grande Valley(德克萨斯大学里奥格兰德谷大学) University of Pittsburgh(匹兹堡大学) NC State University(北卡罗来纳州立大学) University of Washington(华盛顿大学) University of Maryland(马里兰大学)

专题命中 医疗多模态 :biomedical(abstract,comments);分类 cs.CV

AI总结 本研究发现文本推理在医疗多模态决策中优于多模态输入,提出三种策略以提升多模态医疗决策能力。

Comments Accepted by ICDM 2025 the Workshop on Synergy of AI and Multimodal Biomedical Data Mining

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15425 2025-05-26 cs.CV 61%

On the Robustness of Medical Vision-Language Models: Are they Truly Generalizable?

Raza Imam, Rufael Marew, Mohammad Yaqub

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·穆萨大学人工智能学院)

专题命中 医疗多模态 :medical image(abstract,comments);分类 cs.CV

Comments Dataset and Code is available at https://github.com/BioMedIA-MBZUAI/RobustMedCLIP Accepted at: Medical Image Understanding and Analysis (MIUA) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.08125 2025-05-01 cs.LG stat.ML 61%

A Large-scale Multimodal Study for Predicting Mortality Risk Using Minimal and Low Parameter Models and Separable Risk Assessment

Alvaro E. Ulloa Cerna, Marios Pattichis, David P. vanMaanen, Linyuan Jing, Aalpen A. Patel, Joshua V. Stough, Christopher M. Haggerty, Brandon K. Fornwalt

机构 * Department of Translational Data Science and Informatics, Geisinger(转化数据科学与信息学部门,Geisinger) Department of Electrical and Computer Engineering, University of New Mexico(电气与计算机工程系,新墨西哥大学)

专题命中 医疗多模态 :biomedical(abstract,journal_ref);分类 cs.LG

Journal ref IEEE Journal of Biomedical and Health Informatics, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10775 2025-01-22 cs.CV cs.AI 61%

MedFILIP: Medical Fine-grained Language-Image Pre-training

Xinjie Liang, Xiangyu Li, Fanding Li, Jie Jiang, Qing Dong, Wei Wang, Kuanquan Wang, Suyu Dong, Gongning Luo, Shuo Li

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV;biomedical(comments)

Comments 10 pages, 5 figures, IEEE Journal of Biomedical and Health Informatics 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09874 2024-11-18 cs.AI eess.SP 61%

A Hybrid Artificial Intelligence System for Automated EEG Background Analysis and Report Generation

Chin-Sung Tung, Sheng-Fu Liang, Shu-Feng Chang, Chung-Ping Young

专题命中 医疗多模态 :diagnosis(abstract);分类 eess.SP;biomedical(journal_ref)

Comments Example code available at https://github.com/tcs211/AI_EEEG_REPORT

Journal ref IEEE Journal of Biomedical and Health Informatics (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09886 2023-07-20 cs.CV cs.AI 61%

A reinforcement learning approach for VQA validation: an application to diabetic macular edema grading

Tatiana Fountoukidou, Raphael Sznitman

专题命中 医疗多模态 :medical image(abstract,journal_ref);分类 cs.CV

Comments 16 pages (+ 23 pages supplementary material)

Journal ref Medical image analysis 87 (2023): 102822

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.09808 2022-06-15 cs.CV 61%

Integrated Construction of Multimodal Atlases with Structural Connectomes in the Space of Riemannian Metrics

Kristen M. Campbell, Haocheng Dai, Zhe Su, Martin Bauer, P. Thomas Fletcher, Sarang C. Joshi

专题命中 医疗多模态 :MRI(abstract);分类 cs.CV;biomedical(comments)

Comments Accepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) https://www.melba-journal.org/papers/2022:016.html. arXiv admin note: substantial text overlap with arXiv:2103.05730

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.06708 2018-09-18 cs.CV 61%

Speech Map: A Statistical Multimodal Atlas of 4D Tongue Motion During Speech from Tagged and Cine MR Images

Jonghye Woo, Fangxu Xing, Maureen Stone, Jordan Green, Timothy G. Reese, Thomas J. Brady, Van J. Wedeen, Jerry L. Prince, Georges El Fakhri

专题命中 医疗多模态 :MRI(abstract);分类 cs.CV;biomedical(comments)

Comments Accepted at Journal of Computer Methods in Biomechanics and Biomedical Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29903 2026-04-01 q-bio.NC eess.SP 60%

Multimodal Higher-Order Brain Networks: A Topological Signal Processing Perspective

多模态高阶脑网络:拓扑信号处理视角

Breno C. Bispo, Stefania Sardellitti, Juliano B. Lima, Fernando A. N. Santos

专题命中 医疗多模态 :MRI(abstract);分类 q-bio、eess.SP

AI总结 本文提出基于拓扑信号处理的多模态框架,通过高阶拓扑域建模脑部,利用扩散MRI和静息态fMRI学习个体化脑细胞复合体,揭示高阶相互作用的拓扑特征及其与行为的关联。

Comments This paper has been sumbmitted to IEEE Transactions on Medical Imaging (TMI), March 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18314 2026-02-13 q-bio.QM cs.LG q-bio.NC 60%

BrainSymphony: A parameter-efficient multimodal foundation model for brain dynamics with limited data

BrainSymphony: 一种参数高效、多模态的基础模型,用于在有限数据下的脑动态

Moein Khajehnejad, Forough Habibollahi, Devon Stoliker, Adeel Razi

机构 * Turner Institute for Brain and Mental Health(大脑与心理健康Turner研究所) School of Psychological Sciences, Monash University(墨尔本大学心理学科学学院) Cortical Labs(皮层实验室) CIFAR Azrieli Global Scholars Program(CIFAR阿兹里埃利全球学者计划)

专题命中 医疗多模态 :MRI(abstract);分类 cs.LG、q-bio

AI总结 BrainSymphony是一种参数高效、多模态的基础模型,通过整合fMRI和扩散MRI数据,实现有限数据下的脑动态分析,优于更大模型并提升神经科学应用。

Comments 32 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.05748 2026-02-06 q-bio.NC cs.LG 60%

Analyzing heterogeneity in Alzheimer Disease using multimodal normative modeling on imaging-based ATN biomarkers

利用多模态规范建模分析阿尔茨海默病的异质性:基于影像学ATN生物标志物

Sayantan Kumar, Tom Earnest, Braden Yang, Deydeep Kothapalli, Andrew J. Aschenbrenner, Jason Hassenstab, Chengie Xiong, Beau Ances, John Morris, Tammie L. S. Benzinger, Brian A. Gordon, Philip Payne, Aristeidis Sotiras

专题命中 医疗多模态 :MRI(abstract);分类 cs.LG、q-bio

AI总结 本研究利用多模态规范建模分析阿尔茨海默病影像学ATN生物标志物的异质性,揭示了疾病严重程度与认知功能的关系。

Comments Under review in Alzheimer's & Dementia

Journal ref Alzheimer's Dement. 2025; 21:e70143

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13292 2026-01-14 q-bio.QM cs.AI eess.IV 60%

An interpretable generative multimodal neuroimaging-genomics framework for decoding Alzheimer's disease

可解释的生成多模态神经影像-基因组框架用于解码阿尔茨海默病

Giorgio Dolci, Federica Cruciani, Md Abdur Rahaman, Anees Abrol, Jiayu Chen, Zening Fu, Ilaria Boscolo Galazzo, Gloria Menegaz, Vince D. Calhoun

机构 * Department of Computer Science, University of Verona(威尼斯大学计算机科学系) Department of Engineering for Innovation Medicine, University of Verona(威尼斯大学创新医学工程系) Tri-Institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS), Georgia State University, Georgia Institute of Technology, Emory University(神经影像与数据科学转化研究三机构中心(TReNDS),佐治亚州立大学,佐治亚理工学院,埃默里大学)

专题命中 医疗多模态 :MRI(abstract);分类 q-bio、eess.IV

AI总结 本文提出了一种可解释的生成多模态神经影像-基因组框架,用于解码阿尔茨海默病,通过多模态数据和单核苷酸多态性实现AD检测和MCI预测,并揭示了与疾病相关的生物学机制。

Comments 33 pages, 8 figures (main text + supplementary materials), submitted to a journal

Journal ref J. Neural Eng. 22 056021 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02403 2025-10-06 q-bio.QM cs.AI cs.CV 60%

Glaucoma Detection and Structured OCT Report Generation via a Fine-tuned Multimodal Large Language Model

Jalil Jalili, Yashraj Gavhane, Evan Walker, Anna Heinke, Christopher Bowd, Akram Belghith, Massimo A. Fazio, Christopher A. Girkin, C. Gustavo De Moraes, Jeffrey M. Liebmann, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV、q-bio

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07367 2025-07-11 q-bio.BM cs.LG 60%

Platform for Representation and Integration of multimodal Molecular Embeddings

Erika Yilin Zheng, Yu Yan, Baradwaj Simha Sankar, Ethan Ji, Steven Swee, Irsyad Adam, Ding Wang, Alexander Russell Pelletier, Alex Bui, Wei Wang, Peipei Ping

专题命中 医疗多模态 :biomedical(abstract);分类 cs.LG、q-bio

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18253 2025-06-10 cs.LG cs.AI q-bio.QM 60%

Multimodal Integration of Longitudinal Noninvasive Diagnostics for Survival Prediction in Immunotherapy Using Deep Learning

Melda Yeghaian, Zuhir Bodalal, Daan van den Broek, John B A G Haanen, Regina G H Beets-Tan, Stefano Trebeschi, Marcel A J van Gerven

专题命中 医疗多模态 :CT(abstract);分类 cs.LG、q-bio

Journal ref Journal of the American Medical Informatics Association, 2025;, ocaf074

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01456 2025-06-03 q-bio.GN cs.AI cs.LG q-bio.NC 60%

GenDMR: A dynamic multimodal role-swapping network for identifying risk gene phenotypes

Lina Qin, Cheng Zhu, Chuqi Zhou, Yukun Huang, Jiayi Zhu, Ping Liang, Jinju Wang, Yixing Huang, Cheng Luo, Dezhong Yao, Ying Tan

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.LG、q-bio

Comments 31 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15534 2024-09-25 cs.LG cs.AI cs.CL q-bio.QM 60%

Geneverse: A collection of Open-source Multimodal Large Language Models for Genomic and Proteomic Research

Tianyu Liu, Yijia Xiao, Xiao Luo, Hua Xu, W. Jim Zheng, Hongyu Zhao

专题命中 医疗多模态 :biomedical(abstract);分类 cs.LG、q-bio

Comments 8 pages

Journal ref EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15132 2024-07-23 q-bio.NC cs.LG 60%

Deep multimodal saliency parcellation of cerebellar pathways: linking microstructure and individual function through explainable multitask learning

Ari Tchetchenian, Leo Zekelman, Yuqian Chen, Jarrett Rushmore, Fan Zhang, Edward H. Yeterian, Nikos Makris, Yogesh Rathi, Erik Meijering, Yang Song, Lauren J. O'Donnell

专题命中 医疗多模态 :MRI(abstract);分类 cs.LG、q-bio

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09484 2023-07-24 q-bio.BM cs.CE cs.LG physics.chem-ph 60%

MolFM: A Multimodal Molecular Foundation Model

Yizhen Luo, Kai Yang, Massimo Hong, Xing Yi Liu, Zaiqing Nie

专题命中 医疗多模态 :biomedical(abstract);分类 cs.LG、q-bio

Comments 31 pages, 15 figures, and 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16509 2022-12-20 q-bio.GN cs.AI cs.LG q-bio.BM stat.ML 60%

Multimodal Learning for Multi-Omics: A Survey

Sina Tabakhi, Mohammod Naimul Islam Suvon, Pegah Ahadian, Haiping Lu

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.LG、q-bio

Comments 52 pages, 3 figures; Revised matrix factorization fusion section

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15517 2022-03-30 q-bio.NC cs.LG 60%

A multimodal approach for Parkinson disease analysis

Marcos Faundez-Zanuy, Antonio Satue-Villar, Jiri Mekyska, Viridiana Arreola, Pilar Sanz, Carles Paul, Luis Guirao, Mateu Serra, Laia Rofes, Pere Clavé, Enric Sesa-Nogueras, Josep Roure

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.LG、q-bio

Comments 10 pages

Journal ref In: Bassis S., Esposito A., Morabito F. (eds) Advances in Neural Networks: Computational and Theoretical Issues. Smart Innovation, Systems and Technologies, vol 37. Springer, Cham. 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.02121 2020-03-31 physics.med-ph cs.LG physics.bio-ph q-bio.GN 60%

Next Generation Radiogenomics Sequencing for Prediction of EGFR and KRAS Mutation Status in NSCLC Patients Using Multimodal Imaging and Machine Learning Approaches

Isaac Shiri, Hassan Maleki, Ghasem Hajianfar, Hamid Abdollahi, Saeed Ashrafinia, Mathieu Hatt, Mehrdad Oveisi, Arman Rahmim

专题命中 医疗多模态 :CT(abstract);分类 cs.LG、q-bio

Comments 42 pages,3 Figures,4 Tables, 13 Supplemental Figures, 11 Supplemental Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.00936 2016-09-02 cs.GR q-bio.NC 59%

Multimodal Brain Visualization

Saad Nadeem, Arie Kaufman

专题命中 医疗多模态 :MRI(abstract);分类 q-bio;biomedical(comments)

Comments SPIE Medical Imaging 2016, Proc. SPIE Medical Imaging: Biomedical Applications in Molecular, Structural, and Functional Imaging, 2016

Journal ref SPIE Medical Imaging, pp. 97881Y-97881Y. 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12498 2026-08-17 cs.CV 版本更新 57%

NAST: Improving Negation Handling in Medical Vision-Language Models through Negation-Aware Selective Training

针对医学视觉-语言模型中否定处理改进的分层微调

Ali Abbasi, Mehdi Taghipour, Rahmatollah Beheshti

机构 * University of Delaware(德克萨斯大学)

专题命中 医疗多模态 :radiology(abstract);分类 cs.CV

AI总结 本研究提出NAST方法,通过因果追溯效应调节逐层梯度更新,提升医学视觉-语言模型对否定处理的识别能力,同时保持一般视觉-语言对齐。

Comments 16 pages, 6 figures. Accepted to MLHC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11024 2026-08-12 cs.CV 新提交 57%

When Visual Signals Mislead: A Mechanistic Study of Attribute Hallucination in Vision-Language Models

当视觉信号产生误导:视觉语言模型中属性幻觉的机制研究

Yufei Zhang, Chenlu Zhan, Hongwei Wang

机构 * Zhejiang University(浙江大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 针对视觉语言模型的属性幻觉问题,提出VISOR框架,通过VSNR诊断区分失败模式并路由适配操作,在三类模型上减少属性假阳性且不依赖先验主导假设。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08494 2026-08-12 cs.CV 版本更新 57%

Linguistically-Aligned and Visually-Grounded Preference Optimization for Clinically-Augmented Medical Report Generation

面向临床增强型医疗报告生成的语言对齐与视觉 grounded 的偏好优化

Qiang Hu, Yuxuan Luo, Yingjie Guo, Hao Wang, Qimei Wang, Qiang Li, Zhiwei Wang

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 针对现有医疗报告生成方法的临床事实错误与视觉-语言对齐不足问题,提出DPO-Clin框架,通过ECD模块、M²DPO及反事实偏好数据构建提升模型可靠性,在多医学数据集上取得优异性能。

Comments Accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26738 2026-08-06 cs.CV cs.AI cs.CL 版本更新 57%

SleepVLM: A Rule-Grounded Vision-Language Model for Auditable Sleep Staging

SleepVLM:基于视觉语言模型的可解释且规则驱动的睡眠分期

Guifeng Deng, Pan Wang, Mengfan Niu, Jiquan Wang, Shuying Rao, Junyi Xie, Xi'ang Chen, Sha Zhao, Gang Pan, Wanjun Guo, Tao Li, Haiteng Jiang

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 提出SleepVLM,一种基于规则驱动的视觉语言模型,通过多通道PSG波形图像进行睡眠分期,并生成符合AASM评分标准的临床可读解释,在保持高准确率的同时提升可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02692 2026-08-05 cs.LG 新提交 57%

PatTree: a novel approach for automated creation of multimodal, graph-based patient representations for medical classification tasks

PatTree:一种用于医学分类任务的多模态图基患者表示自动构建新方法

Julia Gehrmann, Lars Quakulinski, Hamza Naseem, Oya Beyan

专题命中 医疗多模态 :clinical AI(abstract);分类 cs.LG

AI总结 该研究提出PatTree,一种自动构建的多模态图基患者表示,在ADNI-1队列子集上实现阿尔茨海默病等三分类任务98.5%平衡准确率,可作为临床AI流程的可扩展基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01456 2026-08-04 cs.CV cs.CL 新提交 57%

Long-Horizon Embodied Decision-Making via Multimodal Memory Compression

基于多模态记忆压缩的长视野具身决策

Bingxuan Li, Rui Yang, Cheng Qian, Jiateng Liu, Jeonghwan Kim, Zhenhailong Wang, Manling Li, Tong Zhang, Heng Ji

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Northwestern University(西北大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 该研究提出DunphyBench基准用于评估长视野具身决策,发现当前智能体与人类表现存在差距,设计了MeMento记忆压缩器,使VLM驱动智能体准确率提升7.18%且内存使用降低85.38%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00976 2026-08-04 cs.CV 新提交 57%

Location-Aware Fine-Grained Representation Learning for Medical Vision Foundation Models

面向医学视觉基础模型的位置感知细粒度表示学习

Myeongkyun Kang, Yanting Yang, Xiaoxiao Li

机构 * The University of British Columbia(不列颠哥伦比亚大学) Vector Institute(矢量研究所)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 本研究提出基于位置感知细粒度表示学习的医学视觉基础模型LoFi,构建大规模医学定位数据集MedG,在多项医学视觉任务中性能优于现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏