arXivDaily arXiv每日学术速递 周一至周五更新

科学与医疗

医学 AI

医学智能、临床 AI、医学影像、病理、诊断和医疗健康大模型。

共收录 966 信号源:cs.CV, cs.LG, q-bio, eess.IV, eess.SP

1. 医疗多模态 966 篇

2507.17239 2025-07-24 cs.CV 57%

MaskedCLIP: Bridging the Masked and CLIP Space for Semi-Supervised Medical Vision-Language Pre-training

Lei Zhu, Jun Zhou, Rick Siow Mong Goh, Yong Liu

机构 * Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR)(高性能计算研究所(IHPC)、科技研究局(A*STAR))

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

Comments Accepted to MedAGI 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14738 2025-07-22 cs.CV 57%

MultiRetNet: A Multimodal Vision Model and Deferral System for Staging Diabetic Retinopathy

Jeannie She, Katie Spivakovsky

机构 * MIT(麻省理工学院)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11200 2025-07-21 cs.CV 57%

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study

Che Liu, Jiazhen Pan, Weixiang Shen, Wenjia Bai, Daniel Rueckert, Rossella Arcucci

机构 * Imperial College London, UK(伦敦帝国学院) Technical University of Munich, Germany(慕尼黑技术大学) Ludwig Maximilian University of Munich(慕尼黑路德维希-马克西米利安大学)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09816 2025-07-21 cs.CV 57%

Mind the Modality Gap: Towards a Remote Sensing Vision-Language Model via Cross-modal Alignment

Angelos Zavras, Dimitrios Michail, Begüm Demir, Ioannis Papoutsis

机构 * organization= Orion Lab, National Observatory of Athens \& National Technical University of Athens , country= Greece organization= Department of Informatics \& Telematics, Harokopio University of Athens , country= Greece organization= Faculty of Electrical Engineering organization= BIFOLD - Berlin Institute for the Foundations of Learning

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

Comments Accepted at the ISPRS Journal of Photogrammetry and Remote Sensing. Our code implementation and weights for all experiments are publicly available at https://github.com/Orion-AI-Lab/MindTheModalityGap

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08607 2025-07-14 cs.CV 57%

BayesTTA: Continual-Temporal Test-Time Adaptation for Vision-Language Models via Gaussian Discriminant Analysis

Shuang Cui, Jinglin Xu, Yi Li, Xiongxin Tang, Jiangmeng Li, Jiahuan Zhou, Fanjiang Xu, Fuchun Sun, Hui Xiong

机构 * National Key Laboratory of Space Integrated Information System(国家空间信息集成系统重点实验室) Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学) Wangxuan Institute of Computer Technology(计算机技术王轩研究所) Peking University(北京大学) Department of Computer Science and Technology(计算机科学与技术系) Tsinghua University(清华大学) Thrust of Artificial Intelligence, the Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能方向) Department of Computer Science & Engineering, the Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)

专题命中 医疗多模态 :CT(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08024 2025-07-14 cs.CV 57%

Self-Consistency in Vision-Language Models for Precision Agriculture: Multi-Response Consensus for Crop Disease Management

Mihir Gupta, Abhay Mangla, Ross Greer, Pratik Desai

机构 * The Harker School, USA(哈克尔学校) Dougherty Valley High School, USA(道格拉斯谷高中) Kissan.ai, USA(加州大学梅尔德分校) University of California, Merced, USA

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07902 2025-07-11 cs.CV 57%

MIRA: A Novel Framework for Fusing Modalities in Medical RAG

Jinhong Wang, Tajamul Ashraf, Zongyan Han, Jorma Laaksonen, Rao Mohammad Anwer

机构 * Department of Computer Vision, MBZUAI(视觉计算系,MBZUAI) Department of Computer Science, Aalto University(计算机科学系,阿alto大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02900 2025-07-11 cs.CV 57%

MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine

Yunfei Xie, Ce Zhou, Lang Gao, Juncheng Wu, Xianhang Li, Hong-Yu Zhou, Sheng Liu, Lei Xing, James Zou, Cihang Xie, Yuyin Zhou

机构 * Huazhong University of Science and Technology(华中科技大学) UC Santa Cruz(加州大学圣克ruz分校) Harvard University(哈佛大学) Stanford University(斯坦福大学)

专题命中 医疗多模态 :medical AI(abstract);分类 cs.CV

Comments The dataset is publicly available at https://yunfeixie233.github.io/MedTrinity-25M/. Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02908 2025-07-08 cs.LG cs.AI 57%

Hyperbolic Kernel Graph Neural Networks for Neurocognitive Decline Analysis from Multimodal Brain Imaging

Meimei Yang, Yongheng Sun, Qianqian Wang, Andrea Bozoki, Maureen Kohi, Mingxia Liu

机构 * Department of Radiology and Biomedical Research Imaging Center (BRIC), University of North Carolina at Chapel Hill(放射科与生物医学研究成像中心(BRIC)、北卡罗来纳大学教堂山分校)

专题命中 医疗多模态 :MRI(abstract);分类 cs.LG

Comments 14 pages, 5 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00052 2025-07-02 cs.CV cs.AI 57%

VSF-Med:A Vulnerability Scoring Framework for Medical Vision-Language Models

Binesh Sadanandan, Vahid Behzadan

机构 * SAIL Lab, University of New Haven, West Haven, CT, USA(SAIL实验室,新罕布什尔大学,西哈文,康涅狄格州,美国)

专题命中 医疗多模态 :radiology(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04397 2025-07-01 cs.CL cs.AI cs.LG 57%

Multimodal Medical Code Tokenizer

Xiaorui Su, Shvat Messica, Yepeng Huang, Ruth Johnson, Lukas Fesser, Shanghua Gao, Faryad Sahneh, Marinka Zitnik

机构 * Department of Biomedical Informatics, Harvard Medical School, Boston, MA, USA(生物医学信息学系,哈佛医学院,波士顿,马萨诸塞州,美国)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.LG

Comments ICML'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08870 2025-07-01 cs.CL cs.AI cs.LG 57%

The Limited Impact of Medical Adaptation of Large Language and Vision-Language Models

Daniel P. Jeong, Pranav Mani, Saurabh Garg, Zachary C. Lipton, Michael Oberst

机构 * Carnegie Mellon University(卡内基梅隆大学) Abridge Mistral AI Johns Hopkins University(约翰霍普金斯大学)

专题命中 医疗多模态 :biomedical(abstract);分类 cs.LG

Comments Extended version of EMNLP 2024 paper arXiv:2411.04118. Includes additional results on clinical note QA tasks and supervised fine-tuning evaluations

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05319 2025-06-26 cs.CV cs.AI 57%

Robust Multimodal Learning for Ophthalmic Disease Grading via Disentangled Representation

Xinkun Wang, Yifang Wang, Senwei Liang, Feilong Tang, Chengzhi Liu, Ming Hu, Chao Hu, Junjun He, Zongyuan Ge, Imran Razzak

机构 * MBZUAI, United Arab Emirates(MBZUAI,阿联酋) Monash University, Australia(墨尔本大学,澳大利亚) Liverpool University, United Kingdom(利物浦大学,英国) China Unicom (Shanghai) Industrial Internet Co., Ltd., China(中国联合(上海)工业互联网有限公司,中国) Shanghai AI Lab, China(上海人工智能实验室,中国)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

Comments 10pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19324 2025-06-25 cs.CV 57%

Memory-Augmented Incomplete Multimodal Survival Prediction via Cross-Slide and Gene-Attentive Hypergraph Learning

Mingcheng Qu, Guang Yang, Donglin Di, Yue Gao, Tonghua Su, Yang Song, Lei Fan

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) School of Software, Tsinghua University(清华大学软件学院) School of Computer Science and Engineering, UNSW Sydney(新南威尔士大学计算机科学与工程学院)

专题命中 医疗多模态 :pathology(abstract);分类 cs.CV

Comments accepted by MICCAI2025 code: https://github.com/MCPathology/M2Surv

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05349 2025-06-25 cs.CV 57%

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos

Hanoona Rasheed, Abdelrahman Shaker, Anqi Tang, Muhammad Maaz, Ming-Hsuan Yang, Salman Khan, Fahad Shahbaz Khan

机构 * MBZUAI(穆桑大学人工智能研究所) University of California Merced(加州大学默塞德分校) Google Research(谷歌研究) Australian National University(澳大利亚国立大学) Linköping University(林肯大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

Comments VideoMathQA Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17500 2025-06-24 cs.CV 57%

Few-Shot, Now for Real: Medical VLMs Adaptation without Balanced Sets or Validation

Julio Silva-Rodríguez, Fereshteh Shakeri, Houda Bahig, Jose Dolz, Ismail Ben Ayed

机构 * ÉTS Montréal(ÉTS蒙特利尔) Centre de Recherche du Centre Hospitalier de l’Université de Montréal (CRCHUM)(蒙特利尔大学中心医院研究中心(CRCHUM))

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

Comments MICCAI 2025. Code: https://github.com/jusiro/SS-Text

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15477 2025-06-19 cs.CV 57%

Multimodal Large Language Models for Medical Report Generation via Customized Prompt Tuning

Chunlei Li, Jingyang Hou, Yilei Shi, Jingliang Hu, Xiao Xiang Zhu, Lichao Mou

机构 * MedAI Technology (Wuxi) Co. Ltd.(MedAI技术(无锡)有限公司) Technical University of Munich(慕尼黑技术大学)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12849 2025-06-17 cs.CV 57%

CAPO: Reinforcing Consistent Reasoning in Medical Decision-Making

Songtao Jiang, Yuan Wang, Ruizhe Chen, Yan Zhang, Ruilin Luo, Bohan Lei, Sibo Song, Yang Feng, Jimeng Sun, Jian Wu, Zuozhu Liu

机构 * Zhejiang University(浙江大学) Tsinghua University(清华大学) Alibaba Group(阿里巴巴集团) Angelalign Inc.(Angelalign公司) UIUC(伊利诺伊大学香槟分校)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06600 2025-06-17 cs.CV 57%

RARL: Improving Medical VLM Reasoning and Generalization with Reinforcement Learning and LoRA under Data and Hardware Constraints

Tan-Hanh Pham, Chris Ngo

机构 * Harvard Medical School(哈佛医学院) Knovel Engineering Lab(Knovel工程实验室)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13509 2025-06-17 cs.CL cs.AI cs.LG 57%

ProMedTS: A Self-Supervised, Prompt-Guided Multimodal Approach for Integrating Medical Text and Time Series

Shuai Niu, Jing Ma, Hongzhan Lin, Liang Bai, Zhihua Wang, Wei Bi, Yida Xu, Guo Li, Xian Yang

机构 * Hong Kong Baptist University(香港 Baptist 大学) Shanxi University(山西大学) Shanghai Institute for Advanced Study of Zhejiang University(浙江大学上海研究院) Tencent AI Lab(腾讯人工智能实验室) Manchester Metropolitan University(曼彻斯特 Metropolitan 大学) The University of Manchester(曼彻斯特大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.LG

Comments This paper is accepted by ACL2025(Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06084 2025-06-09 cs.CV 57%

WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management

Bowen Yuan, Selena Song, Javier Fernandez, Yadan Luo, Mahsa Baktashmotlagh, Zijian Wang

机构 * The University of Queensland(昆士兰大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06076 2025-06-09 cs.CV 57%

Full Conformal Adaptation of Medical Vision-Language Models

Julio Silva-Rodríguez, Leo Fillioux, Paul-Henry Cournède, Maria Vakalopoulou, Stergios Christodoulidis, Ismail Ben Ayed, Jose Dolz

机构 * ÉTS Montréal(蒙特利尔ÉTS学院) CentraleSupélec, Université Paris-Saclay(中央理工学院,巴黎萨克雷大学)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

Comments IPMI 2025. Code: https://github.com/jusiro/FCA

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10332 2025-06-05 cs.CV cs.AI 57%

Prescribing the Right Remedy: Mitigating Hallucinations in Large Vision-Language Models via Targeted Instruction Tuning

Rui Hu, Yahan Tu, Shuyu Wei, Dongyuan Lu, Jitao Sang

机构 * Beijing Key Lab of Traffic Data Analysis and Mining(北京交通大数据分析与挖掘重点实验室) Beijing Jiaotong University(北京交通大学) School of Information Technology and Management(信息科学技术学院)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

Comments Accepted in Information Sciences 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01724 2025-06-03 cs.CV 57%

Active Learning via Vision-Language Model Adaptation with Open Data

Tong Wang, Jiaqi Wang, Shu Kong

机构 * University of Macau(澳门大学) Shanghai AI Lab(上海人工智能实验室) Institute of Collaborative Innovation project webpage(协同创新项目研究所)

专题命中 医疗多模态 :CT(abstract);分类 cs.CV

Comments Here is the project webpage: https://leowangtong.github.io/ALOR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00805 2025-06-03 cs.CV cs.CL 57%

HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models

Songtao Jiang, Yan Zhang, Yeying Jin, Zhihang Tang, Yangyang Wu, Yang Feng, Jian Wu, Zuozhu Liu

机构 * Zhejiang University(浙江大学) Byte Dance(字节跳动) National University of Singapore(新加坡国立大学) Angelalign Inc China(Angelalign中国公司) Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence(浙江省医学影像人工智能重点实验室)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24059 2025-06-02 cs.LG 57%

Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition

Sean Foley, Hong Nguyen, Jihwan Lee, Sudarsana Reddy Kadiri, Dani Byrd, Louis Goldstein, Shrikanth Narayanan

机构 * University of Southern CaliforniaUSA(美国南加州大学) Signal Analysis and Interpretation Laboratory(信号分析与解释实验室) Department of Linguistics(语言学系)

专题命中 医疗多模态 :MRI(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19195 2025-05-27 cs.AI cs.CV 57%

CardioCoT: Hierarchical Reasoning for Multimodal Survival Analysis

Shaohao Rui, Haoyang Su, Jinyi Xiang, Lian-Ming Wu, Xiaosong Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院) Fudan University(复旦大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 医疗多模态 :MRI(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17835 2025-05-26 cs.CV 57%

VLM Models and Automated Grading of Atopic Dermatitis

Marc Lalonde, Hamed Ghodrati

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03350 2025-05-07 cs.CV 57%

A Vision-Language Model for Focal Liver Lesion Classification

Song Jian, Hu Yuchang, Wang Hui, Chen Yen-Wei

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

Comments 9 pages,4 figures, 4 tables,Innovation in Medicine and Healthcare Proceedings of 13th KES-InMed 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21051 2025-05-01 cs.LG cs.CL cs.MM 57%

Multimodal Large Language Models for Medicine: A Comprehensive Survey

Jiarui Ye, Hao Tang

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) School of Computer Science, Peking University(北京大学计算机科学学院)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏