arXivDaily arXiv每日学术速递 周一至周五更新

科学与医疗

医学 AI

医学智能、临床 AI、医学影像、病理、诊断和医疗健康大模型。

共收录 966 信号源:cs.CV, cs.LG, q-bio, eess.IV, eess.SP

1. 医疗多模态 966 篇

2603.02754 2026-03-04 cs.CV 57%

Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing

无需训练即可清晰感知:减轻多模态大语言模型在遥感中的幻觉

Yi Liu, Jing Zhang, Di Wang, Xiaoyu Tian, Haonan Guo, Bo Du

机构 * School of Computer Science, Wuhan University, China(武汉大学计算机学院) Zhongguancun Academy, China(中关村学院) School of Computer Science, Chongqing University, China(重庆大学计算机学院)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 RADAR通过利用多模态大语言模型的内在注意力,无需训练即可减少遥感视觉问答中的事实性和逻辑性幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00289 2026-03-03 cs.CV 57%

Seeking Necessary and Sufficient Information from Multimodal Medical Data

从多模态医学数据中寻求必要和充分的信息

Boyu Chen, Weiye Bao, Junjie Liu, Michael Shen, Bo Peng, Paul Taylor, Zhu Li, Mengyue Yang

机构 * University College London, London, UK(伦敦大学学院) Imperial College London, London, UK(伦敦帝国学院) Mingdu Tech, China(明都科技) University of Bristol, Bristol, UK(布里斯托大学)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 本文提出通过概率必要性和充分性学习多模态医学数据中的必要和充分特征,以提升模型性能和鲁棒性。

Comments 11 pages, 1 figure. Submitted to MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25867 2026-02-19 cs.LG 57%

Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs

从医学文献合成高质量的视觉问答系统:基于生成-验证框架的大型多模态模型

Xiaoke Huang, Ningsen Wang, Hui Liu, Xianfeng Tang, Yuyin Zhou

机构 * UC Santa Cruz(加州大学圣克ruz分校) Fudan University(复旦大学) Amazon Research(亚马逊研究)

专题命中 医疗多模态 :biomedical(abstract);分类 cs.LG

AI总结 MedVLSynther通过生成-验证框架从开放文献合成高质量医学VQA数据,提升六个基准测试的准确率,达到77.57的VQA-RAD表现。

Comments Project page, code, data, and models: https://ucsc-vlaa.github.io/MedVLSynther/ ; Accepted by ICLR'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06288 2026-02-17 cs.CV 57%

Unsupervised MR-US Multimodal Image Registration with Multilevel Correlation Pyramidal Optimization

无监督的MR-US多模态图像配准与多级相关金字塔优化

Jiazheng Wang, Zeyu Liu, Min Liu, Xiang Chen, Xinyao Yu, Yaonan Wang, Hang Zhang

机构 * School of Artificial Intelligence and Robotics, Hunan University, Changsha, Hunan, China(人工智能与机器人学院,湖南大学,长沙,湖南,中国) National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, Changsha, Hunan, China(机器人视觉感知与控制技术国家工程研究中心,湖南大学,长沙,湖南,中国) National University of Singapore(新加坡国立大学) Cornell University(康奈尔大学)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 本研究提出基于多级相关金字塔优化的无监督多模态图像配准方法,以解决术前与术中多模态图像的配准问题,实现了在ReMIND2Reg任务中的优异表现。

Comments first-place method of ReMIND2Reg Learn2Reg MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13650 2026-02-17 cs.CV cs.AI cs.CL 57%

KorMedMCQA-V: A Multimodal Benchmark for Evaluating Vision-Language Models on the Korean Medical Licensing Examination

KorMedMCQA-V: 一种用于评估视觉语言模型在韩国医学资格考试中多模态多项选择问答能力的基准测试

Byungjin Choi, Seongsu Bae, Sunjun Kweon, Edward Choi

机构 * Ajou University School of Medicine(阿乔大学医学院) KAIST(韩国科学技术院)

专题命中 医疗多模态 :CT(abstract);分类 cs.CV

AI总结 KorMedMCQA-V是一个用于评估视觉语言模型在韩国医学资格考试中多模态多项选择问答能力的基准测试,展示了不同模型在医疗领域中的表现差异。

Comments 17 pages, 2 figures, 6 tables. (Includes appendix.)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13459 2026-02-17 eess.SP 57%

Towards Causality-Aware Modeling for Multimodal Brain-Muscle Interactions

迈向因果意识的多模态脑-肌相互作用建模

Farwa Abbas, Wei Dai, Zoran Cvetkovic, Verity McClelland

专题命中 医疗多模态 :biomedical(abstract);分类 eess.SP

AI总结 本文提出一种结合几何流形重建与概率时间建模的DBN启发CCM框架,用于多模态脑-肌交互的因果建模,揭示了肌张力障碍中特定频率的通路重组织,并展示了其在生物标志物开发和神经调节干预中的应用潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10624 2026-02-12 cs.CV cs.AI 57%

A Vision-Language Foundation Model for Zero-shot Clinical Collaboration and Automated Concept Discovery in Dermatology

一种用于零样本临床协作和皮肤病自动概念发现的视觉-语言基础模型

Siyuan Yan, Xieji Li, Dan Mo, Philipp Tschandl, Yiwen Jiang, Zhonghua Wang, Ming Hu, Lie Ju, Cristina Vico-Alonso, Yizhen Zheng, Jiahe Liu, Juexiao Zhou, Camilla Chello, Jen G. Cheung, Julien Anriot, Luc Thomas, Clare Primiero, Gin Tan, Aik Beng Ng, Simon See, Xiaoying Tang, Albert Ip, Xiaoyang Liao, Adrian Bowling, Martin Haskett, Shuang Zhao, Monika Janda, H. Peter Soyer, Victoria Mar, Harald Kittler, Zongyuan Ge

机构 * AIM for Health Lab(AIM for Health实验室) Faculty of Information Technology, Monash University(信息技术学院,墨尔本大学) Department of Dermatology, Medical University of Vienna(皮肤科,维也纳医学大学) Faculty of Engineering, Monash University(工程学院,墨尔本大学) Institute of Ophthalmology, University College London(眼科研究所,伦敦大学学院) Dermatology Department. Fundacion Hospital 12 de Octubre(皮肤科部门,12月12日基金会医院) School of Data Science, The Chinese University of Hong Kong, Shenzhen(数据科学学院,香港中文大学(深圳)) Frazer Institute, The University of Queensland(弗拉泽研究所,昆士兰大学) Dermatology Research Centre, Brisbane, Australia(皮肤科研究中心,澳大利亚布里斯班) Department of Medical and Cardiovascular Sciences, Sapienza University of Rome(医学与心血管科学系,罗马萨皮恩扎大学) Victorian Melanoma Service, Alfred Care Group, Bayside Health, Melbourne, Australia(维多利亚黑色素瘤服务,阿尔弗雷德护理集团,墨尔本布里斯班健康中心,澳大利亚) SkIIN Discovery Program, Monash University(SkIIN发现计划,墨尔本大学) Claude Bernard Lyon-1 University(克劳德·伯纳德里昂-1大学) Centre Léon Berard, Lyon, France(莱昂·贝拉德中心,法国里昂) Dermatology department, Hôpital Lyon Sud, Hospices Civils de Lyon, Lyons, France(皮肤科部门,里昂南医院,里昂民事医院,法国里昂) Cancer Research Center of Lyon, Lyons, France(里昂癌症研究中心,法国里昂) eResearch Centre, Monash University(eResearch研究中心,墨尔本大学) NVIDIA AI Techno(NVIDIA AI技术)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 DermFM-Zero是一种用于皮肤病零样本临床协作和自动概念发现的视觉-语言基础模型,通过掩码潜在建模和对比学习训练,实现了无需任务特定适应的高精度诊断和多模态检索能力。

Comments reports

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10143 2026-02-12 cs.CV 57%

MPA: Multimodal Prototype Augmentation for Few-Shot Learning

MPA: 多模态原型增强用于少样本学习

Liwen Wu, Wei Wang, Lei Zhao, Zhan Gao, Qika Lin, Shaowen Yao, Zuozhu Liu, Bin Pu

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 MPA通过多模态原型增强方法,在少样本学习中实现更优性能,通过语义增强、多视图增强和不确定类吸收器提升模型表现。

Comments This paper has been accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06303 2026-02-11 cs.LG 57%

Multimodal Graph Neural Networks for Prognostic Modeling of Brain Network Reorganization

多模态图神经网络用于脑网络重组的预后建模

Preksha Girish, Rachana Mysore, Kiran K. N., Hiranmayee R., Shipra Prashanth, Shrey Kumar

专题命中 医疗多模态 :MRI(abstract);分类 cs.LG

AI总结 本文提出多模态图神经网络用于脑网络重组的预后建模,通过整合多种影像数据,生成可解释的生物标志物以预测认知下降风险。

Comments Fundamental methodological error invalidating results

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08077 2026-02-10 cs.LG cs.AI 57%

Multimodal normative modeling in Alzheimers Disease with introspective variational autoencoders

在阿尔茨海默病中使用反思变分自编码器的多模态规范建模

Sayantan Kumar, Peijie Qiu, Aristeidis Sotiras

机构 * Washington University in St Louis(华盛顿大学圣路易斯分校) Washington University in St Louis School of Medicine(华盛顿大学圣路易斯医学院)

专题命中 医疗多模态 :MRI(abstract);分类 cs.LG

AI总结 本文提出mmSIVAE,通过结合MOPOE聚合提升多模态数据的规范建模效果,提高参考分布保真度和多模态整合能力,用于阿尔茨海默病的偏差分析。

Comments Conference on Health, Inference, and Learning (CHIL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06184 2026-02-09 cs.CV cs.CL 57%

PhenoLIP: Integrating Phenotype Ontology Knowledge into Medical Vision-Language Pretraining

PhenoLIP:将表型本体知识整合到医学视觉-语言预训练中

Cheng Liang, Chaoyi Wu, Weike Zhao, Ya Zhang, Yanfeng Wang, Weidi Xie

机构 * School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 PhenoLIP通过整合表型本体知识提升医学视觉-语言模型的表型识别与跨模态检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07163 2026-02-06 cs.CV 57%

Test-time Adaptive Hierarchical Co-enhanced Denoising Network for Reliable Multimodal Classification

测试时自适应层次联合增强去噪网络用于可靠的多模态分类

Shu Shen, C. L. Philip Chen, Tong Zhang

机构 * The Guangdong Provincial Key Laboratory of Computational Intelligence and Cyberspace Information, the School of Computer Science and Engineering, South China University of Technology(广东省计算智能与网络信息重点实验室、计算机科学与工程学院、华南理工大学) The Pazhou Laboratory(琶洲实验室)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 本文提出TAHCD网络,通过自适应稳定子空间对齐和样本自适应置信度对齐,有效去除多模态噪声,提升多模态分类的鲁棒性和泛化能力。

Comments 14 pages,9 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03910 2026-02-05 eess.IV 57%

CONRep: Uncertainty-Aware Vision-Language Report Drafting Using Conformal Prediction

CONRep:基于置信预测的不确定性感知视觉-语言报告起草

Danial Elyassirad, Benyamin Gheiji, Mahsa Vatanparast, Amir Mahmoud Ahmadzadeh, Seyed Amir Asef Agah, Mana Moassefi, Meysam Tavakoli, Shahriar Faghani

专题命中 医疗多模态 :radiology(abstract);分类 eess.IV

AI总结 CONRep通过置信预测为视觉语言模型生成的放射学报告提供不确定性量化,提升自动化报告系统的透明性和临床实用性。

Comments 17 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20504 2026-02-05 cs.CV 57%

UniVRSE: Unified Vision-conditioned Response Semantic Entropy for Hallucination Detection in Medical Vision-Language Models

UniVRSE: 统一的视觉条件响应语义熵用于医学视觉语言模型中的幻觉检测

Zehui Liao, Shishuai Hu, Ke Zou, Mengyuan Jin, Yanning Zhang, Huazhu Fu, Liangli Zhen, Yong Xia

机构 * National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology, School of Computer Science and Engineering, Northwestern Polytechnical University(集成航空航天地面海洋大数据应用技术国家工程实验室,计算机科学与工程学院,西北工业大学)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 UniVRSE通过统一的视觉条件响应语义熵框架,有效检测医学视觉语言模型中的幻觉,提升临床应用的可靠性。

Comments Under Review. 12 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21479 2026-01-30 cs.CV 57%

Hypernetwork-Based Adaptive Aggregation for Multimodal Multiple-Instance Learning in Predicting Coronary Calcium Debulking

基于超网络的自适应聚合用于多模态多实例学习在预测冠状动脉钙化去除中的应用

Kaito Shiku, Ichika Seo, Tetsuya Matoba, Rissei Hino, Yasuhiro Nakano, Ryoma Bise

机构 * Department of Advanced Information Technology, Kyushu University(九州大学先进信息技术系) Department of Cardiovascular Medicine, Kyushu University(九州大学心血管医学系)

专题命中 医疗多模态 :CT(abstract);分类 cs.CV

AI总结 本文提出HyperAdAgFormer,通过超网络自适应聚合策略,用于多模态多实例学习预测冠状动脉钙化去除的必要性。

Comments Accepted to ISBI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18240 2026-01-27 cs.CV 57%

V-Loop: Visual Logical Loop Verification for Hallucination Detection in Medical Visual Question Answering

V-Loop:用于医学视觉问答中幻觉检测的视觉逻辑循环验证

Mengyuan Jin, Zehui Liao, Yong Xia

机构 * Northwestern Polytechnical University(西北工业大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 V-Loop通过双向推理和视觉逻辑循环验证,提升医学视觉问答中幻觉检测的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12346 2026-01-21 cs.CV 57%

MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents

MMDeepResearch-Bench: 一个多模态深度研究代理的基准

Peizhou Huang, Zixuan Zhong, Zhongwei Wan, Donghao Zhou, Samiul Alam, Xin Wang, Zexin Li, Zhihao Dou, Li Zhu, Jing Xiong, Chaofan Tao, Yan Xu, Dimitrios Dimitriadis, Tuo Zhang, Mi Zhang

机构 * OSU(俄亥俄州立大学) Amazon(亚马逊公司) UMich(密歇根大学) UCL(伦敦大学学院) CUHK(香港中文大学) UCR(加州大学尔湾分校) CWRU(克里夫兰医学中心) HKU(香港大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 MMDeepResearch-Bench提出一个多模态深度研究代理的基准,强调报告式合成与引用证据的结合,揭示多模态完整性对深度研究代理的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03482 2026-01-08 cs.AI cs.LG stat.AP 57%

Personalization of Large Foundation Models for Health Interventions

为健康干预个性化设计大型基础模型

Stefan Konigorski, Johannes E. Vedder, Babajide Alamu Owoyele, İbrahim Özkan

专题命中 医疗多模态 :healthcare AI(abstract);分类 cs.LG

AI总结 本文提出混合框架,结合LFMs和N-of-1试验,以实现个性化医疗中的因果推断与责任AI整合。

Comments Accepted to the AAAI 2026 Workshop on Personalization in the Era of Large Foundation Models (PerFM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00519 2026-01-05 cs.LG 57%

A Sparse-Attention Deep Learning Model Integrating Heterogeneous Multimodal Features for Parkinson's Disease Severity Profiling

一种整合异构多模态特征的稀疏注意力深度学习模型用于帕金森病严重程度评估

Dristi Datta, Tanmoy Debnath, Minh Chau, Manoranjan Paul, Gourab Adhikary, Md Geaur Rahman

机构 * School of Computing, Mathematics, and Engineering, Charles Sturt University(计算、数学与工程学院,查尔斯·斯特劳特大学) School of Dentistry and Medical Sciences, Charles Sturt University(牙科学院与医学科学学院,查尔斯·斯特劳特大学) AI and Cyber Futures Centre (AICF)(人工智能与未来技术中心(AICF)) Hawkins Clinic General Medical Practice(霍金斯诊所普通医疗实践)

专题命中 医疗多模态 :MRI(abstract);分类 cs.LG

AI总结 本文提出了一种整合异构多模态特征的稀疏注意力深度学习模型,用于提高帕金森病严重程度评估的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23109 2025-12-30 cs.LG cs.AI stat.ML 57%

How Much Data Is Enough? Uniform Convergence Bounds for Generative & Vision-Language Models under Low-Dimensional Structure

需要多少数据?在低维结构下生成式与视觉-语言模型的统一收敛界限

Paul M. Thompson

机构 * Stevens Institute for Neuroimaging and Informatics, University of Southern California(神经影像与信息学研究所,南加州大学)

专题命中 医疗多模态 :biomedical(abstract);分类 cs.LG

AI总结 研究在低维结构下生成式和视觉-语言模型的统一收敛界限,探讨数据量与模型校准之间的关系。

Comments 13 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21897 2025-12-29 cs.LG cs.AI 57%

MMCTOP: A Multimodal Textualization and Mixture-of-Experts Framework for Clinical Trial Outcome Prediction

MMCTOP: 一种用于临床试验结果预测的多模态文本化与专家混合框架

Carolina Aparício, Qi Shi, Bo Wen, Tesfaye Yadete, Qiwei Han

机构 * Nova School of Business and Economics(诺瓦商学院) Hogarthian Technologies(霍加斯技术) School of Medicine(医学院) Oregon Health & Science University(俄勒冈健康与科学大学) Cleveland Clinic(克利夫兰诊所) IBM Research(IBM研究院)

专题命中 医疗多模态 :biomedical(abstract);分类 cs.LG

AI总结 MMCTOP通过多模态文本化与专家混合框架,提升临床试验结果预测的精度与稳定性。

Comments 15 pages, 3 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21508 2025-12-29 cs.CV 57%

Fixed-Budget Parameter-Efficient Training with Frozen Encoders Improves Multimodal Chest X-Ray Classification

固定预算参数高效训练结合冻结编码器提升多模态胸片分类

Md Ashik Khan, Md Nahid Siddique

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Kharagpur, India(计算机科学与工程系,印度理工学院Kharagpur分校) Knight Foundation School of Computing and Information Sciences, Florida International University, Florida, USA(骑士基金会计算与信息科学学院,佛罗里达国际大学)

专题命中 医疗多模态 :pathology(abstract);分类 cs.CV

AI总结 本研究通过冻结编码器的参数高效训练策略,在降低计算成本的同时提升了多模态胸片分类的性能。

Comments Accepted at the 2025 28th International Conference on Computer and Information Technology (ICCIT). 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18986 2025-12-23 cs.LG cs.AI 57%

R-GenIMA: Integrating Neuroimaging and Genetics with Interpretable Multimodal AI for Alzheimer's Disease Progression

R-GenIMA:整合神经影像与基因组学的可解释多模态AI用于阿尔茨海默病进展

Kun Zhao, Siyuan Dai, Yingying Zhang, Guodong Liu, Pengfei Gu, Chenghua Lin, Paul M. Thompson, Alex Leow, Heng Huang, Lifang He, Liang Zhan, Haoteng Tang

机构 * Eli and Lilly company(艾利和利公司)

专题命中 医疗多模态 :MRI(abstract);分类 cs.LG

AI总结 R-GenIMA通过整合神经影像与基因组学,利用可解释的多模态AI方法,实现了对阿尔茨海默病进展的精准预测与机制揭示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17121 2025-12-22 cs.LG 57%

The Effect of Negation on CLIP in Medical Imaging: Limitations of Contrastive Language-Image Pretraining

否定对CLIP在医学影像中的影响:对比语言-图像预训练的局限性

Jasmine Vu, Shivanand Sheshappanavar

机构 * Santa Clara University(圣克拉拉大学) University of Wyoming(怀俄明大学)

专题命中 医疗多模态 :medical AI(abstract);分类 cs.LG

AI总结 研究探讨CLIP在医学影像中处理否定短语的局限性,并通过调优方法提升其检索准确性与可靠性。

Comments 10 pages, 7 figures, submitted to WACV Pixels to Patients Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13297 2025-12-16 cs.AI cs.LG 57%

MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data

MedInsightBench: 通过多步骤洞察发现评估医疗分析代理在多模态医疗数据中的能力

Zhenghao Zhu, Chuxue Cao, Sirui Han, Yuanfeng Song, Xing Chen, Caleb Chen Cao, Yike Guo

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) ByteDance(字节跳动)

专题命中 医疗多模态 :medical image(abstract);分类 cs.LG

AI总结 MedInsightBench通过多步骤洞察发现评估医疗分析代理在多模态医疗数据中的能力,提出MedInsightAgent框架提升医疗数据洞察发现性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13072 2025-12-16 cs.CV 57%

Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models

锻造动态记忆:基于检索的持续学习用于通用医学基础模型

Zizhi Chen, Yizhen Gao, Minghao Han, Yizhou Liu, Zhaoyu Chen, Dingkang Yang, Lihua Zhang

机构 * College of Intelligent Robotics and Advanced Manufacturing(智能机器人与先进制造学院) Fudan University(复旦大学) Fysics Intelligence Technologies Co., Ltd. (Fysics AI)(Fysics智能技术有限公司(Fysics AI)) School of Computer Science and Engineering(计算机科学与工程学院) Central South University(中南大学)

专题命中 医疗多模态 :biomedical(abstract);分类 cs.CV

AI总结 本文提出基于检索的持续学习方法,通过动态知识蒸馏和RAG技术,解决多模态医学模型在领域迁移和细粒度特征保留中的核心难题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11558 2025-12-15 cs.CV cs.AI cs.CL 57%

DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry

DentalGPT: 促进牙科多模态复杂推理的激励机制

Zhenyang Cai, Jiaming Zhang, Junjie Zhao, Ziyi Zeng, Yanchao Li, Jingyi Liang, Junying Chen, Yunjin Yang, Jiajun You, Shuzhi Deng, Tongfei Wang, Wanting Chen, Chunxiu Hao, Ruiqi Xie, Zhenwei Wen, Xiangyi Feng, Zou Ting, Jin Zou Lin, Jianquan Li, Guangjun Yu, Liangyi Chen, Junwen Wang, Shan Jiang, Benyou Wang

机构 * Shenzhen Stomatology Hospital (Pingshan) of Southern Medical University(南方医科大学深圳口腔医院(平山)) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) State Key Laboratory of Membrane Biology, Beijing Key Laboratory of Cardiometabolic Molecular Medicine, Institute of Molecular Medicine, National Biomedical Imaging Center, School of Future Technology, Peking University(北京大学膜生物学国家重点实验室、北京心代谢分子医学重点实验室、分子医学研究院、国家生物医学成像中心、未来技术学院) Freedom AI Division of Applied Oral Sciences & Community Dental Care Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院应用口腔科学与社区牙科护理系) Beijing Institute of Collaborative Innovation(北京协同创新研究院) National Health Data Institute, Shenzhen(深圳国家健康数据研究院) Shenzhen Loop Area Institute(深圳河套学院) Shenzhen Institute of Big Data(深圳大数据研究院)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 DentalGPT通过高质量数据和强化学习提升牙科多模态推理能力,实现优于现有模型的诊断性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17982 2025-12-15 cs.CV 57%

Few-Shot Learning from Gigapixel Images via Hierarchical Vision-Language Alignment and Modeling

通过层次化视觉-语言对齐和建模实现高像素图像的少样本学习

Bryan Wong, Jong Woo Kim, Huazhu Fu, Mun Yong Yi

机构 * KAIST(韩国科学技术院) IHPC, A*STAR(A*STAR研究院)

专题命中 医疗多模态 :pathology(abstract);分类 cs.CV

AI总结 HiVE-MIL通过层次化视觉-语言对齐和建模,提升高像素图像的少样本学习性能,实现4.1%的宏F1提升。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10750 2025-12-12 cs.CV 57%

LDP: Parameter-Efficient Fine-Tuning of Multimodal LLM for Medical Report Generation

LDP:多模态大语言模型在医疗报告生成中的参数高效微调

Tianyu Zhou, Junyi Tang, Zehui Li, Dahong Qian, Suncheng Xiang

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 LDP通过多模态大语言模型和参数高效微调技术,提升结肠镜息肉诊断报告的准确性和效率,显著降低训练成本并获得专家评分。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20549 2025-12-09 cs.LG cs.AI 57%

MedGR$^2$: Breaking the Data Barrier for Medical Reasoning via Generative Reward Learning

MedGR$^2$: 通过生成奖励学习突破医学推理的数据壁垒

Weihai Zhi, Jiayan Guo, Shangyang Li

专题命中 医疗多模态 :medical AI(abstract);分类 cs.LG

AI总结 MedGR$^2$通过生成奖励学习解决医学推理中的数据稀缺问题,实现高效训练和泛化,优于现有方法。

Comments 8 pages, 5 figures

Journal ref AAAI'2026 Main Technical Track

详情

展开后加载摘要…

URL PDF HTML 收藏