arXivDaily arXiv每日学术速递 周一至周五更新

科学与医疗

医学 AI

医学智能、临床 AI、医学影像、病理、诊断和医疗健康大模型。

共收录 966 信号源:cs.CV, cs.LG, q-bio, eess.IV, eess.SP

1. 医疗多模态 966 篇

2604.14656 2026-04-17 cs.AI cs.CL cs.CV 57%

Rethinking Patient Education as Multi-turn Multi-modal Interaction

重新思考患者教育作为多轮多模态交互

Zonghai Yao, Zhipeng Tang, Chengtao Lin, Xiong Luo, Benlu Wang, Juncheng Huang, Chin Siang Ong, Hong Yu

机构 * VA Bedford Health Care(VA贝德福德医疗中心) UMass Amherst(马萨诸塞大学阿默斯特分校) UMass Lowell(马萨诸塞大学洛厄尔分校) Yale University(耶鲁大学) National University of Singapore(新加坡国立大学) Yale School of Medicine(耶鲁医学院)

专题命中 医疗多模态 :radiology(abstract);分类 cs.CV

AI总结 本文提出MedImageEdu基准,通过多轮多模态交互提升患者教育效果,评估咨询过程和最终响应质量,发现多模态模型在视觉 grounding、安全性和情绪互动方面存在不足。

Comments Equal contribution for the first two authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21652 2026-04-15 eess.IV cs.AI physics.med-ph 57%

Enabling Ultra-Fast Cardiovascular Imaging Across Heterogeneous Clinical Environments with A Generalist Foundation Model and Multimodal Database

通过通用基础模型和多模态数据库实现跨异构临床环境的超快速心血管成像

Zi Wang, Mingkai Huang, Zhang Shi, Hongjie Hu, Lan Lan, Hui Zhang, Yan Li, Xi Hu, Qing Lu, Zongming Zhu, Qiong Yao, Yuxiang Dai, Fanwen Wang, Yinzhe Wu, Jun Lyu, Qianqian Gao, Guangming Xu, Zhenxuan Zhang, Haosen Zhang, Qing Li, Guangming Wang, Tianxing He, Lizhen Lan, Siyue Li, Le Xue, Mengting Sun, Yuntong Lyu, Junpu Hu, Jiayu Zhu, Rizwan Ahmad, Zhengyu Bu, Xianling Qian, Guanke Cai, Ruiyu Cao, Weirui Cai, Chang Xu, Yuyang Ren, Feidan Yu, Siying Ma, Ziqiang Xu, Xinran Chen, Sha Hua, Daniel Kim, Yajing Zhang, Chen Ouyang, Wenjia Bai, Jing Qin, Yucheng Yang, Daniel Rueckert, He Wang, Qian Tao, Claudia Prieto, Michael Markl, Alistair Young, Lianming Wu, Shuo Wang, Chen Qin, Mengsu Zeng, Xihong Hu, Haibo Xu, Xiaobo Qu, Hao Li, Guang Yang, Chengyan Wang

专题命中 医疗多模态 :diagnosis(abstract);分类 eess.IV

AI总结 本文提出CardioMM模型,通过统一语义理解和物理约束数据一致性,实现跨不同扫描仪、协议和患者情况的稳健重建,验证了24倍加速仍能保持关键心脏表型和诊断图像质量。

Comments Github: https://github.com/wangziblake/CardioMM_MMCMR-427K

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09364 2026-04-14 cs.CV cs.CL 57%

Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts

仲裁失败,而非感知失明:视觉-语言模型如何解决视觉-语言冲突

Farhad Nooralahzadeh, Omid Rohanian, Yi Zhang, Jonathan Fürst, Kurt Stockinger

机构 * Institute of Computer Science, Zurich University of Applied Sciences(苏黎世应用科技大学计算机科学研究所) University of Oxford(牛津大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 研究探讨视觉-语言模型在视觉与语言冲突中的仲裁机制,发现编码与基础的脱节,通过多模态仲裁交叉分析揭示视觉属性在早期层可线性解码,最终层logit与基础结果相关性达0.847,表明需针对性干预提升视觉基础能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09841 2026-04-14 cs.CV cs.AI 57%

Is There Knowledge Left to Extract? Evidence of Fragility in Medically Fine-Tuned Vision-Language Models

还有可以提取的知识吗?医学微调视觉语言模型中的脆弱性证据

Oliver McLaughlin, Daniel Shubin, Carsten Eickhoff, Ritambhara Singh, William Rudman, Michal Golovanevsky

机构 * Brown University(布朗大学) University of Washington(华盛顿大学) University of Tübingen(蒂宾根大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 研究评估了四个医学微调的视觉语言模型在四个医学影像任务中的表现,发现任务难度增加时性能下降至接近随机水平,表明临床推理有限,医学微调无明显优势,模型对提示词敏感。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07814 2026-04-10 cs.CV 57%

AgriChain Visually Grounded Expert Verified Reasoning for Interpretable Agricultural Vision Language Models

AgriChain:基于视觉的专家验证推理用于可解释的农业视觉语言模型

Hazza Mahmood, Yongqiang Yu, Rao Anwer

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 本文提出AgriChain数据集,通过专家验证的推理链提升农业视觉语言模型的准确性和可解释性,实验表明其在植物病害诊断中优于多个基线模型。

Comments 9 pages

Journal ref LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04614 2026-04-09 cs.LG cs.AI 57%

A Clinical Point Cloud Paradigm for In-Hospital Mortality Prediction from Multi-Level Incomplete Multimodal EHRs

一种用于院内死亡预测的临床点云范式:从多级不完整多模态电子健康记录

Bohao Li, Tao Zou, Junchen Ye, Yan Gong, Bowen Du

机构 * Beihang University(北京航空航天大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.LG

AI总结 本文提出HealthPoint范式,通过统一的4D空间表示多级不完整多模态EHRs,引入低秩关系注意力机制和分层交互策略,实现灵活的事件级交互与细粒度自监督,提升模态恢复和未标记数据利用效率,实验表明其在风险预测中性能优异。

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18297 2026-04-08 cs.CV 57%

Image-to-Text for Medical Reports Using Adaptive Co-Attention and Triple-LSTM Module

利用自适应共注意和三LSTM模块的图像到文本生成医学报告

Yishen Liu

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 本文提出CA-TriNet模型,结合Transformer和多LSTM网络,通过自适应权重运算和三LSTM模块提升医学图像与文本生成的准确性与多样性。

Comments arXiv admin note: This submission has been withdrawn by arXiv administrators due to incorrect authorship. Author list truncated

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05748 2026-04-08 cs.CV 57%

SVC 2026: the Second Multimodal Deception Detection Challenge and the First Domain Generalized Remote Physiological Measurement Challenge

SVC 2026: 第二届多模态欺骗检测挑战赛及首个领域通用远程生理测量挑战

Dongliang Zhu, Zhiyi Niu, Bo Zhao, Jiajian Huang, Shuo Ye, Xun Lin, Hui Ma, Taorui Wang, Jiayu Zhang, Chunmei Zhu, Junzhe Cao, Yingjie Ma, Rencheng Song, Albert Clapés, Sergio Escalera, Dan Guo, Zitong Yu

机构 * Wuhan University(武汉大学) Great Bay University(大湾区大学) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学) Sun Yat-sen University(中山大学) Hefei University of Technology(合肥工业大学) University of Barcelona(巴塞罗那大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 本文提出SVC 2026挑战赛,旨在通过多模态欺骗检测和远程脉搏波测速估计任务,推动对细微视觉信号的鲁棒表示学习研究。

Comments Accepted by the SVC workshop @ CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05831 2026-04-08 cs.LG cs.AI 57%

HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding

HeartcareGPT:一种统一的多模态ECG套件,用于双信号-图像建模与理解

Yihan Xie, Sijing Li, Tianwei Lin, Zhuonan Wang, Chenglin Yang, Yu Zhong, Wenjie Yan, Wenqiao Zhang, Xiaogang Guo, Jun Xiao, Yueting Zhuang, Beng Chin Ooi

机构 * Zhejiang University(浙江大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.LG

AI总结 本文提出HeartcareGPT,通过Dual Stream Projection Alignment机制,实现ECG信号与图像的统一建模,提升多视角ECG理解能力,并建立医学多模态大语言模型向生理信号领域扩展的方法论基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04133 2026-04-07 cs.CV cs.AI 57%

Learning Robust Visual Features in Computed Tomography Enables Efficient Transfer Learning for Clinical Tasks

在计算机断层扫描中学习鲁棒的视觉特征以实现临床任务的高效迁移学习

Rubén Moreno-Aguado, Alba Magallón, Victor Moreno, Yingying Fang, Guang Yang

机构 * Bioengineering Department and Imperial-X, Imperial College London(帝国理工学院生物工程系与Imperial-X) Department of Computer Science, University of Manchester(曼彻斯特大学计算机科学系) Oncology Data Analytics Program, Catalan Institute of Oncology(加泰罗尼亚肿瘤研究所肿瘤数据分析项目) Colorectal Cancer Group, ONCOBELL Program, Institut d’Investigació Biomèdica de Bellvitge(贝尔维特奇生物医学研究所ONCOBELL项目结直肠癌研究组) Consortium for Biomedical Research in Epidemiology and Public Health(流行病学与公共卫生生物医学研究联盟) Department of Clinical Sciences, Faculty of Medicine and Health Sciences, Universitat de Barcelona(巴塞罗那大学医学与健康科学学院临床科学系) Institute of Complex Systems, University of Barcelona(巴塞罗那大学复杂系统研究所) National Heart and Lung Institute, Imperial College London(帝国理工学院国家心肺研究所) Cardiovascular Research Centre, Royal Brompton Hospital(皇家布朗普顿医院心血管研究中心) School of Biomedical Engineering & Imaging Sciences, King’s College London(伦敦国王学院生物医学工程与影像科学学院)

专题命中 医疗多模态 :CT(abstract);分类 cs.CV

AI总结 本文提出VoxelFM,一种基于自蒸馏的3D CT基础模型,通过学习语义丰富的特征,实现了在多种临床任务中无需微调即可高效迁移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17514 2026-04-07 cs.CV 57%

EI: Early Intervention for Multimodal Imaging based Disease Recognition

EI:基于多模态影像的疾病识别早期干预

Qijie Wei, Hailan Lin, Xirong Li

机构 * Renmin University of China(中国人民大学) Beijing Key Laboratory for Intelligent Diagnosis of Fundus Diseases and Drug-Device R&D and Translation(眼底疾病智能诊断与药械研发转化北京市重点实验室)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 本文提出EI框架,通过早期干预提升多模态影像疾病识别性能,引入MoR方法优化Vision Foundation Models适应,验证了在三个公开数据集上的有效性。

Comments Accepted to CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02748 2026-04-06 cs.CV 57%

Visual Instruction-Finetuned Language Model for Versatile Brain MR Image Tasks

面向多样化脑部MRI任务的视觉指令微调语言模型

Jonghun Kim, Sinyoung Ra, Hyunjin Park

机构 * Sungkyunkwan University(成均馆大学) Department of Electrical and Computer Engineering(电气与计算机工程系) Department of Artificial Intelligence(人工智能系)

专题命中 医疗多模态 :MRI(abstract);分类 cs.CV

AI总结 本文提出LLaBIT模型,通过视觉指令微调提升语言模型在脑部MRI任务中的表现,通过特征图重用和文本数据生成提升性能,验证了其在报告生成、视觉问答、图像分割和图像翻译中的有效性。

Comments ICPR 2026 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01310 2026-04-03 cs.CV 57%

Sparse Spectral LoRA: Routed Experts for Medical VLMs

稀疏光谱LoRA:用于医学VLMs的路由专家

Omid Nejati Manzari, Hojat Asgariandehkordi, Taha Koleilat, Yiming Xiao, Hassan Rivaz

机构 * Concordia University(康考迪亚大学)

专题命中 医疗多模态 :radiology(abstract);分类 cs.CV

AI总结 本文提出MedQwen,一种参数高效的医学VLM,结合了光谱路由的专家混合(MoE)与理论支持的缩放规则,通过非重叠SVD段初始化专家,减少序列遗忘并提升医疗图像任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26726 2026-03-31 cs.CV cs.AI 57%

A Multimodal Deep Learning Framework for Edema Classification Using HCT and Clinical Data

一种结合头颅CT和临床数据的多模态深度学习框架用于水肿分类

Aram Ansary Ogholbake, Hannah Choi, Spencer Brandenburg, Alyssa Antuna, Zahraa Al-Sharshahi, Makayla Cox, Haseeb Ahmed, Jacqueline Frank, Nathan Millson, Luke Bauerle, Jessica Lee, David Dornbos, Qiang Cheng

机构 * University of Kentucky(肯塔基大学) Department of Computer Science, University of Kentucky(肯塔基大学计算机科学系) Department of Neurological Surgery, University of Kentucky(肯塔基大学神经外科学系) University of Kentucky College of Medicine(肯塔基大学医学院) Department of Neurology, University of Kentucky(肯塔基大学神经学系)

专题命中 医疗多模态 :CT(abstract);分类 cs.CV

AI总结 本文提出AttentionMixer框架,通过融合头颅CT和临床数据实现脑水肿检测,采用自监督Vision Transformer Autoencoder编码CT数据,并利用交叉注意力模块整合临床信息,最终通过MLP-Mixer提升分类性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26008 2026-03-30 cs.CV cs.AI 57%

FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants

FairLLaVA: 大规模视觉-语言助手中的公平性感知参数高效微调

Mahesh Bhosale, Abdul Wasi, Shantam Srivastava, Shifa Latif, Tianyu Luan, Mingchen Gao, David Doermann, Xuan Gong

机构 * University at Buffalo(布法罗大学) University of Kashmir(克什米尔大学) Accenture(埃森哲) Harvard Medical School(哈佛医学院)

专题命中 医疗多模态 :radiology(abstract);分类 cs.CV

AI总结 FairLLaVA通过减少目标属性间的互信息,实现视觉指令微调中的公平性改进,提升医疗影像生成的公平性和自然语言生成质量。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23953 2026-03-27 cs.CV cs.ET 57%

VOLMO: Versatile and Open Large Models for Ophthalmology

VOLMO:面向眼科学的多功能和开放型大模型

Zhenyue Qin, Younjoon Chung, Elijah Lee, Wanyue Feng, Xuguang Ai, Serina Applebaum, Minjie Zou, Yang Liu, Pan Xiao, Mac Singer, Amisha Dave, Aidan Gilson, Tiarnan D. L. Keenan, Emily Y. Chew, Zhiyong Lu, Yih-Chung Tham, Ron Adelman, Luciano V. Del Priore, Qingyu Chen

机构 * Department of Biomedical Informatics & Data Science, Yale University(耶鲁大学生物医学信息学与数据科学系) Ray and Stephanie Lane Computational Biology Department, Carnegie Mellon University(卡内基梅隆大学雷和斯蒂芬妮·兰德计算生物学系) Yong Loo Lin School of Medicine, National University of Singapore(新加坡国立大学杨洛林医学院) Department of Radiology, Washington University in Saint Louis(圣路易斯华盛顿大学放射科) National Eye Institute, National Institutes of Health(国家卫生研究院眼科研究所) National Library of Medicine, National Institutes of Health(国家卫生研究院国家医学图书馆)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 VOLMO提出了一种通用框架,用于开发专门的眼科多模态大语言模型,通过预训练、微调和临床推理三个阶段,提升了眼科疾病筛查和诊断的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13119 2026-03-26 cs.CV cs.AI 57%

Geometry-Guided Camera Motion Understanding in VideoLLMs

基于几何的视频LLMs中相机运动理解

Haoan Feng, Sri Harsha Musunuri, Guan-Ming Su

机构 * University of Maryland, College Park(马里兰大学College Park分校) Dolby Laboratories Inc.(杜比实验室)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 本文提出通过基准测试、诊断和注入框架,解决视频LLMs中相机运动表示不足的问题,通过CameraMotionDataset和CameraMotionVQA基准,提升模型对相机运动的识别能力。

Comments 10 pages, 7 figures, supplementary included CVPR2026 PVUW

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21597 2026-03-25 cs.AI cs.CV 57%

Cerebra: A Multidisciplinary AI Board for Multimodal Dementia Characterization and Risk Assessment

Cerebra:一个多学科AI平台用于多模态痴呆症特征化和风险评估

Sheng Liu, Long Chen, Zeyun Zhao, Qinglin Gou, Qingyue Wei, Arjun Masurkar, Kevin M. Spiegler, Philip Kuball, Stefania C. Bray, Megan Bernath, Deanna R. Willis, Jiang Bian, Lei Xing, Eric Topol, Kyunghyun Cho, Yu Huang, Ruogu Fang, Narges Razavian, James Zou

机构 * Department of Biomedical Data Science, Stanford University(斯坦福大学生物医学数据科学系) Department of Radiation Oncology, Stanford University(斯坦福大学放射肿瘤科) Center for Data Science, New York University(纽约大学数据科学中心) J. Crayton Pruitt Family Department of Biomedical Engineering, University of Florida(佛罗里达大学杰·克雷顿·普瑞特家庭生物医学工程系) Department of Biomedical Engineering and Informatics, Indiana University Indianapolis(印第安纳大学印第安纳波利斯分校生物医学工程与信息学系) Department of Neurology, NYU Grossman School of Medicine(纽约大学格罗斯曼医学院神经科) Department of Neuroscience, NYU Grossman School of Medicine(纽约大学格罗斯曼医学院神经科学系) UF Health Family Medicine – Haile Plantation(佛罗里达大学健康家庭医学-海莱种植园) Department of Family Medicine, Indiana University School of Medicine(印第安纳大学医学院家庭医学系) Department of Biostatistics and Health Data Science, Indiana University School of Medicine(印第安纳大学医学院生物统计学与健康数据科学系) Scripps Research Translational Institute(斯克里普斯研究转化研究所) Courant Institute, New York University(纽约大学柯朗研究所) Center for Biomedical Informatics, Regenstrief Institute(再生斯蒂夫研究所生物医学信息中心) Center for Cognitive Aging and Memory, McKnight Brain Institute, University of Florida(佛罗里达大学麦金托神经研究所认知衰老与记忆中心) Population Health Department, NYU Langone Health(纽约大学兰戈恩健康人口健康部) Radiology Department, NYU Langone Health(纽约大学兰戈恩健康放射科)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 Cerebra通过多智能体协调处理电子病历、临床笔记和医学影像,提供可视化分析和对话界面,提升临床决策支持的可解释性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21700 2026-03-24 cs.CV 57%

PPGL-Swarm: Integrated Multimodal Risk Stratification and Hereditary Syndrome Detection in Pheochromocytoma and Paraganglioma

PPGL-Swarm:集成多模态风险分层与遗传综合征检测的pheochromocytoma和paraganglioma诊断系统

Zelin Liu, Xiangfu Yu, Jie Huang, Ge Wang, Yizhe Yuan, Zhenyu Yi, Jing Xie, Haotian Jiang, Lichi Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学) Ruijin Hospital(瑞金医院) ShanghaiTech University(上海科技大学)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 本文提出PPGL-Swarm系统,通过多模态证据整合实现全面诊断报告,包括自动化GAPP评分、基因风险警报及可追溯的推理路径,解决传统诊断方法的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12469 2026-03-24 cs.CV 57%

Unleashing Video Language Models for Fine-grained HRCT Report Generation

释放视频语言模型以生成细粒度HRCT报告

Yingying Fang, Huichi Zhou, KinHei Lee, Yijia Wang, Zhenxuan Zhang, Jiahao Huang, Guang Yang

机构 * Bioengineering Department, Imperial College London, London, UK School of Biomedical Engineering \& lmaging Sciences, King's College London, London,UK

专题命中 医疗多模态 :CT(abstract);分类 cs.CV

AI总结 本文提出AbSteering框架,通过异常中心方案和直接偏好优化目标,提升视频语言模型在HRCT报告生成中的精度与细粒度区分能力,优于现有领域特定CT基础模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16664 2026-03-18 cs.CV cs.AI 57%

Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation

Kestrel: 为降低LVLM幻觉而引入自反思

Jiawei Mao, Hardy Chen, Haoqin Tu, Yuhan Wang, Letian Zhang, Zeyu Zheng, Huaxiu Yao, Zirui Wang, Cihang Xie, Yuyin Zhou

机构 * UC Santa Cruz(加州大学圣克ruz分校) UC Berkeley(加州大学伯克利分校) UNC-Chapel Hill(北卡罗来纳大学教堂山分校) Apple(苹果公司)

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 Kestrel提出一种无需训练的框架,通过显式视觉 grounding 与证据验证自反思机制减少LVLM幻觉,实验显示在POPE和MME-Hallucination基准上性能提升,同时提供透明的验证轨迹。

Comments 16 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16372 2026-03-18 cs.CV 57%

InViC: Intent-aware Visual Cues for Medical Visual Question Answering

InViC:面向医疗视觉问答的意图感知视觉线索

Zhisong Wang, Ziyang Chen, Zanting Ye, Hongze Zhu, Yefeng Zheng, Yong Xia

机构 * National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology, School of Computer Science and Engineering, Northwestern Polytechnical University, Xi’an 710072, China(集成空天地海大数据应用技术国家工程实验室,计算机科学与工程学院,西北工业大学,西安710072,中国) Westlake University(西湖大学) Southern Medical University(南方医科大学)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 本文提出InViC框架,通过引入Cue Tokens Extraction模块和两阶段微调策略,提升医疗视觉问答中意图对齐的视觉证据利用,验证了瓶颈化训练在提高可信度上的有效性。

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15168 2026-03-17 cs.CV cs.AI 57%

Multimodal Connectome Fusion via Cross-Attention for Autism Spectrum Disorder Classification Using Graph Learning

通过交叉注意力进行多模态连接组融合用于自闭症谱系障碍分类的图学习

Ansar Rahman, Hassan Shojaee-Mend, Sepideh Hatamikia

机构 * Department of Medicine, Danube Private University (DPU)(丹ube私人大学医学系) Department of Medical Physics and Biomedical Engineering, Medical University of Vienna(维也纳医学大学医学物理与生物医学工程系) Department of Medical Informatics, Faculty of Medicine, Gonabad University of Medical Sciences(戈纳巴德医学院医学信息学系) Austrian Center for Medical Innovation and Technology (ACMIT)(奥地利医学创新与技术中心)

专题命中 医疗多模态 :MRI(abstract);分类 cs.CV

AI总结 本文提出一种多模态图学习框架,通过交叉注意力机制融合功能和结构影像数据,提升自闭症谱系障碍分类的准确率。

Comments 29 Pages; 5 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15100 2026-03-17 cs.CV 57%

Learning from Limited and Incomplete Data: A Multimodal Framework for Predicting Pathological Response in NSCLC

从有限和不完整数据中学习:一种多模态框架用于预测非小细胞肺癌的病理反应

Alice Natalina Caragliano, Giulia Farina, Fatih Aksu, Camillo Maria Caruso, Claudia Tacconi, Carlo Greco, Lorenzo Nibid, Edy Ippolito, Michele Fiore, Giuseppe Perrone, Sara Ramella, Paolo Soda, Valerio Guarrasi

机构 * Department of Medicine and Surgery(医学与外科系) Department of Diagnostics and Intervention, Radiation Physics, Biomedical Engineering(诊断与介入系,放射物理,生物医学工程) Umeå University, Umeå, Sweden(乌梅大学,瑞典)

专题命中 医疗多模态 :CT(abstract);分类 cs.CV

AI总结 本文提出一种多模态深度学习框架,整合基础模型的CT特征提取与缺失感知架构,以在有限数据和不完整临床资料下准确预测NSCLC的病理反应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21509 2026-03-17 cs.CV 57%

Curing Semantic Drift: A Dynamic Approach to Grounding Generation in Large Vision-Language Models

消除语义漂移:一种动态方法用于在大视觉-语言模型中进行生成基础

Jiahe Chen, Jiaying He, Qiyuan Chen, Qian Shao, Jiahe Ying, Hongxia Xu, Jintai Chen, Jianwei Zheng, Jian Wu

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 本文提出DLC方法,通过动态校准logits来减少大视觉-语言模型中的语义漂移,提升生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.14084 2026-03-13 cs.AI cs.LG 57%

Multimodal Explainability via Latent Shift applied to COVID-19 stratification

多模态可解释性通过潜在位移应用于COVID-19分层

Valerio Guarrasi, Lorenzo Tronchin, Domenico Albano, Eliodoro Faiella, Deborah Fazzini, Domiziana Santucci, Paolo Soda

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.LG

AI总结 本文提出一种多模态可解释方法,通过潜在位移技术在COVID-19分层中实现决策解释,提升模型可解释性与分类性能。

Journal ref Pattern Recognition 156 (2024) 110825

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09931 2026-03-11 cs.CV cs.AI 57%

Adaptive Clinical-Aware Latent Diffusion for Multimodal Brain Image Generation and Missing Modality Imputation

自适应临床感知潜在扩散用于多模态脑图像生成与缺失模态插值

Rong Zhou, Houliang Zhou, Yao Su, Brian Y. Chen, Yu Zhang, Lifang He, Alzheimer's Disease Neuroimaging Initiative

专题命中 医疗多模态 :diagnosis(abstract);分类 cs.CV

AI总结 ACADiff通过自适应临床感知扩散模型,实现多模态脑图像生成与缺失模态插值,提升阿尔茨海默病诊断的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09101 2026-03-11 cs.CV 57%

MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration

MedKCO:通过知识驱动的认知协调进行医学视觉-语言预训练

Chenran Zhang, Ruiqi Wu, Tao Zhou, Yi Zhou

机构 * School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Ministry of Education, China(教育部新一代人工智能技术及其交叉应用重点实验室) School of Computer Science and Engineering, Nanjing University of Science and Technology, China(南京理工大学计算机科学与工程学院)

专题命中 医疗多模态 :medical image(abstract);分类 cs.CV

AI总结 MedKCO通过知识驱动的认知协调优化医学视觉-语言预训练,提升模型在分布偏移下的泛化能力。

Comments CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06697 2026-03-10 cs.CV cs.AI 57%

Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs

通过目光思考:将眼动作为视觉推理监督用于医学视觉语言模型

Yiwei Li, Zihao Wu, Yanjun Lv, Hanqi Jiang, Weihang You, Zhengliang Liu, Dajiang Zhu, Xiang Li, Quanzheng Li, Tianming Liu, Lin Zhao

机构 * School of Computing, University of Georgia(佐治亚大学计算机学院) Department of Computer Science and Engineering, University of Texas, Arlington(德克萨斯大学阿灵顿分校计算机科学与工程系) Massachusetts General Hospital, Harvard Medical School(麻省总医院哈佛医学院) Department of Biomedical Engineering, New Jersey Institute of Technology(新泽西理工学院生物医学工程系)

专题命中 医疗多模态 :radiology(abstract);分类 cs.CV

AI总结 通过引入眼动标记,利用时间有序的注视轨迹引导医学VLM的视觉推理,提升放射学任务的性能与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06147 2026-03-09 cs.CV 57%

Longitudinal NSCLC Treatment Progression via Multimodal Generative Models

多模态生成模型用于非小细胞肺癌治疗进展的纵向预测

Massimiliano Mantegna, Elena Mulero Ayllón, Alice Natalina Caragliano, Francesco Di Feola, Claudia Tacconi, Michele Fiore, Edy Ippolito, Carlo Greco, Sara Ramella, Philippe C. Cattin, Paolo Soda, Matteo Tortora, Valerio Guarrasi

机构 * Unit of Artificial Intelligence and Computer Systems, Department of Engineering, Università Campus Bio-Medico di Roma, Italy(人工智能与计算机系统单位,工程系,罗马大学生物医学学院) Multi-Specialist Clinical Institute for Orthopaedic Trauma Care (COT), Messina, Italy(骨科创伤护理多学科临床研究所(COT),意大利Messina) Department of Diagnostics and Intervention, Radiation Physics, Biomedical Engineering, Umeå University, Sweden(诊断与介入系,放射物理,生物医学工程,乌梅大学,瑞典) Operative Research Unit of Radiation Oncology, Fondazione Policlinico Universitario Campus Bio-Medico, Rome, Italy(放射肿瘤手术研究单位,大学生物医学学院基金会,罗马,意大利) Research Unit of Radiation Oncology, Department of Medicine and Surgery, Università Campus Bio-Medico di Roma, Italy(放射肿瘤研究单位,医学与外科系,罗马大学生物医学学院,意大利) Department of Biomedical Engineering, University of Basel, Allschwil, Switzerland(生物医学工程系,巴塞尔大学,瑞士Allschwil) Department of Naval, Electrical, Electronics and Telecommunications Engineering, University of Genoa, Italy(海军、电气、电子与电信工程系,热那亚大学,意大利)

专题命中 医疗多模态 :CT(abstract);分类 cs.CV

AI总结 本研究提出虚拟治疗框架,利用多模态生成模型预测NSCLC治疗进展,验证扩散模型在生成稳定肿瘤演变轨迹方面的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏