arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-05-29 至 2026-05-29 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 9 篇

2605.29602 2026-05-29 cs.CV 85%

CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning

CogniVerse: 用认知反思与几何推理革新多模态检索增强生成

Xiang Fang, Wanlong Fang, Changshuo Wang

机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) Nanyang Technological University, Singapore(新加坡南洋理工大学) University College London(伦敦大学学院)

专题命中 跨模态检索 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 提出CogniVerse框架,通过认知反思模块、基于黎曼流形对齐的多模态检索模块和最优传输层次生成模块,解决多模态检索增强生成中的噪声检索、跨模态语义错位和生成不连贯问题。

Comments Accepted in CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30111 2026-05-29 cs.CV cs.AI 84%

xModel-KD: Cross-modal Knowledge Distillation for 3D Scene Perception using LiDAR

xModel-KD:基于LiDAR的3D场景感知跨模态知识蒸馏

Thenukan Pathmanathan, Kanchan Keisham, Thangarajah Akilan

机构 * Dept. of Computer Science Lakehead University Thunder Bay, Canada School of Computer Science Engg. \& Info. Systems Vellore Institute of Technology Tamil Nadu, India Dept. of Software Engg. Lakehead University Thunder Bay, Canada

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 提出跨模态知识蒸馏框架xModel-KD,通过对比学习对齐2D图像纹理与3D点云几何特征,在无额外标注下提升LiDAR点云分割性能。

Comments 3 figures, and 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29900 2026-05-29 cs.LG cs.IT math.IT 82%

OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment

OVA-IB:用于多模态对齐的一对多信息瓶颈

Tianchao Li, Shujian Yu, Xinrui Zu, Zhaolong Wei, Jeremy Gummeson, Jack C. P. Cheng, Robert Jenssen

机构 * Hong Kong University of Science and Technology(香港科学与技术大学) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) UiT – The Arctic University of Norway(挪威北极大学) University of Copenhagen(哥本哈根大学) Norwegian Computing Center(挪威计算中心) University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校)

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract)

AI总结 提出基于信息瓶颈的一对多对齐框架OVA-IB,通过充分性对比下界和最小性正则化实现任意数量模态的对齐,在分类、回归和跨模态检索任务中表现鲁棒。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29628 2026-05-29 cs.SD cs.AI cs.CL cs.LG eess.AS 82%

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings

COMET:音频-文本多模态对比嵌入中模态间隙的概念空间剖析

Yonggang Zhu, Liting Gao, Aidong Men, Wenwu Wang

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Centre for Vision, Speech, and Signal Processing (CVSSP), University of Surrey(Surrey 大学视觉、语音和信号处理中心)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 提出COMET框架,通过PLS-SVD分解揭示CLAP模型中模态间隙主要由少数共享概念轴贡献,并基于谱截断方法无训练地缓解间隙,实现零样本音频字幕接近全监督性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14113 2026-05-29 cs.CV cs.AI cs.LG cs.MA 81%

ProtoMedAgent: Multimodal Clinical Interpretability via Privacy-Aware Agentic Workflows

ProtoMedAgent: 通过隐私感知的智能体工作流实现多模态临床可解释性

Alvaro Lopez Pellicer, Plamen Angelov, Marwan Bukhari, Yi Li, Eduardo Soares, Jemma Kerns

机构 * School of Computing and Communications(计算与通信学校) Lancaster University(兰卡斯特大学) Lancaster Medical School(兰卡斯特医学院) PUC-Rio(里约热内卢联邦大学) Puc-Behring Institute for AI(人工智能皮克林研究所)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 提出ProtoMedAgent框架,通过神经符号瓶颈和反射性Scribe-Critic循环约束生成过程,解决原型网络在临床报告中的语义结构缺失和检索谄媚问题,并引入k-匿名和ℓ-多样性隐私门控。

Comments CVR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29606 2026-05-29 cs.AI cs.IR 79%

HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering

HiKEY: 面向开放域文档问答的分层多模态检索

Joongmin Shin, Gyuho Shim, Jeongbae Park, Jaehyung Seo, Heuiseok Lim

机构 * Human-inspired AI Research, Korea University(韩国大学人机智能研究部) Computer Science and Engineering, Konkuk University(韩国康科大学计算机科学与工程系) Department of Computer Science and Engineering, Korea University(韩国大学计算机科学与工程系)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 提出基于文档层次结构的分层多模态检索框架HiKEY,通过文档层次解析和粗到细的检索策略解决大规模工业语料中的路由失败和证据碎片化问题,在ODQA基准上检索召回率提升达12.9%,端到端QA性能提升达6.8%。

Comments Accepted to ACL2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29956 2026-05-29 cs.IR 78%

Uncertainty Quantification for Multimodal Retrieval Augmented Generation

多模态检索增强生成的不确定性量化

Simon Binz, Heydar Soudani, Faegheh Hasibi

专题命中 跨模态检索 :multimodal(title,abstract)

AI总结 提出 LeMUQ 方法,通过多模态和检索感知的概率信号建模不确定性,提升多模态 RAG 系统的可靠性,AUROC 平均提升 3.8%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29980 2026-05-29 cs.CV cs.AI cs.LG 62%

Genetically Aligned Patient Representations Improve Hematological Diagnosis

基因对齐的患者表示改善血液学诊断

Muhammed Furkan Dasdelen, Fatih Ozlugedik, Ilaria Looser, Rao Muhammad Umer, Christian Pohlkamp, Carsten Marr

机构 * Institute of AI for Health, Helmholtz Munich, Germany International School of Medicine, Istanbul Medipol University, T\"urkiye Munich Leukemia Laboratory, Germany Department of Medicine III, Ludwig-Maximilian-University Hospital, Germany Department of Physics, University of Munich, Germany Munich Center for Machine Learning (MCML), Germany DKTK, German Cancer Consortium, Germany

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 提出一种两阶段框架,通过自监督视觉预训练和监督对比学习对齐白细胞图像与染色体畸变及体细胞突变,提升血液学诊断性能。

Comments Accepted for publication at the 29th International Conference on Medical Image Computing and Computer Assisted Intervention - MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29926 2026-05-29 cs.LG 50%

A Triple-Modal Contrastive Learning Framework with Sequence, Graph, and 3D Features for Drug-Target Interaction Prediction

一种融合序列、图和3D特征的三模态对比学习框架用于药物-靶标相互作用预测

Le Xu, Xi Zhang, Dan Luo, Ting Wang, Xuan Lin

机构 * School of Computer Science, Xiangtan University, Xiangtan 411105, China(湘潭大学计算机科学学院)

专题命中 跨模态检索 :cross-modal(abstract)

AI总结 提出TriMod-DTI框架,通过融合药物和蛋白质的1D序列、2D图和3D结构,并采用三模态对比学习策略对齐潜在空间表示,从而提升药物-靶标相互作用预测性能。

Comments 12 pages, 5 figures, ISBRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏