arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-05-05 至 2026-05-05 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 6 篇

2605.01163 2026-05-05 cs.IR cs.LG 82%

Multimodal Data Curation Through Ranked Retrieval

通过排序检索进行多模态数据整理

Pratyush Muthukumar, Harshil Kotamreddy, Sarah Amiraslani, Tomo Kanazawa, Ramani Akkati, Shaan Jain, Andrew Mathau

机构 * NVIDIA

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文提出通过优化训练对和嵌入模型来提升多模态检索对齐,利用对称核子采样和专家嵌入引擎减少模态驱动的分离,有效缩小模态差距。

Comments ICLR DATA-FM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01393 2026-05-05 cs.CV 57%

Recall to Predict: Grounding Motion Forecasting in Interpretable Motion Bank

回忆以预测:在可解释的运动库中 grounded 运动预测

Abhishek Vivekanandan, Ahmed Abouelazm, J. Marius Zöllner

机构 * FZI Forschungszentrum Informatik(弗劳恩霍夫研究所信息技术研究中心) Karlsruhe Institute of Technology (KIT)(卡尔斯鲁厄理工学院)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出了一种端到端的可区分框架,通过对比学习构建全面的'运动库'来 grounding 运动预测,采用新颖的锚点检索层动态检索显式运动先验,结合 DETR 式解码器和 WTA 动态高斯混合模型,实现可解释的多模态预测。

Comments Sumitted for PeerReview

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01165 2026-05-05 cs.CV 57%

CEZSAR: A Contrastive Embedding Method for Zero-Shot Action Recognition

CEZSAR:一种用于零样本动作识别的对比嵌入方法

Valter Estevam, Rayson Laroca, Helio Pedrini, David Menotti

机构 * Federal Institute of Paraná, Irati, Brasil(巴西南里奥格兰德联邦理工学院) Federal University of Paraná, Curitiba, Brasil(巴西南里奥格兰德联邦大学) Pontifical Catholic University of Paraná, Curitiba, Brasil(巴西南里奥格兰德天主教大学) University of Campinas, Campinas, Brasil(坎皮纳斯大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种基于对比学习的零样本动作识别方法,通过联合嵌入空间对视频和句子进行编码,解决语义鸿沟和领域偏移问题,实验在UCF-101和Kinetics-400数据集上取得最佳效果。

Comments Accepted for presentation at the International Conference on Pattern Recognition (ICPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00925 2026-05-05 cs.LG cs.CV q-bio.QM 57%

Linking spatial biology and clinical histology via Haiku

通过俳句连接空间生物学与临床组织学

Yan Cui, Jacob S. Leiby, Wenhui Lei, Dokyoon Kim, Yanxiang Deng, Aaron T. Mayer, Zhenqin Wu, Alexandro E. Trevino, Zhi Huang

机构 * Department of Pathology and Laboratory Medicine, University of Pennsylvania(宾夕法尼亚大学病理学与实验室医学系) Department of Bioengineering, University of Pennsylvania(宾夕法尼亚大学生物工程系) Department of Biostatistics, Epidemiology & Informatics, University of Pennsylvania(宾夕法尼亚大学生物统计学、流行病学与信息学系) Enable Medicine, Menlo Park, CA, USA(Enable Medicine)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出Haiku模型,整合分子、形态和临床数据,通过三模态对比学习实现跨模态检索,提升分类和预测任务性能,并支持零样本生物标志物推断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00902 2026-05-05 cs.CV cs.IR 57%

Validation of Whole-Slide Foundation Models for Image Retrieval in TCGA Data

对TCGA数据中整张滑动图像检索的整张滑动图像基础模型进行验证

Tianhao Lei, Parsa Esmaeilkhani, Saghir Alfasly, Wataru Uegami, Judy C. Boughey, Matthew P. Goetz, Krishna R. Kalari, H. R. Tizhoosh

机构 * KIMIA Lab, Artificial Intelligence and Informatics, Mayo Clinic, Rochester, MN, USA(KIMIA实验室,人工智能与信息学,梅奥诊所,罗切斯特,明尼苏达州,美国) Department of Neurology, Northwestern University Feinberg School of Medicine, Chicago, IL, USA(神经病学系,北western大学费因伯格医学院,芝加哥,伊利诺伊州,美国) Department of Computer Science, Temple University, Philadelphia, PA, USA(计算机科学系,泰勒大学,费城,宾夕法尼亚州,美国) Department of Breast and Melanoma Surgical Oncology, Comprehensive Cancer Center, Mayo Clinic, Rochester, MN, USA(乳腺和黑色素瘤外科肿瘤学系,综合癌症中心,梅奥诊所,罗切斯特,明尼苏达州,美国) Department of Oncology, Comprehensive Cancer Center, Mayo Clinic, Rochester, MN, USA(肿瘤学系,综合癌症中心,梅奥诊所,罗切斯特,明尼苏达州,美国) Department of Quantitative Health Sciences, Mayo Clinic, Rochester, MN, USA(定量健康科学系,梅奥诊所,罗切斯特,明尼苏达州,美国)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 本研究验证了TCGA数据中整张滑动图像基础模型的检索性能,发现基于片段和监督聚合的方法在Top-1和Top-3准确率上表现相当,但整体性能受器官和诊断差异影响较大,形态学检索存在固有限制,需进一步改进特征表示和多模态框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12778 2026-05-05 cs.CL 57%

HeteroRAG: A Heterogeneous Retrieval-Augmented Generation Framework for Medical Vision Language Tasks

HeteroRAG:一种用于医疗视觉语言任务的异构检索增强生成框架

Zhe Chen, Yusheng Liao, Zhiyuan Zhu, Haolin Li, Hongcheng Liu, Yanfeng Wang, Yu Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

AI总结 本文提出HeteroRAG框架,通过异构知识源增强医疗大视觉语言模型,解决异构数据检索问题,提升事实准确性与可靠性。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏