arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3450 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3450 篇

2503.19474 2025-04-03 cs.CV cs.AI 81%

A-MESS: Anchor based Multimodal Embedding with Semantic Synchronization for Multimodal Intent Recognition

Yaomin Shen, Xiaojian Lin, Wei Fan

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12843 2025-03-27 cs.CV cs.AI 81%

Towards Scalable Foundation Model for Multi-modal and Hyperspectral Geospatial Data

Haozhe Si, Yuxuan Wan, Minh Do, Deepak Vasisht, Han Zhao, Hendrik F. Hamann

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19801 2025-03-26 cs.CV cs.AI 81%

SeLIP: Similarity Enhanced Contrastive Language Image Pretraining for Multi-modal Head MRI

Zhiyang Liu, Dong Yang, Minghao Zhang, Hanyu Sun, Hong Wu, Huiying Wang, Wen Shen, Chao Chai, Shuang Xia

专题命中 跨模态检索 :multi-modal(title);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13071 2025-03-25 cs.CL cs.IR cs.SD eess.AS 81%

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval

Mohammad Mahdi Abootorabi, Ehsaneddin Asgari

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、eess.AS

Comments accepted at ECIR 2025, 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12663 2025-03-18 cs.CV cs.CL cs.LG cs.RO 81%

Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding

Imran Kabir, Md Alimoor Reza, Syed Billah

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10937 2025-03-17 cs.CV cs.AI cs.LG 81%

ChatGPT Encounters Morphing Attack Detection: Zero-Shot MAD with Multi-Modal Large Language Models and General Vision Models

Haoyu Zhang, Raghavendra Ramachandra, Kiran Raja, Christoph Busch

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10166 2025-03-14 cs.IR cs.AI cs.MM 81%

ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning

Pengfei Luo, Jingbo Zhou, Tong Xu, Yuan Xia, Linli Xu, Enhong Chen

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments WWW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03107 2025-03-06 cs.CL cs.AI 81%

External Reliable Information-enhanced Multimodal Contrastive Learning for Fake News Detection

Biwei Cao, Qihang Wu, Jiuxin Cao, Bo Liu, Jie Gui

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments accepted by AAAI'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00036 2025-02-27 cs.CL cs.AI cs.LG 81%

EMERGE: Enhancing Multimodal Electronic Health Records Predictive Modeling with Retrieval-Augmented Generation

Yinghao Zhu, Changyu Ren, Zixiang Wang, Xiaochen Zheng, Shiyun Xie, Junlan Feng, Xi Zhu, Zhoujun Li, Liantao Ma, Chengwei Pan

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments CIKM 2024 Full Research Paper; arXiv admin note: text overlap with arXiv:2402.07016

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15040 2025-02-24 cs.CL cs.AI 81%

Reducing Hallucinations of Medical Multimodal Large Language Models with Visual Retrieval-Augmented Generation

Yun-Wei Chu, Kai Zhang, Christopher Malon, Martin Renqiang Min

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments GenAI4Health - AAAI '25

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11751 2025-02-18 cs.CV cs.AI 81%

Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning

Yuqi Pang, Bowen Yang, Haoqin Tu, Yun Cao, Zeyu Zhang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11442 2025-02-18 cs.IR cs.AI cs.CL cs.LG 81%

Multi-Turn Multi-Modal Question Clarification for Enhanced Conversational Understanding

Kimia Ramezan, Alireza Amiri Bavandpour, Yifei Yuan, Clemencia Siro, Mohammad Aliannejadi

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06720 2025-02-18 cs.CV cs.CL 81%

VP-MEL: Visual Prompts Guided Multimodal Entity Linking

Hongze Mi, Jinyuan Li, Xuying Zhang, Haoran Cheng, Jiahao Wang, Di Sun, Gang Pan

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02458 2025-02-05 cs.CL cs.CV 81%

SAISA: Towards Multimodal Large Language Models with Both Training and Inference Efficiency

Qianhao Yuan, Yanjiang Liu, Yaojie Lu, Hongyu Lin, Ben He, Xianpei Han, Le Sun

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02048 2025-02-05 cs.LG cs.CL cs.CV 81%

Efficient Domain Adaptation of Multimodal Embeddings using Constrastive Learning

Georgios Margaritis, Periklis Petridis, Dimitris J. Bertsimas

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14166 2025-01-27 cs.CV cs.AI 81%

Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation

Cong-Duy Nguyen, Xiaobao Wu, Thong Nguyen, Shuai Zhao, Khoi Le, Viet-Anh Nguyen, Feng Yichao, Anh Tuan Luu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13267 2025-01-27 cs.SD cs.CL eess.AS 81%

CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models

Shangda Wu, Yashan Wang, Ruibin Yuan, Zhancheng Guo, Xu Tan, Ge Zhang, Monan Zhou, Jing Chen, Xuefeng Mu, Yuejie Gao, Yuanliang Dong, Jiafeng Liu, Xiaobing Li, Feng Yu, Maosong Sun

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、eess.AS

Comments 17 pages, 10 figures, 4 tables, accepted by NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13297 2025-01-24 cs.CL cs.AI cs.IR cs.LG 81%

RAMQA: A Unified Framework for Retrieval-Augmented Multi-Modal Question Answering

Yang Bai, Christan Earl Grant, Daisy Zhe Wang

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by NAACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05030 2025-01-10 cs.AI cs.CL 81%

A General Retrieval-Augmented Generation Framework for Multimodal Case-Based Reasoning Applications

Ofir Marom

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16215 2024-12-24 cs.CV cs.AI cs.IR 81%

Zero-Shot Image Moderation in Google Ads with LLM-Assisted Textual Descriptions and Cross-modal Co-embeddings

Enming Luo, Wei Qiao, Katie Warren, Jingxiang Li, Eric Xiao, Krishna Viswanathan, Yuan Wang, Yintao Liu, Jimin Li, Ariel Fuxman

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14475 2024-12-20 cs.CV cs.CL 81%

MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

Junjie Zhou, Zheng Liu, Ze Liu, Shitao Xiao, Yueze Wang, Bo Zhao, Chen Jason Zhang, Defu Lian, Yongping Xiong

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11216 2024-12-20 cs.CV cs.AI cs.IR 81%

Distribution-Consistency-Guided Multi-modal Hashing

Jin-Yu Liu, Xian-Ling Mao, Tian-Yi Che, Rong-Cheng Tu

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11227 2024-12-20 eess.IV cs.AI cs.CV 81%

OCTCube-M: A 3D multimodal optical coherence tomography foundation model for retinal and systemic diseases with cross-cohort and cross-device validation

Zixuan Liu, Hanwen Xu, Addie Woicik, Linda G. Shapiro, Marian Blazes, Yue Wu, Verena Steffen, Catherine Cukras, Cecilia S. Lee, Miao Zhang, Aaron Y. Lee, Sheng Wang

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04594 2024-12-20 cs.CV cs.AI 81%

Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Qirui Jiao, Daoyuan Chen, Yilun Huang, Bolin Ding, Yaliang Li, Ying Shen

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 22 pages, 10 figures, 16 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10455 2024-12-17 cs.CV cs.AI cs.CG 81%

Geo-LLaVA: A Large Multi-Modal Model for Solving Geometry Math Problems with Meta In-Context Learning

Shihao Xu, Yiyang Luo, Wei Shi

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08276 2024-11-27 cs.CV cs.AI 81%

Direction-Oriented Visual-semantic Embedding Model for Remote Sensing Image-text Retrieval

Qing Ma, Jiancheng Pan, Cong Bai

专题命中 跨模态检索 :image-text(title,abstract);分类 cs.CV、cs.AI

Comments 14 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15403 2024-11-26 cs.CV cs.AI 81%

MMDS: A Multimodal Medical Diagnosis System Integrating Image Analysis and Knowledge-based Departmental Consultation

Yi Ren, HanZhi Zhang, Weibin Li, Jun Fu, Diandong Liu, Tianyi Zhang, Jie He, Licheng Jiao

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09747 2024-11-22 cs.CV cs.AI cs.DC cs.LG cs.RO 81%

t-READi: Transformer-Powered Robust and Efficient Multimodal Inference for Autonomous Driving

Pengfei Hu, Yuhang Qian, Tianyue Zheng, Ang Li, Zhe Chen, Yue Gao, Xiuzhen Cheng, Jun Luo

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 14 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09729 2024-11-13 cs.CV cs.AI 81%

MIRAGE: Multimodal Identification and Recognition of Annotations in Indian General Prescriptions

Tavish Mankash, V. S. Chaithanya Kota, Anish De, Praveen Prakash, Kshitij Jadhav

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 5 pages, 9 figures, 3 tables, submitted to ISBI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21943 2024-10-30 cs.CL cs.AI 81%

Beyond Text: Optimizing RAG with Multimodal Inputs for Industrial Applications

Monica Riedler, Stefan Langer

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏