arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3450 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3450 篇

2505.13957 2025-05-21 cs.CR cs.CL 79%

Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation

Jiankun Zhang, Shenglai Zeng, Jie Ren, Tianqi Zheng, Hui Liu, Xianfeng Tang, Hui Liu, Yi Chang

机构 * Michigan State University(密歇根州立大学) Jilin University(吉林大学)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13520 2025-05-21 cs.IR cs.AI 79%

Beyond Retrieval: Joint Supervision and Multimodal Document Ranking for Textbook Question Answering

Hessa Alawwad, Usman Naseem, Areej Alhothali, Ali Alkhathlan, Amani Jamal

机构 * Faculty of Computing and Information Technology, King Abdulaziz University(计算机与信息科技学院,国王阿卜杜勒阿齐兹大学) College of Computer and Information Science, Imam Mohammad Ibn Saud Islamic University (IMSIU)(计算机与信息科学学院,伊玛目穆罕默德·本·萨乌德伊斯兰大学) School of Computing, Macquarie University(计算学院,麦考瑞大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments 14 pages, 16 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13306 2025-05-20 cs.CV cs.IR 79%

GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval

Chengsong Sun, Weiping Li, Xiang Li, Yuankun Liu, Lianlei Shan

机构 * School of Software and Microelectronics, Peking University(软件与微电子学院,北京大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11815 2025-05-20 cs.CV 79%

UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings

Jiajun Qin, Yuan Pu, Zhuolun He, Seunggeun Kim, David Z. Pan, Bei Yu

机构 * The Chinese University of Hong Kong, China(香港中文大学) ChatEDA Tech(ChatEDA科技) University of Texas at Austin, USA(德克萨斯大学奥斯汀分校)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04960 2025-05-09 cs.IR cs.MM 79%

Learning Item Representations Directly from Multimodal Features for Effective Recommendation

Xin Zhou, Xiaoxiong Zhang, Dusit Niyato, Zhiqi Shen

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

Comments Code: https://github.com/enoche/LIRDRec

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18674 2025-05-06 cs.CV cs.LG 79%

Active Data Curation Effectively Distills Large-Scale Multimodal Models

Vishaal Udandarao, Nikhil Parthasarathy, Muhammad Ferjad Naeem, Talfan Evans, Samuel Albanie, Federico Tombari, Yongqin Xian, Alessio Tonioni, Olivier J. Hénaff

机构 * Google(谷歌) Google DeepMind(谷歌DeepMind) Tübingen AI Center, University of Tübingen(图宾根人工智能中心,图宾根大学) University of Cambridge(剑桥大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21028 2025-05-01 cs.CR cs.AI cs.LG 79%

Semantic-Aware Contrastive Fine-Tuning: Boosting Multimodal Malware Classification with Discriminative Embeddings

Ivan Montoya Sanchez, Shaswata Mitra, Aritran Piplai, Sudip Mittal

机构 * dept. name of organization (of Aff.)(机构部门名称)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments 8 pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08535 2025-04-29 cs.IR cs.CV cs.LG 79%

Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking

Tianyu Zhu, Myong Chol Jung, Jesse Clark

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Journal ref The ACM Web Conference 2025 (WWW2025) Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10916 2025-04-24 physics.med-ph cs.CV 79%

Embedding Radiomics into Vision Transformers for Multimodal Medical Image Classification

Zhenyu Yang, Haiming Zhu, Rihui Zhang, Haipeng Zhang, Jianliang Wang, Chunhao Wang, Minbin Chen, Fang-Fang Yin

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 27 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13209 2025-04-21 cs.CR cs.AI 79%

On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks

Ting Bi, Chenghang Ye, Zheyu Yang, Ziyi Zhou, Cui Tang, Jun Zhang, Zui Tao, Kailong Wang, Liting Zhou, Yang Yang, Tianlong Yu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07718 2025-04-11 cs.CV 79%

Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval

Zehong Ma, Hao Chen, Wei Zeng, Limin Su, Shiliang Zhang

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments TMM25

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10702 2025-04-11 cs.MM 79%

Retrieval Augmented Verification for Zero-Shot Detection of Multimodal Disinformation

Arka Ujjal Dey, Artemis Llabrés, Ernest Valveny, Dimosthenis Karatzas

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17408 2025-04-11 cs.CL 79%

P-Transformer: A Prompt-based Multimodal Transformer Architecture For Medical Tabular Data

Yucheng Ruan, Xiang Lan, Daniel J. Tan, Hairil Rizal Abdullah, Mengling Feng

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04303 2025-04-08 cs.CL 79%

Graph-Based Multimodal Contrastive Learning for Chart Question Answering

Yue Dai, Soyeon Caren Han, Wei Liu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07202 2025-04-08 cs.AI 79%

A Zero-shot Learning Method Based on Large Language Models for Multi-modal Knowledge Graph Embedding

Bingchen Liu, Jingchen Li, Yuanyuan Fang, Xin Li

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00338 2025-04-02 cs.LG cs.AI cs.MA cs.SI 79%

Agentic Multimodal AI for Hyperpersonalized B2B and B2C Advertising in Competitive Markets: An AI-Driven Competitive Advertising Framework

Sakhinana Sagar Srinivas, Akash Das, Shivam Gupta, Venkataramana Runkana

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15836 2025-03-27 cs.CR cs.AI 79%

Intelligent Code Embedding Framework for High-Precision Ransomware Detection via Multimodal Execution Path Analysis

Levi Gareth, Maximilian Fairbrother, Peregrine Blackwood, Lucasta Underhill, Benedict Ruthermore

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17408 2025-03-25 cs.LG cs.AI 79%

Leveraging OpenFlamingo for Multimodal Embedding Analysis of C2C Car Parts Data

Maisha Binte Rashid, Pablo Rivas

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments The 26th International Conference on Artificial Intelligence (ICAI'24: July 22-25, 2024; Las Vegas, USA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12485 2025-03-24 cs.CV 79%

Cross-Modal Consistency Learning for Sign Language Recognition

Kepeng Wu, Zecheng Li, Hezhen Hu, Wengang Zhou, Houqiang Li

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13862 2025-03-19 cs.CV cs.LG 79%

HySurvPred: Multimodal Hyperbolic Embedding with Angle-Aware Hierarchical Contrastive Learning and Uncertainty Constraints for Survival Prediction

Jiaqi Yang, Wenting Chen, Xiaohan Xing, Sean He, Xiaoling Luo, Xinheng Lyu, Linlin Shen, Guoping Qiu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments submitted to IJCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10526 2025-03-14 cs.CV 79%

NeighborRetr: Balancing Hub Centrality in Cross-Modal Retrieval

Zengrong Lin, Zheng Wang, Tianwen Qian, Pan Mu, Sixian Chan, Cong Bai

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at CVPR 2025, 18 pages, 7 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19751 2025-03-03 cs.CV 79%

Lightweight Contrastive Distilled Hashing for Online Cross-modal Retrieval

Jiaxing Li, Lin Jiang, Zeqi Ma, Kaihang Jiang, Xiaozhao Fang, Jie Wen

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11742 2025-03-03 cs.CV 79%

Range and Bird's Eye View Fused Cross-Modal Visual Place Recognition

Jianyi Peng, Fan Lu, Bin Li, Yuan Huang, Sanqing Qu, Guang Chen

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17470 2025-02-28 eess.SP cs.AI 79%

MC2SleepNet: Multi-modal Cross-masking with Contrastive Learning for Sleep Stage Classification

Younghoon Na, Hyun Keun Ahn, Hyun-Kyung Lee, Yoongeol Lee, Seung Hun Oh, Hongkwon Kim, Jeong-Gun Lee

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19128 2025-02-27 cs.CV 79%

SCA3D: Enhancing Cross-modal 3D Retrieval via 3D Shape and Caption Paired Data Augmentation

Junlong Ren, Hao Wu, Hui Xiong, Hao Wang

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13277 2025-02-27 cs.LG cs.AI 79%

HyperGCL: Multi-Modal Graph Contrastive Learning via Learnable Hypergraph Views

Khaled Mohammed Saifuddin, Shihao Ji, Esra Akbas

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16451 2025-02-25 cs.CL 79%

Contrastive Learning of English Language and Crystal Graphs for Multimodal Representation of Materials Knowledge

Yang Jeong Park, Mayank Kumaran, Chia-Wei Hsu, Elsa Olivetti, Ju Li

专题命中 跨模态检索 :multimodal(title);cross-modal(abstract);分类 cs.CL

Comments 24 pages, 14 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13732 2025-02-25 cs.CV 79%

Modeling Multi-modal Cross-interaction for Multi-label Few-shot Image Classification Based on Local Feature Selection

Kun Yan, Zied Bouraoui, Fangyun Wei, Chang Xu, Ping Wang, Shoaib Jameel, Steven Schockaert

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments In Transactions on Multimedia Computing Communications and Applications. arXiv admin note: text overlap with arXiv:2112.01037

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04616 2025-02-24 cs.CL 79%

Can LLMs Improve Multimodal Fact-Checking by Asking Relevant Questions?

Alimohammad Beigi, Bohan Jiang, Dawei Li, Zhen Tan, Pouya Shaeri, Tharindu Kumarage, Amrita Bhattacharjee, Huan Liu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13954 2025-02-20 cs.CL cs.LG 79%

Latent Distribution Decoupling: A Probabilistic Framework for Uncertainty-Aware Multimodal Emotion Recognition

Jingwang Huang, Jiang Zhong, Qin Lei, Jinpeng Gao, Yuming Yang, Sirui Wang, Peiguang Li, Kaiwen Wei

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏