arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3454 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3454 篇

2507.21179 2025-09-25 cs.LG cs.AI 74%

CANDLE: A Cross-Modal Agentic Knowledge Distillation Framework for Interpretable Sarcopenia Diagnosis

Yuqi Jin, Zhenhao Shuai, Zihan Hu, Weiteng Zhang, Weihao Xie, Jianwei Shuai, Xian Shen, Zhen Feng

专题命中 跨模态检索 :cross-modal(title);分类 cs.AI

Comments 11 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18672 2025-09-24 cs.HC cs.AI 74%

NaviSense: A Multimodal Assistive Mobile application for Object Retrieval by Persons with Visual Impairment

Ajay Narayanan Sridhar, Fuli Qiao, Nelson Daniel Troncoso Aldas, Yanpei Shi, Mehrdad Mahdavi, Laurent Itti, Vijaykrishnan Narayanan

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Independent Researcher(独立研究者) University of Southern California(南加州大学)

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05883 2025-09-09 cs.CR cs.AI 74%

Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs

Andrew Yeo, Daeseon Choi

机构 * Ranchview High School(拉文斯维尔高中) Soongsil University(松山大学)

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

Comments 8 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08781 2025-08-13 cs.CV 74%

SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)

Trong-Thuan Nguyen, Viet-Tham Huynh, Quang-Thuc Nguyen, Hoang-Phuc Nguyen, Long Le Bao, Thai Hoang Minh, Minh Nguyen Anh, Thang Nguyen Tien, Phat Nguyen Thuan, Huy Nguyen Phong, Bao Huynh Thai, Vinh-Tiep Nguyen, Duc-Vu Nguyen, Phu-Hoa Pham, Minh-Huy Le-Hoang, Nguyen-Khang Le, Minh-Chinh Nguyen, Minh-Quan Ho, Ngoc-Long Tran, Hien-Long Le-Hoang, Man-Khoi Tran, Anh-Duong Tran, Kim Nguyen, Quan Nguyen Hung, Dat Phan Thanh, Hoang Tran Van, Tien Huynh Viet, Nhan Nguyen Viet Thien, Dinh-Khoi Vo, Van-Loc Nguyen, Trung-Nghia Le, Tam V. Nguyen, Minh-Triet Tran

机构 * University of Science, VNU-HCM(越南胡志明市科学大学) University of Information Technology, VNU-HCM(越南胡志明市信息技术大学) Vietnam National University(越南国家大学) University of Dayton(戴维森大学)

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03729 2025-08-07 cs.LG cs.HC cs.MM 74%

Privileged Contrastive Pretraining for Multimodal Affect Modelling

Kosmas Pinitas, Konstantinos Makantasis, Georgios N. Yannakakis

机构 * Institute of Digital Games University of Malta(数字游戏研究所马耳他大学) Department of Artificial Intelligence University of Malta(人工智能系马耳他大学)

专题命中 跨模态检索 :multimodal(title);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05079 2025-07-08 cs.AI eess.SP 74%

Multimodal-to-Text Prompt Engineering in Large Language Models Using Feature Embeddings for GNSS Interference Characterization

Harshith Manjunath, Lucas Heublein, Tobias Feigl, Felix Ott

机构 * Fraunhofer Institute for Integrated Circuits IIS(弗劳恩霍夫集成电路研究所)

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

Journal ref IEEE Wireless Communications and Networking Conference (WCNC), March 2025, Milan, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.04612 2025-06-09 cs.CV 74%

Self-Supervised Generative-Contrastive Learning of Multi-Modal Euclidean Input for 3D Shape Latent Representations: A Dynamic Switching Approach

Chengzhi Wu, Julius Pfrommer, Mingyuan Zhou, Jürgen Beyerer

机构 * Institute for Anthropomatics and Robotics(人机化与机器人研究所) Karlsruhe Institute of Technology(卡尔斯鲁厄理工大学) Fraunhofer Institute of Optronics, System Technologies and Image Exploitation IOSB(弗劳恩霍夫光学、系统技术与图像利用研究所)

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14468 2025-04-22 cs.CL cs.LG eess.SP q-bio.NC 74%

sEEG-based Encoding for Sentence Retrieval: A Contrastive Learning Approach to Brain-Language Alignment

Yijun Liu

专题命中 跨模态检索 :multimodal(abstract,comments);multimodal foundation model(abstract,comments);分类 cs.CL

Comments Accepted for poster presentation at the CVPR 2025 Workshop on Multimodal Foundation Models (MMFM3)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11235 2025-02-12 cs.CL 74%

GT2Vec: Large Language Models as Multi-Modal Encoders for Text and Graph-Structured Data

Jiacheng Lin, Kun Qian, Haoyu Han, Nurendra Choudhary, Tianxin Wei, Zhongruo Wang, Sahika Genc, Edward W Huang, Sheng Wang, Karthik Subbian, Danai Koutra, Jimeng Sun

专题命中 跨模态检索 :multi-modal(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02527 2025-01-07 cs.CV 74%

Vision-Driven Prompt Optimization for Large Language Models in Multimodal Generative Tasks

Leo Franklin, Apiradee Boonmee, Kritsada Wongsuwan

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00318 2025-01-03 cs.CV cs.LG 74%

Improving Text-based Person Search via Part-level Cross-modal Correspondence

Jicheol Park, Boseung Jeong, Dongwon Kim, Suha Kwak

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19547 2024-12-23 cs.LG cs.CV 74%

CLIPLoss and Norm-Based Data Selection Methods for Multimodal Contrastive Learning

Yiping Wang, Yifang Chen, Wendan Yan, Alex Fang, Wenjing Zhou, Kevin Jamieson, Simon Shaolei Du

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments This paper supercedes our previous VAS paper (arXiv:2402.02055). It's accepted by NeurIPS2024 as spotlight paper. DataComp benchmark: https://www.datacomp.ai/dcclip/leaderboard.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07619 2024-12-11 cs.CL 74%

DRUM: Learning Demonstration Retriever for Large MUlti-modal Models

Ellen Yi-Ge, Jiechao Gao, Wei Han, Wei Zhu

专题命中 跨模态检索 :multi-modal(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23437 2024-11-01 cs.LG cs.CL cs.IR 74%

Mind the Gap: A Generalized Approach for Cross-Modal Embedding Alignment

Arihan Yadav, Alan McMillan

专题命中 跨模态检索 :cross-modal(title);分类 cs.CL

Comments 18 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06827 2024-09-12 cs.CV 74%

Cross-Modal Self-Supervised Learning with Effective Contrastive Units for LiDAR Point Clouds

Mu Cai, Chenxu Luo, Yong Jae Lee, Xiaodong Yang

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

Comments IROS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07925 2024-07-12 cs.IR cs.AI cs.SI 74%

Enhancing Social Media Personalization: Dynamic User Profile Embeddings and Multimodal Contextual Analysis Using Transformer Models

Pranav Vachharajani

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

Comments 21 pages, 13 figures. Mentor: Prof Pritam Ranjan

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17615 2024-06-26 cs.SE cs.AI cs.LG 74%

Aligning Programming Language and Natural Language: Exploring Design Choices in Multi-Modal Transformer-Based Embedding for Bug Localization

Partha Chakraborty, Venkatraman Arumugam, Meiyappan Nagappan

专题命中 跨模态检索 :multi-modal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12024 2024-01-23 cs.RO cs.AI cs.LG 74%

Multimodal Visual-Tactile Representation Learning through Self-Supervised Contrastive Pre-Training

Vedant Dave, Fotios Lygerakis, Elmar Rueckert

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03043 2024-01-09 cs.CV 74%

Learning Multimodal Volumetric Features for Large-Scale Neuron Tracing

Qihua Chen, Xuejin Chen, Chenxuan Wang, Yixiong Liu, Zhiwei Xiong, Feng Wu

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments 9 pages, 6 figures, AAAI 2024 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00693 2023-12-12 cs.AI 74%

On Task-personalized Multimodal Few-shot Learning for Visually-rich Document Entity Retrieval

Jiayi Chen, Hanjun Dai, Bo Dai, Aidong Zhang, Wei Wei

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

Comments Paper published at Findings of the Association for Computational Linguistics: EMNLP, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.06286 2023-11-07 eess.SP cs.CV 74%

Automated Cardiovascular Record Retrieval by Multimodal Learning between Electrocardiogram and Clinical Report

Jielin Qiu, Jiacheng Zhu, Shiqi Liu, William Han, Jingqi Zhang, Chaojing Duan, Michael Rosenberg, Emerson Liu, Douglas Weber, Ding Zhao

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments Accepted to the ML4H 2023 Proceedings track

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.16386 2023-09-01 cs.CV 74%

RGB-T Tracking via Multi-Modal Mutual Prompt Learning

Yang Luo, Xiqing Guo, Hui Feng, Lei Ao

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV

Comments 9 pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14523 2023-07-28 cs.CV 74%

Towards multi-modal anatomical landmark detection for ultrasound-guided brain tumor resection with contrastive learning

Soorena Salari, Amirhossein Rasoulian, Hassan Rivaz, Yiming Xiao

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV

Comments Accepted in MICCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13805 2023-05-24 cs.CL 74%

Towards Zero-shot Relation Extraction in Web Mining: A Multimodal Approach with Relative XML Path

Zilong Wang, Jingbo Shang

专题命中 跨模态检索 :multimodal(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08660 2023-04-19 cs.RO cs.AI 74%

(LC)$^2$: LiDAR-Camera Loop Constraints For Cross-Modal Place Recognition

Alex Junho Lee, Seungwon Song, Hyungtae Lim, Woojoo Lee, Hyun Myung

专题命中 跨模态检索 :cross-modal(title);分类 cs.AI

Comments 8 pages, 11 figures, Accepted to IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.15770 2023-03-29 eess.IV cs.CV physics.med-ph 74%

DDMM-Synth: A Denoising Diffusion Model for Cross-modal Medical Image Synthesis with Sparse-view Measurement Embedding

Xiaoyue Li, Kai Shang, Gaoang Wang, Mark D. Butala

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

Comments llncs.cls v2.20,12 pages with 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.01824 2022-11-04 cs.CL 74%

Human in the loop approaches in multi-modal conversational task guidance system development

Ramesh Manuvinakurike, Sovan Biswas, Giuseppe Raffa, Richard Beckwith, Anthony Rhodes, Meng Shi, Gesem Gudino Mejia, Saurav Sahay, Lama Nachman

专题命中 跨模态检索 :multi-modal(title);分类 cs.CL

Comments SCAI @ SIGIR

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.03838 2022-10-11 cs.CV 74%

Learning to embed semantic similarity for joint image-text retrieval

Noam Malali, Yosi Keller

专题命中 跨模态检索 :image-text(title);分类 cs.CV

Comments in IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.02329 2022-09-07 cs.CV cs.LG 74%

Multimodal contrastive learning for remote sensing tasks

Umangi Jain, Alex Wilson, Varun Gulshan

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07734 2022-07-26 q-bio.GN cs.AI cs.GL 74%

COEM: Cross-Modal Embedding for MetaCell Identification

Haiyi Mao, Minxue Jia, Jason Xiaotian Dou, Haotian Zhang, Panayiotis V. Benos

专题命中 跨模态检索 :cross-modal(title);分类 cs.AI

Comments 5 pages, 2 figures, ICML workshop on computational biology

详情

展开后加载摘要…

URL PDF HTML 收藏