arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3450 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3450 篇

2309.16741 2024-01-03 cs.LG cs.AI cs.HC 79%

Multi-Modal Financial Time-Series Retrieval Through Latent Space Projections

Tom Bamford, Andrea Coletta, Elizabeth Fons, Sriram Gopalakrishnan, Svitlana Vyetrenko, Tucker Balch, Manuela Veloso

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted to ICAIF 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15021 2023-12-27 cs.CL 79%

Towards a Unified Multimodal Reasoning Framework

Abhinav Arun, Dipendra Singh Mal, Mehul Soni, Tomohiro Sawada

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments 6 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10547 2023-12-15 cs.CV cs.CY 79%

Rethinking Multimodal Content Moderation from an Asymmetric Angle with Mixed-modality

Jialin Yuan, Ye Yu, Gaurav Mittal, Matthew Hall, Sandra Sajeev, Mei Chen

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00857 2023-12-05 cs.LG cs.AI cs.HC eess.SP 79%

Latent Space Explorer: Visual Analytics for Multimodal Latent Space Exploration

Bum Chul Kwon, Samuel Friedman, Kai Xu, Steven A Lubitz, Anthony Philippakis, Puneet Batra, Patrick T Ellinor, Kenney Ng

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11009 2023-11-21 cs.CL 79%

Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimodal Emotion Recognition

Dongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu Okumura

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04678 2023-11-14 cs.CV q-bio.QM 79%

Weakly supervised cross-modal learning in high-content screening

Watkinson Gabriel, Cohen Ethan, Bourriez Nicolas, Bendidi Ihab, Bollot Guillaume, Genovesio Auguste

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.03964 2023-11-08 cs.CV 79%

Enhancing Multimodal Compositional Reasoning of Visual Language Models with Generative Negative Mining

Ugur Sahin, Hang Li, Qadeer Khan, Daniel Cremers, Volker Tresp

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to WACV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.13214 2023-11-08 cs.LG cs.AI 79%

FedMEKT: Distillation-based Embedding Knowledge Transfer for Multimodal Federated Learning

Huy Q. Le, Minh N. H. Nguyen, Chu Myaet Thwal, Yu Qiao, Chaoning Zhang, Choong Seon Hong

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18770 2023-10-31 cs.IR cs.MM 79%

Leveraging Multimodal Features and Item-level User Feedback for Bundle Construction

Yunshan Ma, Xiaohao Liu, Yinwei Wei, Zhulin Tao, Xiang Wang, Tat-Seng Chua

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

Journal ref WSDM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13265 2023-10-23 cs.CL 79%

MoqaGPT : Zero-Shot Multi-modal Open-domain Question Answering with Large Language Model

Le Zhang, Yihong Wu, Fengran Mo, Jian-Yun Nie, Aishwarya Agrawal

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted into EMNLP2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.11029 2023-10-19 cs.SD cs.IR eess.AS 79%

CLaMP: Contrastive Language-Music Pre-training for Cross-Modal Symbolic Music Information Retrieval

Shangda Wu, Dingyao Yu, Xu Tan, Maosong Sun

专题命中 跨模态检索 :cross-modal(title,abstract);分类 eess.AS

Comments 11 pages, 5 figures, 5 tables, accepted by ISMIR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.12403 2023-10-16 cs.CV cs.LG cs.RO 79%

ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings

Arjun Majumdar, Gunjan Aggarwal, Bhavika Devnani, Judy Hoffman, Dhruv Batra

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments code: https://github.com/gunagg/zson

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08032 2023-10-13 cs.AI 79%

Incorporating Domain Knowledge Graph into Multimodal Movie Genre Classification with Self-Supervised Attention and Contrastive Learning

Jiaqi Li, Guilin Qi, Chuanyi Zhang, Yongrui Chen, Yiming Tan, Chenlong Xia, Ye Tian

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07668 2023-10-12 cs.LG cs.AI 79%

GRaMuFeN: Graph-based Multi-modal Fake News Detection in Social Media

Makan Kananian, Fatima Badiei, S. AmirAli Gh. Ghahramani

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.08671 2023-10-10 cs.CR cs.AI 79%

Deep Cross-Modal Steganography Using Neural Representations

Gyojin Han, Dong-Jae Lee, Jiwan Hur, Jaehyun Choi, Junmo Kim

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.AI

Comments ICIP 2023 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.11880 2023-10-06 cs.MM 79%

Adaptive Marginalized Semantic Hashing for Unpaired Cross-Modal Retrieval

Kaiyi Luo, Chao Zhang, Huaxiong Li, Xiuyi Jia, Chunlin Chen

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11553 2023-09-26 cs.CV 79%

MuMUR : Multilingual Multimodal Universal Retrieval

Avinash Madasu, Estelle Aflalo, Gabriela Ben Melech Stan, Shachar Rosenman, Shao-Yen Tseng, Gedas Bertasius, Vasudev Lal

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments This is an extension of the previous MKTVR paper (for which you can find a reference here : https://dl.acm.org/doi/abs/10.1007/978-3-031-28244-7_42 or in a previous version on arxiv). This version was published to the Information Retrieval Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05032 2023-09-12 cs.CV 79%

Unified Contrastive Fusion Transformer for Multimodal Human Action Recognition

Kyoung Ok Yang, Junho Koh, Jun Won Choi

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08330 2023-09-12 cs.CV 79%

Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival Prediction

Yingxue Xu, Hao Chen

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 4 figures, accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01262 2023-09-06 cs.CV cs.HC cs.LG eess.SP 79%

Multimodal Contrastive Learning with Hard Negative Sampling for Human Activity Recognition

Hyeongju Choi, Apoorva Beedu, Irfan Essa

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.16649 2023-09-01 cs.CV 79%

Learning with Multi-modal Gradient Attention for Explainable Composed Image Retrieval

Prateksha Udhayanan, Srikrishna Karanam, Balaji Vasan Srinivasan

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14613 2023-08-29 cs.CV 79%

MS-Net: A Multi-modal Self-supervised Network for Fine-Grained Classification of Aircraft in SAR Images

Bingying Yue, Jianhao Li, Hao Shi, Yupei Wang, Honghu Zhong

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08306 2023-08-22 cs.MM cs.LG 79%

Towards Balanced Active Learning for Multimodal Classification

Meng Shen, Yizheng Huang, Jianxiong Yin, Heqing Zou, Deepu Rajan, Simon See

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

Comments 12 pages, accepted by ACMMM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.08558 2023-08-21 q-fin.ST cs.AI 79%

BIRP: Bitcoin Information Retrieval Prediction Model Based on Multimodal Pattern Matching

Minsuk Kim, Byungchul Kim, Junyeong Yong, Jeongwoo Park, Gyeongmin Kim

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments 5 pages, 2 figures, KDD 2023 Machine Learning in Finance workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.05948 2023-08-14 cs.CV 79%

Uncertainty-Aware Cross-Modal Transfer Network for Sketch-Based 3D Shape Retrieval

Yiyang Cai, Jiaming Lu, Jiewen Wang, Shuang Liang

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments 6 pages, 7 figures; To be published in IEEE International Conference on Multimedia and Expo 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.06350 2023-07-31 cs.CV 79%

VITR: Augmenting Vision Transformers with Relation-Focused Learning for Cross-Modal Information Retrieval

Yan Gong, Georgina Cosma, Axel Finke

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14244 2023-07-27 cs.MM 79%

Neural-based Cross-modal Search and Retrieval of Artwork

Yan Gong, Georgina Cosma, Axel Finke

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16761 2023-07-25 cs.CV 79%

Improving Cross-Modal Retrieval with Set of Diverse Embeddings

Dongwon Kim, Namyup Kim, Suha Kwak

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2023 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10782 2023-07-21 cs.CV 79%

See More and Know More: Zero-shot Point Cloud Segmentation via Multi-modal Visual Data

Yuhang Lu, Qi Jiang, Runnan Chen, Yuenan Hou, Xinge Zhu, Yuexin Ma

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
1411.7798 2023-07-19 cs.CV 79%

Cross-Modal Learning via Pairwise Constraints

Ran He, Man Zhang, Liang Wang, Ye Ji, Qiyue Yin

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments 12 pages, 5 figures, 70 references

详情

展开后加载摘要…

URL PDF HTML 收藏