arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3450 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3450 篇

2304.11618 2023-04-25 cs.CL cs.AI 81%

Modality-Aware Negative Sampling for Multi-modal Knowledge Graph Embedding

Yichi Zhang, Mingyang Chen, Wen Zhang

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by IJCNN2023. Code is released in https://github.com/zjukg/MANS

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03717 2023-04-10 cs.LG cs.CL cs.CV 81%

On the Importance of Contrastive Loss in Multimodal Learning

Yunwei Ren, Yuanzhi Li

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00448 2023-04-04 cs.CV cs.MM 81%

The style transformer with common knowledge optimization for image-text retrieval

Wenrui Li, Zhengyu Ma, Jinqiao Shi, Xiaopeng Fan

专题命中 跨模态检索 :image-text(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.07740 2023-03-15 cs.CV cs.CL 81%

Efficient Image-Text Retrieval via Keyword-Guided Pre-Screening

Min Cao, Yang Bai, Jingyao Wang, Ziqiang Cao, Liqiang Nie, Min Zhang

专题命中 跨模态检索 :image-text(title,abstract);分类 cs.CV、cs.CL

Comments 11 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.01781 2023-03-06 cs.CV cs.CL 81%

Meme Sentiment Analysis Enhanced with Multimodal Spatial Encoding and Facial Embedding

Muzhaffar Hazman, Susan McKeever, Josephine Griffith

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Published as chapter in ISBN:978-3-031-26438-2

Journal ref In: Longo, L., OReilly, R. (eds) Artificial Intelligence and Cognitive Science. AICS 2022. Communications in Computer and Information Science, vol 1662. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06468 2023-02-15 cs.AI cs.CL cs.LG 81%

Contrastive Multimodal Learning for Emergence of Graphical Sensory-Motor Communication

Tristan Karch, Yoann Lemesle, Romain Laroche, Clément Moulin-Frier, Pierre-Yves Oudeyer

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.04366 2023-01-12 cs.CL cs.IR cs.LG cs.MM 81%

Multimodal Inverse Cloze Task for Knowledge-based Visual Question Answering

Paul Lerner, Olivier Ferret, Camille Guinaudeau

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments Accepted at ECIR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.14322 2023-01-02 cs.IR cs.AI cs.MM 81%

BagFormer: Better Cross-Modal Retrieval via bag-wise interaction

Haowen Hou, Xiaopeng Yan, Yigeng Zhang, Fengzong Lian, Zhanhui Kang

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.AI、cs.MM

Comments 8 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11319 2022-10-21 cs.CV cs.MM 81%

Image-Text Retrieval with Binary and Continuous Label Supervision

Zheng Li, Caili Guo, Zerun Feng, Jenq-Neng Hwang, Ying Jin, Yufeng Zhang

专题命中 跨模态检索 :image-text(title,abstract);分类 cs.CV、cs.MM

Comments 13 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03162 2022-10-21 cs.CV cs.CL 81%

Embedding Arithmetic of Multimodal Queries for Image Retrieval

Guillaume Couairon, Matthieu Cord, Matthijs Douze, Holger Schwenk

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments accepted at O-DRUM (CVPR workshop 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05040 2022-10-06 cs.CL cs.AI 81%

SANCL: Multimodal Review Helpfulness Prediction with Selective Attention and Natural Contrastive Learning

Wei Han, Hui Chen, Zhen Hai, Soujanya Poria, Lidong Bing

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted as a long paper at COLING 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12599 2022-09-27 cs.CV cs.AI 81%

Deep Manifold Hashing: A Divide-and-Conquer Approach for Semi-Paired Unsupervised Cross-Modal Retrieval

Yufeng Shi, Xinge You, Jiamiao Xu, Feng Zheng, Qinmu Peng, Weihua Ou

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.08012 2022-09-26 cs.CL cs.AI cs.LG 81%

CascadER: Cross-Modal Cascading for Knowledge Graph Link Prediction

Tara Safavi, Doug Downey, Tom Hope

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CL、cs.AI

Comments AKBC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.06515 2022-09-20 cs.CV cs.MM 81%

Learning to Evaluate Performance of Multi-modal Semantic Localization

Zhiqiang Yuan, Wenkai Zhang, Chongyang Li, Zhaoying Pan, Yongqiang Mao, Jialiang Chen, Shouke Li, Hongqi Wang, Xian Sun

专题命中 跨模态检索 :multi-modal(title);cross-modal(abstract);分类 cs.CV、cs.MM

Comments 19 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.07084 2022-09-16 cs.AI cs.CL 81%

Knowledge Graph Completion with Pre-trained Multimodal Transformer and Twins Negative Sampling

Yichi Zhang, Wen Zhang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by KDD 2022 Undergraduate Consortium

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.00682 2022-09-05 cs.CV cs.AI cs.GR cs.LG 81%

Zero-Shot Multi-Modal Artist-Controlled Retrieval and Exploration of 3D Object Sets

Kristofer Schlachter, Benjamin Ahlbrand, Zhu Wang, Valerio Ortenzi, Ken Perlin

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.14326 2022-08-31 cs.CV cs.AI 81%

GaitFi: Robust Device-Free Human Identification via WiFi and Vision Multimodal Learning

Lang Deng, Jianfei Yang, Shenghai Yuan, Han Zou, Chris Xiaoxuan Lu, Lihua Xie

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 12 pages, 8 figures, accepted by IEEE Internet of Things Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.03666 2022-08-19 cs.MM cs.CV cs.HC 81%

See What You See: Self-supervised Cross-modal Retrieval of Visual Stimuli from Brain Activity

Zesheng Ye, Lina Yao, Yu Zhang, Sylvia Gustin

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.03969 2022-07-05 cs.LG cs.CL cs.CV 81%

Multimodal Representations Learning Based on Mutual Information Maximization and Minimization and Identity Embedding for Multimodal Sentiment Analysis

Jiahao Zheng, Sen Zhang, Xiaoping Wang, Zhigang Zeng

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.07190 2022-06-16 cs.CL cs.AI cs.LG 81%

Codec at SemEval-2022 Task 5: Multi-Modal Multi-Transformer Misogynous Meme Classification Framework

Ahmed Mahran, Carlo Alessandro Borella, Konstantinos Perifanos

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted for publication at the 16th International Workshop on Semantic Evaluation, Task 5: MAMI - Multimedia Automatic Misogyny Identification co-located with NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.14643 2022-05-31 cs.CV cs.AI cs.LG 81%

Micro-Expression Recognition Based on Attribute Information Embedding and Cross-modal Contrastive Learning

Yanxin Song, Jianzong Wang, Tianbo Wu, Zhangcheng Huang, Jing Xiao

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments This paper has been accepted by IJCNN2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.08365 2022-05-18 cs.LG cs.AI cs.CV eess.IV 81%

Deep Supervised Information Bottleneck Hashing for Cross-modal Retrieval based Computer-aided Diagnosis

Yufeng Shi, Shuhuang Chen, Xinge You, Qinmu Peng, Weihua Ou, Yue Zhao

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 7 pages, 1 figure

Journal ref The AAAI-22 Workshop on Information Theory for Deep Learning (IT4DL).2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09860 2022-04-22 cs.CV cs.IR cs.MM 81%

Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local Information

Zhiqiang Yuan, Wenkai Zhang, Changyuan Tian, Xuee Rong, Zhengyuan Zhang, Hongqi Wang, Kun Fu, Xian Sun

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.MM

Journal ref in IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-16, 2022, Art no. 5620616

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.07668 2022-04-20 cs.CV cs.CL 81%

Dual-Key Multimodal Backdoors for Visual Question Answering

Matthew Walmer, Karan Sikka, Indranil Sur, Abhinav Shrivastava, Susmit Jha

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Published as conference paper at CVPR 2022. 22 pages, 11 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.07302 2022-04-18 cs.CV cs.CL 81%

Improving Cross-Modal Understanding in Visual Dialog via Contrastive Learning

Feilong Chen, Xiuyi Chen, Shuang Xu, Bo Xu

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.10852 2022-03-22 cs.CV cs.AI 81%

Multi-modal learning for predicting the genotype of glioma

Yiran Wei, Xi Chen, Lei Zhu, Lipei Zhang, Carola-Bibiane Schönlieb, Stephen J. Price, Chao Li

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.03060 2022-01-31 cs.LG cs.CL cs.CV eess.IV 81%

Contrastive Cross-Modal Pre-Training: A General Strategy for Small Sample Medical Imaging

Gongbo Liang, Connor Greenwell, Yu Zhang, Xiaoqin Wang, Ramakanth Kavuluru, Nathan Jacobs

专题命中 跨模态检索 :cross-modal(title);image-text(abstract);分类 cs.CV、cs.CL

Comments This work is accepted to the IEEE Journal of Biomedical and Health Informatics

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.12248 2021-12-15 cs.CV cs.CL 81%

Multi-Modal Answer Validation for Knowledge-Based VQA

Jialin Wu, Jiasen Lu, Ashish Sabharwal, Roozbeh Mottaghi

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Comments AAAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.13424 2021-11-29 cs.CV cs.AI cs.LG 81%

ContIG: Self-supervised Multimodal Contrastive Learning for Medical Imaging with Genetics

Aiham Taleb, Matthias Kirchler, Remo Monti, Christoph Lippert

专题命中 跨模态检索 :multimodal(title);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.11338 2021-11-29 cs.CV cs.CL cs.IR 81%

VLDeformer: Vision-Language Decomposed Transformer for Fast Cross-Modal Retrieval

Lisai Zhang, Hongfa Wu, Qingcai Chen, Yimeng Deng, Zhonghua Li, Dejiang Kong, Zhao Cao, Joanna Siebert, Yunpeng Han

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏