arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3454 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3454 篇

2411.02537 2024-11-12 cs.CV cs.AI cs.CL cs.IR 67%

INQUIRE: A Natural World Text-to-Image Retrieval Benchmark

Edward Vendrow, Omiros Pantazis, Alexander Shepard, Gabriel Brostow, Kate E. Jones, Oisin Mac Aodha, Sara Beery, Grant Van Horn

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published in NeurIPS 2024, Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00683 2024-11-04 cs.CV cs.AI cs.CL cs.LG 67%

TaxaBind: A Unified Embedding Space for Ecological Applications

Srikumar Sastry, Subash Khanal, Aayush Dhakal, Adeel Ahmad, Nathan Jacobs

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01073 2024-09-04 cs.CV cs.AI cs.CL 67%

SCOPE: Sign Language Contextual Processing with Embedding from LLMs

Yuqi Liu, Wenqian Zhang, Sihan Ren, Chengyu Huang, Jingyi Yu, Lan Xu

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17274 2024-07-25 cs.MM cs.AI cs.CV 67%

Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation

Yongqi Li, Hongru Cai, Wenjie Wang, Leigang Qu, Yinwei Wei, Wenjie Li, Liqiang Nie, Tat-Seng Chua

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07532 2024-07-16 cs.CV cs.AI cs.CL 67%

Interfacing Foundation Models' Embeddings

Xueyan Zou, Linjie Li, Jianfeng Wang, Jianwei Yang, Mingyu Ding, Junyi Wei, Zhengyuan Yang, Feng Li, Hao Zhang, Shilong Liu, Arul Aravinthan, Yong Jae Lee, Lijuan Wang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments CODE: https://github.com/UX-Decoder/FIND

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09541 2024-07-16 cs.CL cs.AI cs.CV 67%

MATE: Meet At The Embedding -- Connecting Images with Long Texts

Young Kyun Jang, Junmo Kang, Yong Jae Lee, Donghyun Kim

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.13854 2024-04-16 cs.CV cs.AI cs.CL 67%

ComCLIP: Training-Free Compositional Image and Text Matching

Kenan Jiang, Xuehai He, Ruize Xu, Xin Eric Wang

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.07699 2024-02-13 cs.CL cs.AI cs.CV 67%

Retrieval-based Disentangled Representation Learning with Natural Language Supervision

Jiawei Zhou, Xiaoguang Li, Lifeng Shang, Xin Jiang, Qun Liu, Lei Chen

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.03320 2024-01-22 cs.LG 67%

BioBridge: Bridging Biomedical Foundation Models via Knowledge Graphs

Zifeng Wang, Zichen Wang, Balasubramaniam Srinivasan, Vassilis N. Ioannidis, Huzefa Rangwala, Rishita Anubhai

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract)

Comments ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09602 2023-12-19 cs.IR 67%

Multi-Modality is All You Need for Transferable Recommender Systems

Youhua Li, Hanwen Du, Yongxin Ni, Pengpeng Zhao, Qi Guo, Fajie Yuan, Xiaofang Zhou

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract)

Comments ICDE'24 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07931 2023-10-13 cs.LG cs.AI cs.CL cs.CV 67%

D2 Pruning: Message Passing for Balancing Diversity and Difficulty in Data Pruning

Adyasha Maharana, Prateek Yadav, Mohit Bansal

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 17 pages (Our code is available at https://github.com/adymaharana/d2pruning)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13923 2023-08-08 cs.CV cs.CL cs.MM 67%

Retrieval-based Knowledge Augmented Vision Language Pre-training

Jiahua Rao, Zifei Shan, Longpo Liu, Yao Zhou, Yuedong Yang

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments arXiv admin note: text overlap with arXiv:2210.09338 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.16395 2023-08-01 cs.CV cs.AI cs.CL cs.LG 67%

Bridging the Gap: Exploring the Capabilities of Bridge-Architectures for Complex Visual Reasoning Tasks

Kousik Rajesh, Mrigank Raman, Mohammed Asad Karim, Pranit Chawla

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05665 2023-06-01 cs.CV cs.AI cs.LG cs.MM 67%

ImageBind: One Embedding Space To Bind Them All

Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, Ishan Misra

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments CVPR 2023 (Highlighted Paper). Website: https://imagebind.metademolab.com/ Code/Models: https://github.com/facebookresearch/ImageBind

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12646 2023-02-21 cs.IR 67%

MAKE: Vision-Language Pre-training based Product Retrieval in Taobao Search

Xiaoyang Zheng, Zilong Wang, Ke Xu, Sen Li, Tao Zhuang, Qingwen Liu, Xiaoyi Zeng

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract)

Comments 5 pages, accepted to The Industry Track of the Web Conference 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04269 2023-02-09 cs.LG cs.AI cs.CL cs.CV 67%

Diagnosing and Rectifying Vision Models using Language

Yuhui Zhang, Jeff Z. HaoChen, Shih-Cheng Huang, Kuan-Chieh Wang, James Zou, Serena Yeung

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published at ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.07094 2023-01-18 cs.CV cs.AI cs.CL cs.LG 67%

Learning Customized Visual Models with Retrieval-Augmented Knowledge

Haotian Liu, Kilho Son, Jianwei Yang, Ce Liu, Jianfeng Gao, Yong Jae Lee, Chunyuan Li

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.02495 2022-09-16 cs.CV cs.AI cs.CL 67%

Sign Language Video Retrieval with Free-Form Textual Queries

Amanda Duarte, Samuel Albanie, Xavier Giró-i-Nieto, Gül Varol

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01058 2022-07-05 cs.AI cs.CV cs.HC cs.MM 67%

Chat-to-Design: AI Assisted Personalized Fashion Design

Weiming Zhuang, Chongjie Ye, Ying Xu, Pengzhi Mao, Shuai Zhang

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.10787 2022-02-23 cs.CL cs.AI cs.CV cs.LG 67%

VU-BERT: A Unified framework for Visual Dialog

Tong Ye, Shijing Si, Jianzong Wang, Rui Wang, Ning Cheng, Jing Xiao

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 5 pages, 2 figures, accepted by 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.07074 2021-12-06 cs.LG cs.AI cs.CL cs.CV 67%

Memotion Analysis through the Lens of Joint Embedding

Nethra Gunti, Sathyanarayanan Ramamoorthy, Parth Patwa, Amitava Das

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted as Student Abstract at AAAI-22

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.09150 2021-10-22 cs.CL cs.AI cs.CV 67%

VisualSem: A High-quality Knowledge Graph for Vision and Language

Houda Alberts, Teresa Huang, Yash Deshpande, Yibo Liu, Kyunghyun Cho, Clara Vania, Iacer Calixto

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted for publication at the 1st Multilingual Representation Learning workshop (MRL 2021) co-located with EMNLP 2021. 15 pages, 8 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.09109 2021-02-19 cs.CV cs.AI cs.MM 67%

Understanding and Creating Art with AI: Review and Outlook

Eva Cetinic, James She

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 17 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.15086 2021-01-01 cs.CL cs.AI cs.CV 67%

Accurate Word Representations with Universal Visual Guidance

Zhuosheng Zhang, Haojie Yu, Hai Zhao, Rui Wang, Masao Utiyama

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.03687 2020-05-26 cs.LG stat.ML 67%

COBRA: Contrastive Bi-Modal Representation Algorithm

Vishaal Udandarao, Abhishek Maiti, Deepak Srivatsav, Suryatej Reddy Vyalla, Yifang Yin, Rajiv Ratn Shah

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract)

Comments 13 Pages, 6 Figures and 10 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.11449 2020-04-27 cs.MM cs.CL cs.CV cs.LG 67%

Upgrading the Newsroom: An Automated Image Selection System for News Articles

Fangyu Liu, Rémi Lebret, Didier Orel, Philippe Sordet, Karl Aberer

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments Accepted to ACM Transactions on Multimedia Computing Communications and Applications (ACM TOMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.00850 2019-11-05 cs.AI cs.CL cs.CV 67%

Scene Graph based Image Retrieval -- A case study on the CLEVR Dataset

Sahana Ramnath, Amrita Saha, Soumen Chakrabarti, Mitesh M. Khapra

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 3 pages including references, Accepted at the ICCV 2019 Workshop - 'Linguistics Meets Image and Video Retrieval' (received Best Paper Award)

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.08454 2018-12-27 cs.CL cs.AI cs.CV cs.LG 67%

Attention Based Natural Language Grounding by Navigating Virtual Environment

Akilesh B, Abhishek Sinha, Mausoom Sarkar, Balaji Krishnamurthy

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at WACV 2019. Also at NeurIPS 2017 workshop on Visually-Grounded Interaction and Language (ViGIL)

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.01606 2018-07-02 cs.MM cs.AI cs.CV cs.LG 67%

Multimedia Semantic Integrity Assessment Using Joint Embedding Of Images And Text

Ayush Jaiswal, Ekraam Sabir, Wael AbdAlmageed, Premkumar Natarajan

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments *Ayush Jaiswal and Ekraam Sabir contributed equally to the work in this paper

Journal ref In Proceedings of the 2017 ACM on Multimedia Conference, pp. 1465-1471. ACM, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.02717 2016-08-10 cs.CV cs.AI cs.CL cs.LG 67%

Mean Box Pooling: A Rich Image Representation and Output Embedding for the Visual Madlibs Task

Ashkan Mokarian, Mateusz Malinowski, Mario Fritz

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to BMVC'16

详情

展开后加载摘要…

URL PDF HTML 收藏