arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45832 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4634 篇

2304.09172 2024-01-19 cs.CV cs.LG 83%

Hyperbolic Image-Text Representations

Karan Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson, Ramakrishna Vedantam

专题命中 图文多模态 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV

Comments ICML 2023 (v3: Add link to code in abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09725 2024-01-19 cs.IR cs.MM 83%

Enhancing Image-Text Matching with Adaptive Feature Aggregation

Zuhui Wang, Yunting Yin, I. V. Ramakrishnan

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.MM

Comments Accepted by ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17468 2024-01-09 cs.CV cs.LG 83%

Cross-modal Active Complementary Learning with Self-refining Correspondence

Yang Qin, Yuan Sun, Dezhong Peng, Joey Tianyi Zhou, Xi Peng, Peng Hu

专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments This paper is accepted by NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01736 2024-01-05 cs.CV 83%

Few-shot Adaptation of Multi-modal Foundation Models: A Survey

Fan Liu, Tianshu Zhang, Wenwen Dai, Wenwen Cai, Xiaocong Zhou, Delong Chen

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14233 2023-12-25 cs.CV 83%

VCoder: Versatile Vision Encoders for Multimodal Large Language Models

Jitesh Jain, Jianwei Yang, Humphrey Shi

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Project Page: https://praeclarumjj3.github.io/vcoder/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08084 2023-12-18 cs.AI 83%

A Novel Energy based Model Mechanism for Multi-modal Aspect-Based Sentiment Analysis

Tianshuo Peng, Zuchao Li, Ping Wang, Lefei Zhang, Hai Zhao

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.AI

Comments AAAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.04364 2023-12-12 cs.MM 83%

ITportrait: Image-Text Coupled 3D Portrait Domain Adaptation

Xiangwen Deng, Yufeng Wang, Yuanhao Cai, Jingxiang Sun, Yebin Liu, Haoqian Wang

专题命中 图文多模态 :image-text(title,abstract);multi-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05863 2023-11-13 cs.CR cs.CV 83%

Watermarking Vision-Language Pre-trained Models for Multi-modal Embedding as a Service

Yuanmin Tang, Jing Yu, Keke Gai, Xiangyan Qu, Yue Hu, Gang Xiong, Qi Wu

专题命中 图文多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05437 2023-11-10 cs.CV cs.AI cs.CL cs.LG cs.MM 83%

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang, Jianfeng Gao, Chunyuan Li

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 25 pages, 25M file size. Project Page: https://llava-vl.github.io/llava-plus/

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10350 2023-10-27 cs.LG cs.CV 83%

Improving Multimodal Datasets with Image Captioning

Thao Nguyen, Samir Yitzhak Gadre, Gabriel Ilharco, Sewoong Oh, Ludwig Schmidt

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted at NeurIPS 2023 Datasets & Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08368 2023-10-13 cs.CV 83%

Mapping Memes to Words for Multimodal Hateful Meme Classification

Giovanni Burbi, Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini, Alberto Del Bimbo

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments ICCV2023 CLVL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.01615 2023-09-19 cs.CV 83%

ConTEXTual Net: A Multimodal Vision-Language Model for Segmentation of Pneumothorax

Zachary Huemann, Xin Tie, Junjie Hu, Tyler J. Bradshaw

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03921 2023-09-11 cs.CV 83%

C-CLIP: Contrastive Image-Text Encoders to Close the Descriptive-Commentative Gap

William Theisen, Walter Scheirer

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV

Comments 11 Pages, 5 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10313 2023-09-06 cs.CL 83%

Beyond Triplet: Leveraging the Most Data for Multimodal Machine Translation

Yaoming Zhu, Zewei Sun, Shanbo Cheng, Luyang Huang, Liwei Wu, Mingxuan Wang

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL

Comments 8 pages, ACL 2023 Finding

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16527 2023-08-22 cs.IR cs.CV 83%

OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander M. Rush, Douwe Kiela, Matthieu Cord, Victor Sanh

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13181 2023-08-15 cs.LG cs.CV 83%

Sample-Specific Debiasing for Better Image-Text Models

Peiqi Wang, Yingcheng Liu, Ching-Yun Ko, William M. Wells, Seth Berkowitz, Steven Horng, Polina Golland

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Machine Learning for Healthcare Conference 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04832 2023-07-07 cs.MM cs.SI 83%

Unifying Multimodal Source and Propagation Graph for Rumour Detection on Social Media with Missing Features

Tsun-Hin Cheung, Kin-Man Lam

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15195 2023-07-04 cs.CV 83%

Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, Rui Zhao

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.01758 2023-05-26 cs.CV 83%

Improving Zero-shot Generalization and Robustness of Multi-modal Models

Yunhao Ge, Jie Ren, Andrew Gallagher, Yuxiao Wang, Ming-Hsuan Yang, Hartwig Adam, Laurent Itti, Balaji Lakshminarayanan, Jiaping Zhao

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12029 2023-05-12 cs.CV 83%

VLCDoC: Vision-Language Contrastive Pre-Training Model for Cross-Modal Document Classification

Souhail Bakkali, Zuheng Ming, Mickael Coustaty, Marçal Rusiñol, Oriol Ramos Terrades

专题命中 图文多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted at PR

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04530 2023-05-09 cs.CL 83%

A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual Clues

Yunxin Li, Baotian Hu, Xinyu Chen, Yuxin Ding, Lin Ma, Min Zhang

专题命中 图文多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CL

Comments Accepted to ACL 2023 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03307 2023-04-10 cs.CV eess.IV 83%

Vita-CLIP: Video and text adaptive CLIP via Multimodal Prompting

Syed Talal Wasim, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan, Mubarak Shah

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted at CVPR-2023. Codes/models available at https://github.com/TalalWasim/Vita-CLIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.03117 2023-04-04 cs.CV 83%

MaPLe: Multi-modal Prompt Learning

Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, Fahad Shahbaz Khan

专题命中 图文多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted at CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00182 2023-03-28 cs.CV 83%

Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language Models

Wenhao Wu, Xiaohan Wang, Haipeng Luo, Jingdong Wang, Yi Yang, Wanli Ouyang

专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.06591 2023-03-14 cs.CV 83%

Accommodating Audio Modality in CLIP for Multimodal Processing

Ludan Ruan, Anwen Hu, Yuqing Song, Liang Zhang, Sipeng Zheng, Qin Jin

专题命中 图文多模态 :multimodal(title,abstract);vision-language-audio(abstract);分类 cs.CV

Comments Accepted by AAAI2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.06430 2023-03-03 cs.CV 83%

CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Hongwei Xue, Yuchong Sun, Bei Liu, Jianlong Fu, Ruihua Song, Houqiang Li, Jiebo Luo

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01887 2023-02-02 cs.CV 83%

Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

Sunan He, Taian Guo, Tao Dai, Ruizhi Qiao, Bo Ren, Shu-Tao Xia

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments AAAI 2023 (Oral presentation paper). Updated version

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.12570 2022-12-19 cs.LG cs.CV 83%

Learning Multimodal VAEs through Mutual Supervision

Tom Joy, Yuge Shi, Philip H. S. Torr, Tom Rainforth, Sebastian M. Schmon, N. Siddharth

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.14713 2022-11-21 cs.IR cs.AI 83%

Image-text Retrieval: A Survey on Recent Research and Development

Min Cao, Shiping Li, Juntao Li, Liqiang Nie, Min Zhang

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.AI

Comments Accpted by IJCAI'2022 survey track

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.13188 2022-10-25 cs.CV 83%

Dissecting Deep Metric Learning Losses for Image-Text Retrieval

Hong Xuan, Xi Chen

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

Comments arXiv admin note: text overlap with arXiv:2201.11307

Journal ref WACV2023

详情

展开后加载摘要…

URL PDF HTML 收藏