arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45986 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4644 篇

2405.12456 2024-05-22 eess.IV cs.CV cs.LG 79%

Mutual Information Analysis in Multimodal Learning Systems

Hadi Hadizadeh, S. Faegheh Yeganli, Bahador Rashidi, Ivan V. Bajić

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 6 pages, 7 figures, IEEE MIPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00384 2024-05-22 cs.CV 79%

TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias

Sanghyun Jo, Soohyun Ryu, Sungyub Kim, Eunho Yang, Kyungsu Kim

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11496 2024-05-21 cs.CV cs.IR 79%

DEMO: A Statistical Perspective for Efficient Image-Text Matching

Fan Zhang, Xian-Sheng Hua, Chong Chen, Xiao Luo

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09981 2024-05-17 cs.CV 79%

Adversarial Robustness for Visual Grounding of Multimodal Large Language Models

Kuofeng Gao, Yang Bai, Jiawang Bai, Yong Yang, Shu-Tao Xia

专题命中 图文多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments ICLR 2024 Workshop on Reliable and Responsible Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05949 2024-05-10 cs.CV 79%

CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts

Jiachen Li, Xinyao Wang, Sijie Zhu, Chia-Wen Kuo, Lu Xu, Fan Chen, Jitesh Jain, Humphrey Shi, Longyin Wen

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.00962 2024-05-03 cs.CV 79%

FITA: Fine-grained Image-Text Aligner for Radiology Report Generation

Honglong Yang, Hui Tang, Xiaomeng Li

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.00029 2024-05-02 cs.CV cs.IR 79%

Automatic Creative Selection with Cross-Modal Matching

Alex Kim, Jia Huang, Rob Monarch, Jerry Kwac, Anikesh Kamath, Parmeshwar Khurd, Kailash Thiyagarajan, Goodman Gu

专题命中 图文多模态 :cross-modal(title);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15655 2024-04-25 cs.CV 79%

Multi-Modal Proxy Learning Towards Personalized Visual Multiple Clustering

Jiawei Yao, Qi Qian, Juhua Hu

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2024. Project page: https://github.com/Alexander-Yao/Multi-MaP

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11864 2024-04-25 cs.CV 79%

Progressive Multi-modal Conditional Prompt Tuning

Xiaoyu Qiu, Hao Feng, Yuechen Wang, Wengang Zhou, Houqiang Li

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09797 2024-04-16 cs.CV 79%

TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Bozhi Luan, Hao Feng, Hong Chen, Yonghui Wang, Wengang Zhou, Houqiang Li

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14991 2024-04-15 cs.CV 79%

FoodLMM: A Versatile Food Assistant using Large Multi-modal Model

Yuehao Yin, Huiyan Qi, Bin Zhu, Jingjing Chen, Yu-Gang Jiang, Chong-Wah Ngo

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04676 2024-04-10 cs.CV eess.IV 79%

Stacked Cross-modal Feature Consolidation Attention Networks for Image Captioning

Mozhgan Pourkeshavarz, Shahabedin Nabavi, Mohsen Ebrahimi Moghaddam, Mehrnoush Shamsfard

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Journal ref Multimedia Tools and Applications, Volume 83, pages 12209-12233, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04960 2024-04-09 cs.CV 79%

PairAug: What Can Augmented Image-Text Pairs Do for Radiology?

Yutong Xie, Qi Chen, Sinuo Wang, Minh-Son To, Iris Lee, Ee Win Khoo, Kerolos Hendy, Daniel Koh, Yong Xia, Qi Wu

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted to CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.07633 2024-04-09 cs.CL cs.LG 79%

Interpretable Detection of Out-of-Context Misinformation with Neural-Symbolic-Enhanced Large Multimodal Model

Yizhou Zhang, Loc Trinh, Defu Cao, Zijun Cui, Yan Liu

专题命中 图文多模态 :multimodal(title);cross-modal(abstract);分类 cs.CL

Comments 9 Pages, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04231 2024-04-08 cs.CV 79%

Image-Text Co-Decomposition for Text-Supervised Semantic Segmentation

Ji-Jia Wu, Andy Chia-Hao Chang, Chieh-Yu Chuang, Chun-Pei Chen, Yu-Lun Liu, Min-Hung Chen, Hou-Ning Hu, Yung-Yu Chuang, Yen-Yu Lin

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01779 2024-04-01 cs.CV 79%

HallE-Control: Controlling Object Hallucination in Large Multimodal Models

Bohan Zhai, Shijia Yang, Chenfeng Xu, Sheng Shen, Kurt Keutzer, Chunyuan Li, Manling Li

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Our code is publicly available at https://github.com/bronyayang/HallE_Control

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02918 2024-03-21 cs.CV 79%

Multimodal Prompt Perceiver: Empower Adaptiveness, Generalizability and Fidelity for All-in-One Image Restoration

Yuang Ai, Huaibo Huang, Xiaoqiang Zhou, Jiexiang Wang, Ran He

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 13 pages, 8 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11593 2024-03-19 cs.CV 79%

End-to-end multi-modal product matching in fashion e-commerce

Sándor Tóth, Stephen Wilson, Alexia Tsoukara, Enric Moreu, Anton Masalovich, Lars Roemheld

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 9 pages, submitted to SIGKDD

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12596 2024-03-18 cs.CV 79%

UniHDA: A Unified and Versatile Framework for Multi-Modal Hybrid Domain Adaptation

Hengjia Li, Yang Liu, Yuqi Lin, Zhanwei Zhang, Yibo Zhao, weihang Pan, Tu Zheng, Zheng Yang, Yuchun Jiang, Boxi Wu, Deng Cai

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09027 2024-03-15 cs.CV 79%

VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework

Chris Kelly, Luhui Hu, Bang Yang, Yu Tian, Deshun Yang, Cindy Yang, Zaoshan Huang, Zihao Li, Jiayin Hu, Yuexian Zou

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 17 pages, 5 figures, and 1 table. arXiv admin note: substantial text overlap with arXiv:2311.10125

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11570 2024-03-13 cs.CV 79%

Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model

Shuailei Ma, Chen-Wei Xie, Ying Wei, Siyang Sun, Jiaqi Fan, Xiaoyi Bao, Yuxin Guo, Yun Zheng

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments We find that the statistical information in Figure 2 neglect the statistics for tSOS, so we make corrections. Additionally, we change the statistical samples to those where CLIP misidentify, but prompt tuning identify correctly. At the same time, we also revise some of the descriptions. The changes to the supplementary materials will be updated shortly. arXiv admin note: text overlap with arXiv:2307.06948 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09957 2024-03-13 cs.CV 79%

Self-paced Multi-grained Cross-modal Interaction Modeling for Referring Expression Comprehension

Peihan Miao, Wei Su, Gaoang Wang, Xuewei Li, Xi Li

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by TIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06295 2024-03-12 cs.CV 79%

A streamlined Approach to Multimodal Few-Shot Class Incremental Learning for Fine-Grained Datasets

Thang Doan, Sima Behpour, Xin Li, Wenbin He, Liang Gou, Liu Ren

专题命中 图文多模态 :multimodal(title);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05141 2024-03-11 cs.CV 79%

Med3DInsight: Enhancing 3D Medical Image Understanding with 2D Multi-Modal Large Language Models

Qiuhui Chen, Huping Ye, Yi Hong

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08825 2024-03-11 cs.CV 79%

From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Dongsheng Jiang, Yuchen Liu, Songlin Liu, Jin'e Zhao, Hao Zhang, Zhen Gao, Xiaopeng Zhang, Jin Li, Hongkai Xiong

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12105 2024-03-06 cs.LG cs.CL cs.SE 79%

Mass-Producing Failures of Multimodal Systems with Language Models

Shengbang Tong, Erik Jones, Jacob Steinhardt

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.00249 2024-03-04 cs.CV 79%

Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training

Haowei Liu, Yaya Shi, Haiyang Xu, Chunfeng Yuan, Qinghao Ye, Chenliang Li, Ming Yan, Ji Zhang, Fei Huang, Bing Li, Weiming Hu

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17680 2024-02-28 cs.CV 79%

MCF-VC: Mitigate Catastrophic Forgetting in Class-Incremental Learning for Multimodal Video Captioning

Huiyu Xiong, Lanxiao Wang, Heqian Qiu, Taijin Zhao, Benliu Qiu, Hongliang Li

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11325 2024-02-28 cs.CV 79%

ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models

Zhenghang Yuan, Zhitong Xiong, Lichao Mou, Xiao Xiang Zhu

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05496 2024-02-27 cs.CL 79%

Exploiting Pseudo Image Captions for Multimodal Summarization

Chaoya Jiang, Rui Xie, Wei Ye, Jinan Sun, Shikun Zhang

专题命中 图文多模态 :multimodal(title);cross-modal(abstract);分类 cs.CL

Comments Accepted at ACL2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏