arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45986 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4644 篇

2309.04790 2023-09-12 cs.CL 79%

MMHQA-ICL: Multimodal In-context Learning for Hybrid Question Answering over Text, Tables and Images

Weihao Liu, Fangyu Lei, Tongxu Luo, Jiahe Lei, Shizhu He, Jun Zhao, Kang Liu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.18232 2023-08-16 cs.CV 79%

DIME-FM: DIstilling Multimodal and Efficient Foundation Models

Ximeng Sun, Pengchuan Zhang, Peizhao Zhang, Hardik Shah, Kate Saenko, Xide Xia

专题命中 图文多模态 :multimodal(title);image-text(abstract);分类 cs.CV

Comments Accepted to ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.00847 2023-08-15 cs.CV 79%

MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the Wild

Yuanyuan Liu, Wei Dai, Chuanxu Feng, Wenbin Wang, Guanghao Yin, Jiabei Zeng, Shiguang Shan

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments This paper has been accepted by ACM MM'22

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06262 2023-08-14 cs.LG cs.AI 79%

Foundation Model is Efficient Multimodal Multitask Model Selector

Fanqing Meng, Wenqi Shao, Zhanglin Peng, Chonghe Jiang, Kaipeng Zhang, Yu Qiao, Ping Luo

专题命中 图文多模态 :multimodal(title);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12964 2023-08-07 cs.CV 79%

Text-based Person Search without Parallel Image-Text Data

Yang Bai, Jingyao Wang, Min Cao, Chen Chen, Ziqiang Cao, Liqiang Nie, Min Zhang

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.05920 2023-07-13 eess.IV cs.CV cs.LG 79%

Unified Medical Image-Text-Label Contrastive Learning With Continuous Prompt

Yuhao Wang

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.02677 2023-07-07 cs.CV 79%

Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Teng Wang, Jinrui Zhang, Junjie Fei, Hao Zheng, Yunlong Tang, Zhe Li, Mingqi Gao, Shanshan Zhao

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Tech-report

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07198 2023-06-13 cs.CL 79%

A Survey of Vision-Language Pre-training from the Lens of Multimodal Machine Translation

Jeremy Gwinnup, Kevin Duh

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02348 2023-06-06 cs.CL 79%

Leverage Points in Modality Shifts: Comparing Language-only and Multimodal Word Representations

Aleksey Tikhonov, Lisa Bylinina, Denis Paperno

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted for StarSEM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12256 2023-05-26 cs.CL 79%

Scene Graph as Pivoting: Inference-time Image-free Unsupervised Multimodal Machine Translation with Visual Scene Hallucination

Hao Fei, Qian Liu, Meishan Zhang, Min Zhang, Tat-Seng Chua

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04524 2023-05-09 cs.CV 79%

Scene Text Recognition with Image-Text Matching-guided Dictionary

Jiajun Wei, Hongjian Zhan, Xiao Tu, Yue Lu, Umapada Pal

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted at ICDAR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05173 2023-04-12 cs.CV cs.LG 79%

Improving Image Recognition by Retrieving from Web-Scale Image-Text Data

Ahmet Iscen, Alireza Fathi, Cordelia Schmid

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00785 2023-03-28 cs.CV 79%

Learning to Generate Text-grounded Mask for Open-world Semantic Segmentation from Only Image-Text Pairs

Junbum Cha, Jonghwan Mun, Byungseok Roh

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments CVPR 2023 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13340 2023-03-24 cs.LG cs.CV 79%

Increasing Textual Context Size Boosts Medical Image-Text Matching

Idan Glassberg, Tom Hope

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12997 2023-03-24 cs.CV 79%

FER-former: Multi-modal Transformer for Facial Expression Recognition

Yande Li, Mingjie Wang, Minglun Gong, Yonggang Lu, Li Liu

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07122 2023-03-14 cs.CV 79%

ContextCLIP: Contextual Alignment of Image-Text pairs on CLIP visual representations

Chanda Grover, Indra Deep Mastan, Debayan Gupta

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments 11 Pages, 7 Figures, 2 Tables, ICVGIP

Journal ref ICVGIP, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07254 2023-03-03 cs.CV cs.LG 79%

The Role of Local Alignment and Uniformity in Image-Text Contrastive Learning on Medical Images

Philip Müller, Georgios Kaissis, Daniel Rueckert

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments NeurIPS 2022 Workshop: Self-Supervised Learning - Theory and Practice (Reason for updated version: correction of a typo in Eq. (2) and (3))

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02908 2023-02-07 cs.CV cs.IR 79%

LexLIP: Lexicon-Bottlenecked Language-Image Pre-Training for Large-Scale Image-Text Retrieval

Ziyang luo, Pu Zhao, Can Xu, Xiubo Geng, Tao Shen, Chongyang Tao, Jing Ma, Qingwen lin, Daxin Jiang

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.05453 2023-02-07 cs.CL 79%

It's Just a Matter of Time: Detecting Depression with Time-Enriched Multimodal Transformers

Ana-Maria Bucur, Adrian Cosma, Paolo Rosso, Liviu P. Dinu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at ECIR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.06844 2023-01-18 cs.CV 79%

USER: Unified Semantic Enhancement with Momentum Contrast for Image-Text Retrieval

Yan Zhang, Zhong Ji, Di Wang, Yanwei Pang, Xuelong Li

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.09297 2022-12-20 cs.CV 79%

Transferring General Multimodal Pretrained Models to Text Recognition

Junyang Lin, Xuancheng Ren, Yichang Zhang, Gao Liu, Peng Wang, An Yang, Chang Zhou

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13763 2022-12-01 cs.AI 79%

Clustering-Induced Generative Incomplete Image-Text Clustering (CIGIT-C)

Dongjin Guo, Xiaoming Su, Jiatai Wang, Limin Liu, Zhiyong Pei, Zhiwei Xu

专题命中 图文多模态 :image-text(title,abstract);分类 cs.AI

Comments 13 pages,12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.11223 2022-09-23 cs.CV cs.GR 79%

UniColor: A Unified Framework for Multi-Modal Colorization with Transformer

Zhitong Huang, Nanxuan Zhao, Jing Liao

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by SIGGRAPH Asia 2022. Project page: https://luckyhzt.github.io/unicolor

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.11333 2022-09-22 cs.CV 79%

Multi-modal Understanding and Generation for Medical Images and Text via Vision-Language Pre-Training

Jong Hak Moon, Hyungyung Lee, Woncheol Shin, Young-Hak Kim, Edward Choi

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Journal of Biomedical and Health Informatics

Journal ref IEEE Journal of Biomedical and Health Informatics 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06458 2022-09-07 cs.CV 79%

Cross-Modal Graph with Meta Concepts for Video Captioning

Hao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan Miao

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.04361 2022-08-10 cs.CV 79%

Semi-Supervised Cross-Modal Salient Object Detection with U-Structure Networks

Yunqing Bao, Hang Dai, Abdulmotaleb Elsaddik

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.09307 2022-05-20 cs.CV 79%

Support-set based Multi-modal Representation Enhancement for Video Captioning

Xiaoya Chen, Jingkuan Song, Pengpeng Zeng, Lianli Gao, Heng Tao Shen

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.00843 2022-04-07 cs.CV 79%

X-Trans2Cap: Cross-Modal Knowledge Transfer using Transformer for 3D Dense Captioning

Zhihao Yuan, Xu Yan, Yinghong Liao, Yao Guo, Guanbin Li, Zhen Li, Shuguang Cui

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments To appear in CVPR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.09173 2022-03-18 cs.CL 79%

On Vision Features in Multimodal Machine Translation

Bei Li, Chuanhao Lv, Zefan Zhou, Tao Zhou, Tong Xiao, Anxiang Ma, JingBo Zhu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments Long paper accepted by ACL2022 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.07519 2022-03-18 cs.CL 79%

Leveraging Visual Knowledge in Language Tasks: An Empirical Study on Intermediate Pre-training for Cross-modal Knowledge Transfer

Woojeong Jin, Dong-Ho Lee, Chenguang Zhu, Jay Pujara, Xiang Ren

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CL

Comments Accepted to ACL 2022, 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏