arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45832 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4634 篇

2412.08802 2025-04-25 cs.CL cs.CV cs.IR 84%

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Andreas Koukounas, Georgios Mastrapas, Sedigheh Eslami, Bo Wang, Mohammad Kalim Akram, Michael Günther, Isabelle Mohr, Saba Sturua, Nan Wang, Han Xiao

机构 * Jina AI GmbH(Jina AI公司)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments 30 pages, 1-10 main paper, 10-12 refs, 12-30 benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12902 2025-04-16 cs.AI cs.CL cs.LG 84%

IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities

Bin Wang, Chunyu Xie, Dawei Leng, Yuhui Yin

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10049 2025-04-15 cs.CV cs.CL 84%

Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure

Théo Gigant, Camille Guinaudeau, Frédéric Dufaux

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08269 2025-04-14 cs.CV cs.CL 84%

VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering

Qi Zhi Lim, Chin Poo Lee, Kian Ming Lim, Kalaiarasi Sonai Muthu Anbananthen

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07336 2025-04-11 cs.CV cs.AI 84%

Zeus: Zero-shot LLM Instruction for Union Segmentation in Multimodal Medical Imaging

Siyuan Dai, Kai Ye, Guodong Liu, Haoteng Tang, Liang Zhan

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 21 pages, 4 figures, In Press by a journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21839 2025-03-31 cs.CV cs.AI cs.LG 84%

M-DocSum: Do LVLMs Genuinely Comprehend Interleaved Image-Text in Document Summarization?

Haolong Yan, Kaijun Tan, Yeqing Shen, Xin Huang, Zheng Ge, Xiangyu Zhang, Si Li, Daxin Jiang

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09220 2025-03-26 cs.CL cs.AI 84%

LingYi: Medical Conversational Question Answering System based on Multi-modal Knowledge Graphs

Fei Xia, Bin Li, Yixuan Weng, Shizhu He, Kang Liu, Bin Sun, Shutao Li, Jun Zhao

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09091 2025-03-21 cs.CV cs.AI 84%

Multi-Modal Foundation Models for Computational Pathology: A Survey

Dong Li, Guihong Wan, Xintao Wu, Xinyu Wu, Xiaohui Chen, Yi He, Christine G. Lian, Peter K. Sorger, Yevgeniy R. Semenov, Chen Zhao

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04543 2025-03-07 cs.CL cs.AI 84%

Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model

Wenke Huang, Jian Liang, Xianda Guo, Yiyang Fang, Guancheng Wan, Xuankun Rong, Chi Wen, Zekun Shi, Qingyun Li, Didi Zhu, Yanbiao Ma, Ke Liang, Bin Yang, He Li, Jiawei Shao, Mang Ye, Bo Du

专题命中 图文多模态 :multimodal(title);multi-modal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20277 2025-02-28 cs.CV cs.AI 84%

Explainable, Multi-modal Wound Infection Classification from Images Augmented with Generated Captions

Palawat Busaranuvong, Emmanuel Agu, Reza Saadati Fard, Deepak Kumar, Shefalika Gautam, Bengisu Tulu, Diane Strong

专题命中 图文多模态 :multi-modal(title);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01703 2025-02-03 cs.CL cs.AI cs.LG 84%

UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models

Sejoon Oh, Yiqiao Jin, Megha Sharma, Donghyun Kim, Eric Ma, Gaurav Verma, Srijan Kumar

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10489 2024-12-25 cs.CV cs.AI eess.SP 84%

CognitionCapturer: Decoding Visual Stimuli From Human EEG Signal With Multimodal Information

Kaifan Zhang, Lihuo He, Xin Jiang, Wen Lu, Di Wang, Xinbo Gao

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07292 2024-12-11 cs.MM cs.CL 84%

Multimodal Sentiment Analysis Based on Causal Reasoning

Fuhai Chen, Pengpeng Huang, Xuri Ge, Jie Huang, Zishuo Bao

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07112 2024-12-11 cs.CV cs.CL 84%

Maya: An Instruction Finetuned Multilingual Multimodal Model

Nahid Alam, Karthik Reddy Kanjula, Surya Guthikonda, Timothy Chung, Bala Krishna S Vegesna, Abhipsha Das, Anthony Susevski, Ryan Sze-Yin Chan, S M Iftekhar Uddin, Shayekh Bin Islam, Roshan Santhosh, Snegha A, Drishti Sharma, Chen Liu, Isha Chaturvedi, Genta Indra Winata, Ashvanth. S, Snehanshu Mukherjee, Alham Fikri Aji

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05939 2024-12-10 cs.CV cs.CL cs.LG 84%

Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models

Xiao Xu, Tianhao Niu, Yuxi Xie, Libo Qin, Wanxiang Che, Min-Yen Kan

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments A manuscript that should have been Arxived in May :)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21414 2024-10-30 cs.CL cs.AI 84%

CT2C-QA: Multimodal Question Answering over Chinese Text, Table and Chart

Bowen Zhao, Tianhao Cheng, Yuejie Zhang, Ying Cheng, Rui Feng, Xiaobo Zhang

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02026 2024-10-01 cs.CV cs.AI 84%

Multimodal-Enhanced Objectness Learner for Corner Case Detection in Autonomous Driving

Lixing Xiao, Ruixiao Shi, Xiaoyang Tang, Yi Zhou

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted to 2024 IEEE International Conference on Image Processing (ICIP) as oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19404 2024-09-23 cs.CV cs.CL 84%

EAMA : Entity-Aware Multimodal Alignment Based Approach for News Image Captioning

Junzhe Zhang, Huixuan Zhang, Xunjian Yin, Xiaojun Wan

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06731 2024-08-29 cs.CV cs.AI 84%

Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator

Henry Hengyuan Zhao, Pan Zhou, Mike Zheng Shou

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted by ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13461 2024-08-27 cs.CV cs.AI 84%

Probing the Robustness of Vision-Language Pretrained Models: A Multimodal Adversarial Attack Approach

Jiwei Guan, Tianyu Ding, Longbing Cao, Lei Pan, Chen Wang, Xi Zheng

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03149 2024-08-07 cs.CV cs.CL 84%

Leveraging Entity Information for Cross-Modality Correlation Learning: The Entity-Guided Multimodal Summarization

Yanghai Zhang, Ye Liu, Shiwei Wu, Kai Zhang, Xukai Liu, Qi Liu, Enhong Chen

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments In ACL-Findings 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13488 2024-07-19 cs.CV cs.MM 84%

Similarity over Factuality: Are we making progress on multimodal out-of-context misinformation detection?

Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, Panagiotis C. Petrantonakis

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12498 2024-07-18 cs.CL cs.CV 84%

Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning

Mustafa Dogan, Ilker Kesen, Iacer Calixto, Aykut Erdem, Erkut Erdem

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Preprint. 33 pages, 17 Figures, 3 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08521 2024-07-17 cs.CV cs.CL 84%

Emergent Visual-Semantic Hierarchies in Image-Text Representations

Morris Alper, Hadar Averbuch-Elor

专题命中 图文多模态 :image-text(title);multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL

Comments Accepted to ECCV 2024. Project page: https://hierarcaps.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08418 2024-07-15 cs.CV cs.AI 84%

OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Qingyun Li, Zhe Chen, Weiyun Wang, Wenhai Wang, Shenglong Ye, Zhenjiang Jin, Guanzhou Chen, Yinan He, Zhangwei Gao, Erfei Cui, Jiashuo Yu, Hao Tian, Jiasheng Zhou, Chao Xu, Bin Wang, Xingjian Wei, Wei Li, Wenjian Zhang, Bo Zhang, Pinlong Cai, Licheng Wen, Xiangchao Yan, Zhenxiang Li, Pei Chu, Yi Wang, Min Dou, Changyao Tian, Xizhou Zhu, Lewei Lu, Yushi Chen, Junjun He, Zhongying Tu, Tong Lu, Yali Wang, Limin Wang, Dahua Lin, Yu Qiao, Botian Shi, Conghui He, Jifeng Dai

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10883 2024-07-09 cs.CV cs.CR cs.MM 84%

Improving Adversarial Transferability of Vision-Language Pre-training Models through Collaborative Multimodal Interaction

Jiyuan Fu, Zhaoyu Chen, Kaixun Jiang, Haijing Guo, Jiafeng Wang, Shuyong Gao, Wenqiang Zhang

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.MM

Comments This work won first place in CVPR 2024 Workshop Challenge: Black-box Adversarial Attacks on Vision Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16862 2024-06-24 cs.CV cs.CL 84%

TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

Zhengqing Yuan, Zhaoxu Li, Weiran Huang, Yanfang Ye, Lichao Sun

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments Accepted by ICML workshop 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08637 2024-06-04 cs.CL cs.AI 84%

TextBind: Multi-turn Interleaved Multimodal Instruction-following in the Wild

Huayang Li, Siheng Li, Deng Cai, Longyue Wang, Lemao Liu, Taro Watanabe, Yujiu Yang, Shuming Shi

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI

Comments Findings of ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09989 2024-05-30 cs.CV cs.CL 84%

LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition

Jinyuan Li, Han Li, Di Sun, Jiahao Wang, Wenkun Zhang, Zan Wang, Gang Pan

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted to Findings of ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17753 2024-04-30 cs.CV cs.AI 84%

Leveraging Cross-Modal Neighbor Representation for Improved CLIP Classification

Chao Yi, Lu Ren, De-Chuan Zhan, Han-Jia Ye

专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏