arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45832 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4634 篇

2411.05023 2024-11-11 cs.CL cs.LG quant-ph 83%

Multimodal Quantum Natural Language Processing: A Novel Framework for using Quantum Methods to Analyse Real Data

Hala Hawashin

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL

Comments This thesis, awarded a distinction by the Department of Computer Science at University College London, was successfully defended by the author in September 2024 in partial fulfillment of the requirements for an MSc in Emerging Digital Technologies

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13046 2024-11-01 cs.CV 83%

MoVA: Adapting Mixture of Vision Experts to Multimodal Context

Zhuofan Zong, Bingqi Ma, Dazhong Shen, Guanglu Song, Hao Shao, Dongzhi Jiang, Hongsheng Li, Yu Liu

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22736 2024-10-31 cs.CL 83%

Constructing Multimodal Datasets from Scratch for Rapid Development of a Japanese Visual Language Model

Keito Sasagawa, Koki Maeda, Issa Sugiura, Shuhei Kurita, Naoaki Okazaki, Daisuke Kawahara

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08773 2024-10-29 cs.CV cs.AI cs.CL cs.MM 83%

Veagle: Advancements in Multimodal Representation Learning

Rajat Chawla, Arkajit Datta, Tushar Verma, Adarsh Jha, Anmol Gautam, Ayush Vatsal, Sukrit Chaterjee, Mukunda NS, Ishaan Bhola

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13120 2024-10-17 cs.CV cs.LG 83%

RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering

Yuduo Wang, Pedram Ghamisi

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Journal ref IEEE Trans. Geosci. Remote Sens., vol. 62, pp. 1-13, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11829 2024-10-16 cs.CV 83%

MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding

Yue Cao, Yangzhou Liu, Zhe Chen, Guangchen Shi, Wenhai Wang, Danhuai Zhao, Tong Lu

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 11 pages, 6 figures, technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.06891 2024-10-01 cs.CV 83%

UKnow: A Unified Knowledge Protocol with Multimodal Knowledge Graph Datasets for Reasoning and Vision-Language Pre-Training

Biao Gong, Shuai Tan, Yutong Feng, Xiaoying Xie, Yuyuan Li, Chaochao Chen, Kecheng Zheng, Yujun Shen, Deli Zhao

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15310 2024-09-25 cs.LG cs.CV 83%

Visual Prompting in Multimodal Large Language Models: A Survey

Junda Wu, Zhehao Zhang, Yu Xia, Xintong Li, Zhaoyang Xia, Aaron Chang, Tong Yu, Sungchul Kim, Ryan A. Rossi, Ruiyi Zhang, Subrata Mitra, Dimitris N. Metaxas, Lina Yao, Jingbo Shang, Julian McAuley

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03412 2024-09-06 cs.CV physics.med-ph 83%

TG-LMM: Enhancing Medical Image Segmentation Accuracy through Text-Guided Large Multi-Modal Model

Yihao Zhao, Enhao Zhong, Cuiyun Yuan, Yang Li, Man Zhao, Chunxia Li, Jun Hu, Chenbin Liu

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments 11 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.08395 2024-09-05 cs.CV cs.AI cs.CL cs.LG cs.MM 83%

A Multimodal Memes Classification: A Survey and Open Research Issues

Tariq Habib Afridi, Aftab Alam, Muhammad Numan Khan, Jawad Khan, Young-Koo Lee

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments This is a survey paper on recent state of the art VL models that can be used for memes classification. it has 15 pages and 2 figures

Journal ref SCA 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15556 2024-08-29 cs.CV 83%

Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Wenbin Wang, Liang Ding, Minyan Zeng, Xiabin Zhou, Li Shen, Yong Luo, Dacheng Tao

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14594 2024-08-28 cs.CV 83%

MMR: Evaluating Reading Ability of Large Multimodal Models

Jian Chen, Ruiyi Zhang, Yufan Zhou, Ryan Rossi, Jiuxiang Gu, Changyou Chen

专题命中 图文多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18626 2024-07-29 cs.CL cs.AI cs.CV cs.DL cs.MM 83%

Every Part Matters: Integrity Verification of Scientific Figures Based on Multimodal Large Language Models

Xiang Shi, Jiawei Liu, Yinpeng Liu, Qikai Cheng, Wei Lu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 28 pages, 11 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12709 2024-07-18 cs.CV 83%

MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

Leyang Shen, Gongwei Chen, Rui Shao, Weili Guan, Liqiang Nie

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Github: https://github.com/JiuTian-VL/MoME

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.00219 2024-07-12 cs.CV 83%

Multi-modal Attribute Prompting for Vision-Language Models

Xin Liu, Jiamin Wu, and Wenfei Yang, Xu Zhou, Tianzhu Zhang

专题命中 图文多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted for Publication in IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01884 2024-07-03 cs.CV cs.HC 83%

EIT-1M: One Million EEG-Image-Text Pairs for Human Visual-textual Recognition and More

Xu Zheng, Ling Wang, Kanghao Chen, Yuanhuiyi Lyu, Jiazhou Zhou, Lin Wang

专题命中 图文多模态 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00203 2024-07-02 cs.CV 83%

PathGen-1.6M: 1.6 Million Pathology Image-text Pairs Generation through Multi-agent Collaboration

Yuxuan Sun, Yunlong Zhang, Yixuan Si, Chenglu Zhu, Zhongyi Shui, Kai Zhang, Jingxiong Li, Xingheng Lyu, Tao Lin, Lin Yang

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18579 2024-06-28 cs.CV cs.IR 83%

Hire: Hybrid-modal Interaction with Multiple Relational Enhancements for Image-Text Matching

Xuri Ge, Fuhai Chen, Songpei Xu, Fuxiang Tao, Jie Wang, Joemon M. Jose

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 22pages, 5 Figures, 6 tables, the extension of CMSEI in WACV23, and submitted to ACM TIST. arXiv admin note: text overlap with arXiv:2210.08908

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10954 2024-05-21 cs.CV 83%

Multimodal CLIP Inference for Meta-Few-Shot Image Classification

Constance Ferragu, Philomene Chagniot, Vincent Coyette

专题命中 图文多模态 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18171 2024-04-10 cs.CV cs.LG 83%

Improved Probabilistic Image-Text Representations

Sanghyuk Chun

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

Comments ICLR 2024 camera-ready; Code: https://github.com/naver-ai/pcmepp. Project page: https://naver-ai.github.io/pcmepp/. 30 pages, 2.2 MB

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08498 2024-04-09 cs.CV 83%

Extending CLIP's Image-Text Alignment to Referring Image Segmentation

Seoyeon Kim, Minguk Kang, Dongwon Kim, Jaesik Park, Suha Kwak

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

Comments NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01700 2024-04-04 cs.CV 83%

MotionChain: Conversational Motion Controllers via Multimodal Prompts

Biao Jiang, Xin Chen, Chi Zhang, Fukun Yin, Zhuoyuan Li, Gang YU, Jiayuan Fan

专题命中 图文多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 14 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10185 2024-03-27 cs.CV 83%

ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights

Weixian Lei, Yixiao Ge, Jianfeng Zhang, Dylan Sun, Kun Yi, Ying Shan, Mike Zheng Shou

专题命中 图文多模态 :omni-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments 19 pages, 4 figures and 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10887 2024-03-19 cs.CV 83%

LuoJiaHOG: A Hierarchy Oriented Geo-aware Image Caption Dataset for Remote Sensing Image-Text Retrival

Yuanxin Zhao, Mi Zhang, Bingnan Yang, Zhan Zhang, Jiaju Kang, Jianya Gong

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.02110 2024-03-12 cs.CV 83%

Sieve: Multimodal Dataset Pruning Using Image Captioning Models

Anas Mahmoud, Mostafa Elhoushi, Amro Abbas, Yu Yang, Newsha Ardalani, Hugh Leather, Ari Morcos

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted in CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04547 2024-03-08 cs.LG cs.AI 83%

CLIP the Bias: How Useful is Balancing Data in Multimodal Learning?

Ibrahim Alabdulmohsin, Xiao Wang, Andreas Steiner, Priya Goyal, Alexander D'Amour, Xiaohua Zhai

专题命中 图文多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments 32 pages, 20 figures, 7 tables

Journal ref ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03003 2024-03-06 cs.CV 83%

Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Gen Luo, Yiyi Zhou, Yuxin Zhang, Xiawu Zheng, Xiaoshuai Sun, Rongrong Ji

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15654 2024-02-27 cs.CL 83%

Exploring Failure Cases in Multimodal Reasoning About Physical Dynamics

Sadaf Ghaffari, Nikhil Krishnaswamy

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments 10 pages, 10 figures, Proceedings of AAAI Spring Symposium: Empowering Machine Learning and Large Language Models with Domain and Commonsense Knowledge (MAKE). AAAI (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08360 2024-02-14 cs.CV 83%

Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks

Jusung Lee, Sungguk Cha, Younghyun Lee, Cheoljong Yang

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04655 2024-02-07 cs.CR cs.CV 83%

VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models

Ziyi Yin, Muchao Ye, Tianrong Zhang, Tianyu Du, Jinguo Zhu, Han Liu, Jinghui Chen, Ting Wang, Fenglong Ma

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2023, 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏