arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45832 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4634 篇

2311.03356 2024-06-04 cs.CV cs.AI 81%

GLaMM: Pixel Grounding Large Multimodal Model

Hanoona Rasheed, Muhammad Maaz, Sahal Shaji Mullappilly, Abdelrahman Shaker, Salman Khan, Hisham Cholakkal, Rao M. Anwer, Erix Xing, Ming-Hsuan Yang, Fahad S. Khan

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14563 2024-05-24 cs.CV cs.AI 81%

Concept Visualization: Explaining the CLIP Multi-modal Embedding Using WordNet

Loris Giulivi, Giacomo Boracchi

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted for publication at IJCNN 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08964 2024-04-16 cs.CV cs.AI cs.LG 81%

Understanding Multimodal Deep Neural Networks: A Concept Selection View

Chenming Shang, Hengyuan Zhang, Hao Wen, Yujiu Yang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08856 2024-04-16 cs.CL cs.AI cs.LG 81%

On Speculative Decoding for Multimodal Large Language Models

Mukul Gagrani, Raghavv Goel, Wonseok Jeon, Junyoung Park, Mingu Lee, Christopher Lott

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted as a spotlight paper to ELVM workshop at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01322 2024-04-03 cs.CL cs.AI 81%

A Review of Multi-Modal Large Language and Vision Models

Kilian Carolan, Laura Fennelly, Alan F. Smeaton

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments 33 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01232 2024-04-03 cs.CL cs.CV 81%

Open-Vocabulary Federated Learning with Multimodal Prototyping

Huimin Zeng, Zhenrui Yue, Dong Wang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.18447 2024-03-28 cs.CL cs.CV cs.LG cs.RO 81%

Can Language Beat Numerical Regression? Language-Based Multimodal Trajectory Prediction

Inhwan Bae, Junoh Lee, Hae-Gon Jeon

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17556 2024-03-27 cs.CL cs.AI 81%

m3P: Towards Multimodal Multilingual Translation with Multimodal Prompt

Jian Yang, Hongcheng Guo, Yuwei Yin, Jiaqi Bai, Bing Wang, Jiaheng Liu, Xinnian Liang, Linzheng Cahi, Liqun Yang, Zhoujun Li

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14003 2024-03-22 cs.CV cs.CL cs.LG 81%

Multi-Modal Hallucination Control by Visual Information Grounding

Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pramuditha Perera, Alessandro Achille, Ashwin Swaminathan, Stefano Soatto

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

Journal ref IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08776 2024-03-15 cs.CV cs.AI 81%

Leveraging Chat-Based Large Vision Language Models for Multimodal Out-Of-Context Detection

Fatma Shalabi, Hichem Felouat, Huy H. Nguyen, Isao Echizen

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 13 pages, 6 figures , conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.09209 2024-03-12 cs.CV cs.CL 81%

Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation

Razvan-George Pasca, Alexey Gavryushin, Muhammad Hamza, Yen-Ling Kuo, Kaichun Mo, Luc Van Gool, Otmar Hilliges, Xi Wang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09778 2024-03-07 cs.CV cs.CL cs.LG 81%

Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models

Navid Rajabi, Jana Kosecka

专题命中 图文多模态 :multi-modal(title);image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted to DMLR @ ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02084 2024-02-27 cs.CV cs.CL cs.IR 81%

ITEm: Unsupervised Image-Text Embedding Learning for eCommerce

Baohao Liao, Michael Kozielski, Sanjika Hewavitharana, Jiangbo Yuan, Shahram Khadivi, Tomer Lancewicki

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05946 2024-02-27 cs.CV cs.CL 81%

OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network

Tiancheng Zhao, Peng Liu, Kyusong Lee

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Published at IET CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12058 2024-02-20 cs.CV cs.CL 81%

Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models

Xuanyu Lei, Zonghan Yang, Xinrui Chen, Peng Li, Yang Liu

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.02317 2024-01-24 cs.CL cs.CV 81%

Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings

Daniel Rose, Vaishnavi Himakunthala, Andy Ouyang, Ryan He, Alex Mei, Yujie Lu, Michael Saxon, Chinmay Sonar, Diba Mirza, William Yang Wang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11633 2024-01-23 cs.CV cs.AI 81%

Zoom-shot: Fast and Efficient Unsupervised Zero-Shot Transfer of CLIP to Vision Encoders with Multimodal Loss

Jordan Shipard, Arnold Wiliem, Kien Nguyen Thanh, Wei Xiang, Clinton Fookes

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.04751 2024-01-09 cs.CV cs.AI 81%

Multimodal Parameter-Efficient Few-Shot Class Incremental Learning

Marco D'Alessandro, Alberto Alonso, Enrique Calabrés, Mikel Galar

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Paris, France, 2023, pp. 3385-3395

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02347 2024-01-05 cs.CV cs.AI 81%

Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training

Longtian Qiu, Shan Ning, Xuming He

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV、cs.AI

Comments AAAI 2024.Open sourced, Code and Model Available

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11171 2023-12-19 cs.CV cs.AI 81%

UniDCP: Unifying Multiple Medical Vision-language Tasks via Dynamic Cross-modal Learnable Prompts

Chenlu Zhan, Yufei Zhang, Yu Lin, Gaoang Wang, Hongwei Wang

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09812 2023-12-18 cs.CV cs.AI 81%

Structural Information Guided Multimodal Pre-training for Vehicle-centric Perception

Xiao Wang, Wentao Wu, Chenglong Li, Zhicheng Zhao, Zhe Chen, Yukai Shi, Jin Tang

专题命中 图文多模态 :multimodal(title);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted by AAAI-2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08213 2023-11-15 cs.CV cs.CL 81%

Unlock the Power: Competitive Distillation for Multi-Modal Large Language Models

Xinwei Li, Li Lin, Shuai Wang, Chen Qian

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.01855 2023-11-14 cs.CV cs.AI 81%

Multimodal Data Augmentation for Image Captioning using Diffusion Models

Changrong Xiao, Sean Xin Xu, Kunpeng Zhang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.16741 2023-11-03 cs.AI cs.CV 81%

Socratis: Are large multimodal models emotionally aware?

Katherine Deng, Arijit Ray, Reuben Tan, Saadia Gabriel, Bryan A. Plummer, Kate Saenko

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments ICCV 2023 WECIA

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.20159 2023-11-01 cs.CV cs.AI 81%

Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts

Deepanway Ghosal, Navonil Majumder, Roy Ka-Wei Lee, Rada Mihalcea, Soujanya Poria

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.06939 2023-10-31 cs.CV cs.CL 81%

Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Wanrong Zhu, Jack Hessel, Anas Awadalla, Samir Yitzhak Gadre, Jesse Dodge, Alex Fang, Youngjae Yu, Ludwig Schmidt, William Yang Wang, Yejin Choi

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments NeurIPS D&B 2023. Project homepage: https://github.com/allenai/mmc4

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13812 2023-10-26 cs.CL cs.CV 81%

Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language Compositionality

Harman Singh, Pengchuan Zhang, Qifan Wang, Mengjiao Wang, Wenhan Xiong, Jingfei Du, Yu Chen

专题命中 图文多模态 :image-text(title);multimodal(abstract);分类 cs.CV、cs.CL

Comments EMNLP 2023 (long paper, main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14616 2023-10-25 cs.CL cs.CV 81%

Exploring Affordance and Situated Meaning in Image Captions: A Multimodal Analysis

Pin-Er Chen, Po-Ya Angela Wang, Hsin-Yu Chou, Yu-Hsiang Tseng, Shu-Kai Hsieh

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14581 2023-10-24 cs.CV cs.AI 81%

Leveraging Image-Text Similarity and Caption Modification for the DataComp Challenge: Filtering Track and BYOD Track

Shuhei Yokoo, Peifei Zhu, Yuchi Ishikawa, Mikihiro Tanaka, Masayoshi Kondo, Hirokatsu Kataoka

专题命中 图文多模态 :image-text(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at the ICCV 2023 Workshop on Towards the Next Generation of Computer Vision Datasets: DataComp Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11173 2023-10-18 cs.CV cs.AI 81%

Knowledge Extraction and Distillation from Large-Scale Image-Text Colonoscopy Records Leveraging Large Language and Vision Models

Shuo Wang, Yan Zhu, Xiaoyuan Luo, Zhiwei Yang, Yizhe Zhang, Peiyao Fu, Manning Wang, Zhijian Song, Quanlin Li, Pinghong Zhou, Yike Guo

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏