arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4644 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4644 篇

2402.14767 2024-02-23 cs.CV 79%

DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models

Yuhang Cao, Pan Zhang, Xiaoyi Dong, Dahua Lin, Jiaqi Wang

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06659 2024-02-21 cs.CL 79%

WisdoM: Improving Multimodal Sentiment Analysis by Fusing Contextual World Knowledge

Wenbin Wang, Liang Ding, Li Shen, Yong Luo, Han Hu, Dacheng Tao

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12048 2024-02-20 cs.CL 79%

Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models

Didi Zhu, Zhongyi Sun, Zexi Li, Tao Shen, Ke Yan, Shouhong Ding, Kun Kuang, Chao Wu

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08670 2024-02-14 cs.AI 79%

Rec-GPT4V: Multimodal Recommendation with Large Vision-Language Models

Yuqing Liu, Yu Wang, Lichao Sun, Philip S. Yu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06092 2024-02-12 cs.CV cs.RO 79%

CLIP-Loc: Multi-modal Landmark Association for Global Localization in Object-based Maps

Shigemichi Matsuzaki, Takuma Sugino, Kazuhito Tanaka, Zijun Sha, Shintaro Nakaoka, Shintaro Yoshizawa, Kazuhiro Shintani

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 7 pages, 7 figures. Accepted to IEEE International Conference on Robotics and Automation (ICRA) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05472 2024-02-09 cs.CV 79%

Question Aware Vision Transformer for Multimodal Reasoning

Roy Ganz, Yair Kittenplon, Aviad Aberdam, Elad Ben Avraham, Oren Nuriel, Shai Mazor, Ron Litman

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17862 2024-02-01 cs.CV 79%

Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis

Jianing Li, Xi Nan, Ming Lu, Li Du, Shanghang Zhang

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 15 pages,version 1

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13307 2024-01-25 cs.CV 79%

ChatterBox: Multi-round Multimodal Referring and Grounding

Yunjie Tian, Tianren Ma, Lingxi Xie, Jihao Qiu, Xi Tang, Yuan Zhang, Jianbin Jiao, Qi Tian, Qixiang Ye

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 17 pages, 6 tables, 9 figurs. Code, data, and model are available at: https://github.com/sunsmarterjie/ChatterBox

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03851 2024-01-09 cs.CV q-bio.NC 79%

Aligned with LLM: a new multi-modal training paradigm for encoding fMRI activity in visual cortex

Shuxiao Ma, Linyuan Wang, Senbao Hou, Bin Yan

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01911 2024-01-05 cs.CV cs.LG 79%

Backdoor Attack on Unpaired Medical Image-Text Foundation Models: A Pilot Study on MedCLIP

Ruinan Jin, Chun-Yin Huang, Chenyu You, Xiaoxiao Li

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Paper Accepted at the 2nd IEEE Conference on Secure and Trustworthy Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01076 2024-01-04 cs.CL 79%

DialCLIP: Empowering CLIP as Multi-Modal Dialog Retriever

Zhichao Yin, Binyuan Hui, Min Yang, Fei Huang, Yongbin Li

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CL

Comments ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16602 2023-12-29 cs.CV 79%

Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Jiaxing Huang, Jingyi Zhang, Kai Jiang, Han Qiu, Shijian Lu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14465 2023-12-25 cs.CV 79%

FM-OV3D: Foundation Model-based Cross-modal Knowledge Blending for Open-Vocabulary 3D Detection

Dongmei Zhang, Chang Li, Ray Zhang, Shenghao Xie, Wei Xue, Xiaodong Xie, Shanghang Zhang

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2024. Code will be released at https://github.com/dmzhang0425/FM-OV3D.git

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08548 2023-12-15 cs.CV 79%

EVP: Enhanced Visual Perception using Inverse Multi-Attentive Feature Refinement and Regularized Image-Text Alignment

Mykola Lavreniuk, Shariq Farooq Bhat, Matthias Müller, Peter Wonka

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03777 2023-12-11 cs.CV 79%

On the Robustness of Large Multimodal Models Against Image Adversarial Attacks

Xuanming Cui, Alejandro Aparcedo, Young Kyun Jang, Ser-Nam Lim

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01191 2023-12-05 cs.CV 79%

Bootstrapping Interactive Image-Text Alignment for Remote Sensing Image Captioning

Cong Yang, Zuchao Li, Lefei Zhang

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00438 2023-12-04 cs.CV 79%

Dolphins: Multimodal Language Model for Driving

Yingzi Ma, Yulong Cao, Jiachen Sun, Marco Pavone, Chaowei Xiao

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments The project page is available at https://vlm-driver.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17583 2023-11-30 cs.CV 79%

CLIPC8: Face liveness detection algorithm based on image-text pairs and contrastive learning

Xu Liu, Shu Zhou, Yurong Song, Wenzhe Luo, Xin Zhang

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13847 2023-11-29 cs.CV cs.IT eess.IV math.IT 79%

Perceptual Image Compression with Cooperative Cross-Modal Side Information

Shiyu Qin, Bin Chen, Yujun Huang, Baoyi An, Tao Dai, Shu-Tao Xia

专题命中 图文多模态 :cross-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05425 2023-11-10 cs.CV 79%

Active Mining Sample Pair Semantics for Image-text Matching

Yongfeng Chena, Jin Liua, Zhijing Yang, Ruihan Chena, Junpeng Tan

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.20343 2023-11-06 cs.IR cs.MM 79%

Large Multi-modal Encoders for Recommendation

Zixuan Yi, Zijun Long, Iadh Ounis, Craig Macdonald, Richard Mccreadie

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13959 2023-10-27 cs.CV 79%

Dynamic MDETR: A Dynamic Multimodal Transformer Decoder for Visual Grounding

Fengyuan Shi, Ruopeng Gao, Weilin Huang, Limin Wang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) in October 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13398 2023-10-23 cs.CV 79%

OpenAnnotate3D: Open-Vocabulary Auto-Labeling System for Multi-modal 3D Data

Yijie Zhou, Likun Cai, Xianhui Cheng, Zhongxue Gan, Xiangyang Xue, Wenchao Ding

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments The source code will be released at https://github.com/Fudan-ProjectTitan/OpenAnnotate3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14539 2023-10-12 cs.CR cs.CL 79%

Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models

Erfan Shayegani, Yue Dong, Nael Abu-Ghazaleh

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05109 2023-10-10 cs.CV 79%

Lightweight In-Context Tuning for Multimodal Unified Models

Yixin Chen, Shuai Zhang, Boran Han, Jiaya Jia

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01936 2023-10-04 cs.CV 79%

Constructing Image-Text Pair Dataset from Books

Yamato Okamoto, Haruto Toyonaga, Yoshihisa Ijiri, Hirokatsu Kataoka

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2023 workshop, Towards the Next Generation of Computer Vision Datasets: General DataCentric Submission Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00108 2023-10-03 cs.LG cs.CV 79%

Practical Membership Inference Attacks Against Large-Scale Multi-Modal Models: A Pilot Study

Myeongseob Ko, Ming Jin, Chenguang Wang, Ruoxi Jia

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments International Conference on Computer Vision (ICCV) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18009 2023-09-26 cs.CV 79%

Multi-Modal Face Stylization with a Generative Prior

Mengtian Li, Yi Dong, Minxuan Lin, Haibin Huang, Pengfei Wan, Chongyang Ma

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12110 2023-09-22 cs.CV 79%

Exploiting CLIP-based Multi-modal Approach for Artwork Classification and Retrieval

Alberto Baldrati, Marco Bertini, Tiberio Uricchio, Alberto Del Bimbo

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments Proc. of Florence Heri-Tech 2022: The Future of Heritage Science and Technologies: ICT and Digital Heritage, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10206 2023-09-20 cs.CV 79%

Image-Text Pre-Training for Logo Recognition

Mark Hubenthal, Suren Kumar

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments 8 pages, 5 figures, 4 tables

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2023, pp. 1145-1154

详情

展开后加载摘要…

URL PDF HTML 收藏