arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2308.11217 2023-08-25 cs.LG cs.AI 79%

Federated Learning in Big Model Era: Domain-Specific Multimodal Large Models

Zengxiang Li, Zhaoxiang Hou, Hui Liu, Ying Wang, Tongzhi Li, Longfei Xie, Chao Shi, Chengyi Yang, Weishan Zhang, Zelei Liu, Liang Xu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10741 2023-08-22 cs.LG cs.AI cs.CR 79%

On the Adversarial Robustness of Multi-Modal Foundation Models

Christian Schlarmann, Matthias Hein

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments ICCV AROW 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07971 2023-08-17 cs.CL cs.LG 79%

MultiSChuBERT: Effective Multimodal Fusion for Scholarly Document Quality Prediction

Gideon Maillette de Buy Wenniger, Thomas van Dongen, Lambert Schomaker

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06794 2023-08-16 cs.CV 79%

MMF-Track: Multi-modal Multi-level Fusion for 3D Single Object Tracking

Zhiheng Li, Yubo Cui, Yu Lin, Zheng Fang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 11 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06530 2023-08-15 cs.CV 79%

BEV-DG: Cross-Modal Learning under Bird's-Eye View for Domain Generalization of 3D Semantic Segmentation

Miaoyu Li, Yachao Zhang, Xu MA, Yanyun Qu, Yun Fu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04502 2023-08-15 cs.CL 79%

Revisiting Disentanglement and Fusion on Modality and Context in Conversational Multimodal Emotion Recognition

Bobo Li, Hao Fei, Lizi Liao, Yu Zhao, Chong Teng, Tat-Seng Chua, Donghong Ji, Fei Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03113 2023-08-08 cs.IR cs.MM 79%

Semantic-Guided Feature Distillation for Multimodal Recommendation

Fan Liu, Huilin Chen, Zhiyong Cheng, Liqiang Nie, Mohan Kankanhalli

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments ACM Multimedia 2023 Accepted

Journal ref In Proceedings of the 31st ACM International Conference on Multimedia (MM '23), 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01430 2023-08-04 cs.CL 79%

FinVis-GPT: A Multimodal Large Language Model for Financial Chart Analysis

Ziao Wang, Yuhang Li, Junda Wu, Jaehyeon Soon, Xiaofeng Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments (FinLLM 2023)@IJCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.12180 2023-07-25 eess.IV cs.CV cs.LG 79%

Prototype-Driven and Multi-Expert Integrated Multi-Modal MR Brain Tumor Image Segmentation

Yafei Zhang, Zhiyuan Li, Huafeng Li, Dapeng Tao

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09155 2023-07-19 cs.CV 79%

MLF-DET: Multi-Level Fusion for Cross-Modal 3D Object Detection

Zewei Lin, Yanqing Shen, Sanping Zhou, Shitao Chen, Nanning Zheng

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.08471 2023-07-18 cs.RO cs.AI 79%

Clarifying the Half Full or Half Empty Question: Multimodal Container Classification

Josua Spisak, Matthias Kerzel, Stefan Wermter

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Preprint for ICANN 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04907 2023-07-12 cs.CL cs.LG 79%

SimpleMTOD: A Simple Language Model for Multimodal Task-Oriented Dialogue with Symbolic Scene Representation

Bhathiya Hemanthage, Christian Dondrup, Phil Bartie, Oliver Lemon

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03591 2023-07-10 cs.AI cs.IR 79%

Structure Guided Multi-modal Pre-trained Transformer for Knowledge Graph Reasoning

Ke Liang, Sihang Zhou, Yue Liu, Lingyuan Meng, Meng Liu, Xinwang Liu

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessed

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03425 2023-07-10 cs.CV 79%

Registration-Free Hybrid Learning Empowers Simple Multimodal Imaging System for High-quality Fusion Detection

Yinghan Guan, Haoran Dai, Zekuan Yu, Shouyu Wang, Yuanjie Gu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00865 2023-07-06 cs.CV 79%

AMIGO: Sparse Multi-Modal Graph Transformer with Shared-Context Processing for Representation Learning of Giga-pixel Images

Ramin Nakhli, Puria Azadi Moghadam, Haoyang Mi, Hossein Farahani, Alexander Baras, Blake Gilks, Ali Bashashati

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.01486 2023-07-06 eess.IV cs.CV 79%

H-DenseFormer: An Efficient Hybrid Densely Connected Transformer for Multimodal Tumor Segmentation

Jun Shi, Hongyu Kan, Shulan Ruan, Ziqi Zhu, Minfan Zhao, Liang Qiao, Zhaohui Wang, Hong An, Xudong Xue

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 2 figures. This paper has been accepted by Medical Image Computing and Computer-Assisted Intervention(MICCAI) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16991 2023-06-30 cs.CV 79%

Integrating Large Pre-trained Models into Multimodal Named Entity Recognition with Evidential Fusion

Weide Liu, Xiaoyang Zhong, Jingwen Hou, Shaohua Li, Haozhe Huang, Yuming Fang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15765 2023-06-29 cs.CV 79%

A Novel Two Stream Decision Level Fusion of Vision and Inertial Sensors Data for Automatic Multimodal Human Activity Recognition System

Santosh Kumar Yadav, Muhtashim Rafiqi, Egna Praneeth Gummana, Kamlesh Tiwari, Hari Mohan Pandey, Shaik Ali Akbara

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11746 2023-06-22 cs.SI cs.MM 79%

Focusing on Relevant Responses for Multi-modal Rumor Detection

Jun Li, Yi Bin, Liang Peng, Yang Yang, Yangyang Li, Hao Jin, Zi Huang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.MM

Comments Submitted to TKDE

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.16952 2023-06-21 cs.CV cs.LG eess.IV 79%

Multimodal Fusion Transformer for Remote Sensing Image Classification

Swalpa Kumar Roy, Ankur Deria, Danfeng Hong, Behnood Rasti, Antonio Plaza, Jocelyn Chanussot

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Published in IEEE Transactions on Geoscience and Remote Sensing

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11020 2023-06-21 cs.CL 79%

Dual-Gated Fusion with Prefix-Tuning for Multi-Modal Relation Extraction

Qian Li, Shu Guo, Cheng Ji, Xutan Peng, Shiyao Cui, Jianxin Li

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.10146 2023-06-21 cs.CV eess.IV 79%

Multi-task 3D building understanding with multi-modal pretraining

Shicheng Xu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 8 pages, 9 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.05265 2023-06-21 cs.CV 79%

Multi-Sem Fusion: Multimodal Semantic Fusion for 3D Object Detection

Shaoqing Xu, Fang Li, Ziying Song, Jin Fang, Sifen Wang, Zhi-Xin Yang

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.03128 2023-06-16 cs.CV 79%

PointMCD: Boosting Deep Point Cloud Encoders via Multi-view Cross-modal Distillation for 3D Shape Recognition

Qijian Zhang, Junhui Hou, Yue Qian

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to TMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12423 2023-06-16 cs.LG cs.CV 79%

RadioPathomics: Multimodal Learning in Non-Small Cell Lung Cancer for Adaptive Radiotherapy

Matteo Tortora, Ermanno Cordelli, Rosa Sicilia, Lorenzo Nibid, Edy Ippolito, Giuseppe Perrone, Sara Ramella, Paolo Soda

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10773 2023-06-13 cs.CL 79%

MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

Zhiyang Xu, Ying Shen, Lifu Huang

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CL

Comments ACL 2023, dataset url: https://github.com/VT-NLP/MultiInstruct

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18721 2023-06-12 cs.CV 79%

LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document Understanding

Yi Tu, Ya Guo, Huan Chen, Jinyang Tang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ACL 2023 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02673 2023-06-06 eess.IV cs.CV cs.LG 79%

Cross-Modal Vertical Federated Learning for MRI Reconstruction

Yunlu Yan, Hong Wang, Yawen Huang, Nanjun He, Lei Zhu, Yuexiang Li, Yong Xu, Yefeng Zheng

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01523 2023-06-05 cs.CV cs.LG 79%

Transformer-based Multi-Modal Learning for Multi Label Remote Sensing Image Classification

David Hoffmann, Kai Norman Clasen, Begüm Demir

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at IEEE International Geoscience and Remote Sensing Symposium 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00792 2023-06-02 cs.CV 79%

Learning Across Decentralized Multi-Modal Remote Sensing Archives with Federated Learning

Barış Büyüktaş, Gencer Sumbul, Begüm Demir

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at IEEE International Geoscience and Remote Sensing Symposium (IGARSS) 2023. Our code is available at https://git.tu-berlin.de/rsim/MM-FL

详情

展开后加载摘要…

URL PDF HTML 收藏