arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2211.02269 2022-11-07 cs.CL 79%

Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and Analysis

Changyuan Qiu, Winston Wu, Xinliang Frederick Zhang, Lu Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.01353 2022-11-03 eess.IV cs.CV cs.LG 79%

Fourier Disentangled Multimodal Prior Knowledge Fusion for Red Nucleus Segmentation in Brain MRI

Guanghui Fu, Gabriel Jimenez, Sophie Loizillon, Rosana El Jurdi, Lydia Chougar, Didier Dormont, Romain Valabregue, Ninon Burgos, Stéphane Lehéricy, Daniel Racoceanu, Olivier Colliot, the ICEBERG Study Group

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00207 2022-11-02 cs.CV 79%

GMF: General Multimodal Fusion Framework for Correspondence Outlier Rejection

Xiaoshui Huang, Wentao Qu, Yifan Zuo, Yuming Fang, Xiaowei Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE RAL

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.12079 2022-11-01 cs.CV 79%

Bridging the View Disparity Between Radar and Camera Features for Multi-modal Fusion 3D Object Detection

Taohua Zhou, Yining Shi, Junjie Chen, Kun Jiang, Mengmeng Yang, Diange Yang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 12 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.12078 2022-10-27 cs.LG cs.AI eess.SP 79%

Multimodal sensor data fusion for in-situ classification of animal behavior using accelerometry and GNSS data

Reza Arablouei, Ziwei Wang, Greg J. Bishop-Hurley, Jiajun Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.05406 2022-10-26 cs.IR cs.MM 79%

Disentangled Multimodal Representation Learning for Recommendation

Fan Liu, Huilin Chen, Zhiyong Cheng, Anan Liu, Liqiang Nie, Mohan Kankanhalli

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments IEEE Transactions on Multimedia (TMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13530 2022-10-25 cs.CV 79%

Multimodal Pre-training Based on Graph Attention Network for Document Understanding

Zhenrong Zhang, Jiefeng Ma, Jun Du, Licheng Wang, Jianshu Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.09550 2022-10-19 cs.CL 79%

Probing Cross-modal Semantics Alignment Capability from the Textual Perspective

Zheng Ma, Shi Zong, Mianzhi Pan, Jianbing Zhang, Shujian Huang, Xinyu Dai, Jiajun Chen

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CL

Comments Findings of EMNLP2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04843 2022-10-11 cs.LG cs.CV 79%

Multi-Modal Fusion by Meta-Initialization

Matthew T. Jackson, Shreshth A. Malik, Michael T. Matthews, Yousuf Mohamed-Ahmed

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.05124 2022-10-11 cs.CV 79%

PCNet: A Structure Similarity Enhancement Method for Multispectral and Multimodal Image Registration

Si-Yuan Cao, Beinan Yu, Lun Luo, Shu-Jie Chen, Chunguang Li, Hui-Liang Shen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 33 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.02252 2022-10-05 cs.CV 79%

Channel Exchanging Networks for Multimodal and Multitask Dense Image Prediction

Yikai Wang, Fuchun Sun, Wenbing Huang, Fengxiang He, Dacheng Tao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by TPAMI 2022. Code is available at https://github.com/yikaiw/CEN. arXiv admin note: text overlap with arXiv:2011.05005

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13822 2022-10-04 cs.CV 79%

TokenFlow: Rethinking Fine-grained Cross-modal Alignment in Vision-Language Retrieval

Xiaohan Zou, Changqiao Wu, Lele Cheng, Zhongyuan Wang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13801 2022-09-29 cs.CV 79%

Translation, Scale and Rotation: Cross-Modal Alignment Meets RGB-Infrared Vehicle Detection

Maoxun Yuan, Yinyan Wang, Xingxing Wei

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.09768 2022-09-21 cs.CL 79%

An Efficient End-to-End Transformer with Progressive Tri-modal Attention for Multi-modal Emotion Recognition

Yang Wu, Pai Peng, Zhenyu Zhang, Yanyan Zhao, Bing Qin

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.03765 2022-09-09 eess.SP cs.CV cs.HC cs.LG 79%

Self-Supervised Multimodal Fusion Transformer for Passive Activity Recognition

Armand K. Koupai, Mohammud J. Bocus, Raul Santos-Rodriguez, Robert J. Piechocki, Ryan McConville

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 9 pages, 7 figures, submitted to IET Wireless Sensor Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.02368 2022-09-07 cs.CV 79%

Finger Multimodal Feature Fusion and Recognition Based on Channel Spatial Attention

Jian Guo, Jiaxiang Tu, Hengyi Ren, Chong Han, Lijuan Sun

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.01728 2022-09-07 cs.AI 79%

Features Fusion Framework for Multimodal Irregular Time-series Events

Peiwang Tang, Xianchao Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.00979 2022-09-07 eess.IV cs.CV cs.LG 79%

Multimodal Information Fusion for Glaucoma and DR Classification

Yihao Li, Mostafa El Habib Daho, Pierre-Henri Conze, Hassan Al Hajj, Sophie Bonnin, Hugang Ren, Niranchana Manivannan, Stephanie Magazzeni, Ramin Tadayoni, Béatrice Cochener, Mathieu Lamard, Gwenolé Quellec

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted preprint for presentation at MICCAI-OMIA

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.06180 2022-08-19 cs.CV cs.RO 79%

Multi-modal Depression Estimation based on Sub-attentional Fusion

Ping-Cheng Wei, Kunyu Peng, Alina Roitberg, Kailun Yang, Jiaming Zhang, Rainer Stiefelhagen

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to ECCV 2022 ACVR Workshop. Code is publicly available at https://github.com/PingCheng-Wei/DepressionEstimation

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.05621 2022-08-12 cs.CV 79%

ARMANI: Part-level Garment-Text Alignment for Unified Cross-Modal Fashion Design

Xujie Zhang, Yu Sha, Michael C. Kampffmeyer, Zhenyu Xie, Zequn Jie, Chengwen Huang, Jianqing Peng, Xiaodan Liang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by ACMMM22

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.03845 2022-07-27 cs.CV eess.IV 79%

Multi-modal land cover mapping of remote sensing images using pyramid attention and gated fusion networks

Qinghui Liu, Michael Kampffmeyer, Robert Jenssen, Arnt-Børre Salberg

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 24 pages, 11 figures, submitted to IJRS

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.02904 2022-07-27 cs.CV 79%

Multimodal Object Detection via Probabilistic Ensembling

Yi-Ting Chen, Jinghao Shi, Zelin Ye, Christoph Mertz, Deva Ramanan, Shu Kong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments camera-ready with supplement for ECCV2022 (oral presentation); open-source code at https://github.com/Jamie725/RGBT-detection

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.08238 2022-07-22 cs.CV cs.LG 79%

AXM-Net: Implicit Cross-Modal Feature Alignment for Person Re-identification

Ammarah Farooq, Muhammad Awais, Josef Kittler, Syed Safwan Khalid

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments AAAI-2022 (Oral Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08721 2022-07-18 cs.CV 79%

Multimodal Token Fusion for Vision Transformers

Yikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang, Fuchun Sun, Yunhe Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.09324 2022-07-07 cs.CL 79%

Supervised Visual Attention for Simultaneous Multimodal Machine Translation

Veneta Haralampieva, Ozan Caglayan, Lucia Specia

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to Journal of Artificial Intelligence Research (JAIR)

Journal ref Journal of Artificial Intelligence Research 74 (2022) 1059-1089

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.14699 2022-07-01 cs.CV eess.IV 79%

Fast computation of mutual information in the frequency domain with applications to global multimodal image alignment

Johan Öfverstedt, Joakim Lindblad, Nataša Sladoje

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 7 pages, 4 figures, 2 tables. The article is under consideration at Pattern Recognition Letters

Journal ref Pattern Recognition Letters, Vol. 159, pp. 196-203, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.01681 2022-06-30 cs.CV 79%

SoloGAN: Multi-domain Multimodal Unpaired Image-to-Image Translation via a Single Generative Adversarial Network

Shihua Huang, Cheng He, Ran Cheng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments pages 14, 15 figures

Journal ref IEEE Transactions on Artificial Intelligence 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.13256 2022-06-28 cs.MM 79%

A Topic-Attentive Transformer-based Model For Multimodal Depression Detection

Yanrong Guo, Chenyang Zhu, Shijie Hao, Richang Hong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.12714 2022-06-28 cs.CV cs.CR cs.LG 79%

Defending Multimodal Fusion Models against Single-Source Adversaries

Karren Yang, Wan-Yi Lin, Manash Barman, Filipe Condessa, Zico Kolter

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.12489 2022-06-22 eess.IV cs.CV 79%

Multi-modal and frequency-weighted tensor nuclear norm for hyperspectral image denoising

Xiaozhen Xie, Sheng Liu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments This version modifies the list of authors, due to the changes of some authors' affiliations and the affiliations' requirements

详情

展开后加载摘要…

URL PDF HTML 收藏