arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2303.00403 2023-03-02 cs.CV cs.LG 79%

Can representation learning for multimodal image registration be improved by supervision of intermediate layers?

Elisabeth Wetzer, Joakim Lindblad, Nataša Sladoje

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 15 Pages + 9 Pages Appendix, 10 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.09318 2023-02-21 cs.LG cs.AI 79%

Effective Multimodal Reinforcement Learning with Modality Alignment and Importance Enhancement

Jinming Ma, Feng Wu, Yingfeng Chen, Xianpeng Ji, Yu Ding

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 10 pages, 12 figures, This article is an extended version of the Extended Abstract accepted by the International Conference on Autonomous Agents and Multi-Agent Systems (AAMAS-2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08670 2023-02-20 cs.CV cs.IR 79%

Cascaded information enhancement and cross-modal attention feature fusion for multispectral pedestrian detection

Yang Yang, Kaixiong Xu, Kaizheng Wang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08326 2023-02-17 cs.CL 79%

NUAA-QMUL-AIIT at Memotion 3: Multi-modal Fusion with Squeeze-and-Excitation for Internet Meme Emotion Analysis

Xiaoyu Guo, Jing Ma, Arkaitz Zubiaga

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.05744 2023-02-14 cs.CV 79%

Rethinking Vision Transformer and Masked Autoencoder in Multimodal Face Anti-Spoofing

Zitong Yu, Rizhao Cai, Yawen Cui, Xin Liu, Yongjian Hu, Alex Kot

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12097 2023-02-10 cs.IR cs.MM 79%

Enhancing Dyadic Relations with Homogeneous Graphs for Multimodal Recommendation

Hongyu Zhou, Xin Zhou, Lingzi Zhang, Zhiqi Shen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments 17 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.02277 2023-02-08 cs.CV 79%

R2FD2: Fast and Robust Matching of Multimodal Remote Sensing Image via Repeatable Feature Detector and Rotation-invariant Feature Descriptor

Bai Zhu, Chao Yang, Jinkun Dai, Jianwei Fan, Yuanxin Ye

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 33 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.00941 2023-02-08 cs.CV eess.IV eess.SP 79%

Unsupervised Multimodal Change Detection Based on Structural Relationship Graph Representation Learning

Hongruixuan Chen, Naoto Yokoya, Chen Wu, Bo Du

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08320 2023-02-03 cs.CV 79%

Autoencoders as Cross-Modal Teachers: Can Pretrained 2D Image Transformers Help 3D Representation Learning?

Runpei Dong, Zekun Qi, Linfeng Zhang, Junbo Zhang, Jianjian Sun, Zheng Ge, Li Yi, Kaisheng Ma

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.07502 2023-01-24 cs.LG cs.CV 79%

Multimodal Side-Tuning for Document Classification

Stefano Pio Zingaro, Giuseppe Lisanti, Maurizio Gabbrielli

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 2020 25th International Conference on Pattern Recognition (ICPR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.03033 2023-01-10 cs.CV 79%

RGB-T Multi-Modal Crowd Counting Based on Transformer

Zhengyi Liu, Wei Wu, Yacheng Tan, Guanghui Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Journal ref BMVC2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00933 2023-01-02 cs.CV 79%

Deep Multimodal Fusion for Generalizable Person Re-identification

Suncheng Xiang, Hao Chen, Wei Ran, Zefang Yu, Ting Liu, Dahong Qian, Yuzhuo Fu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.02862 2022-12-27 cs.CV 79%

Farewell to Mutual Information: Variational Distillation for Cross-Modal Person Re-Identification

Xudong Tian, Zhizhong Zhang, Shaohui Lin, Yanyun Qu, Yuan Xie, Lizhuang Ma

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2022 as Oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10030 2022-12-21 cs.AI 79%

InterMulti:Multi-view Multimodal Interactions with Text-dominated Hierarchical High-order Fusion for Emotion Analysis

Feng Qiu, Wanzeng Kong, Yu Ding

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 9 pages, 3 figures. arXiv admin note: text overlap with arXiv:2212.08661

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.11088 2022-12-21 cs.CV 79%

EPNet++: Cascade Bi-directional Fusion for Multi-Modal 3D Object Detection

Zhe Liu, Tengteng Huang, Bingling Li, Xiwu Chen, Xi Wang, Xiang Bai

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by TPAMI-2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16509 2022-12-20 q-bio.GN cs.AI cs.LG q-bio.BM stat.ML 79%

Multimodal Learning for Multi-Omics: A Survey

Sina Tabakhi, Mohammod Naimul Islam Suvon, Pegah Ahadian, Haiping Lu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 52 pages, 3 figures; Revised matrix factorization fusion section

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.07039 2022-12-15 cs.CV 79%

Multi-Modal Domain Fusion for Multi-modal Aerial View Object Classification

Sumanth Udupa, Aniruddh Sikdar, Suresh Sundaram

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 7 pages,2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04661 2022-12-12 cs.CV 79%

An Attention-based Multi-Scale Feature Learning Network for Multimodal Medical Image Fusion

Meng Zhou, Xiaolan Xu, Yuxuan Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 8 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.03943 2022-12-09 cs.CV 79%

Learning Polysemantic Spoof Trace: A Multi-Modal Disentanglement Network for Face Anti-spoofing

Kaicheng Li, Hongyu Yang, Binghui Chen, Pengyu Li, Biao Wang, Di Huang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.02444 2022-12-06 cs.CV physics.optics 79%

Multi-modal Non-line-of-sight Passive Imaging

Andre Beckus, Alexandru Tamasan, George K. Atia

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Journal ref IEEE Transactions on Image Processing, vol. 28, no. 7, pp. 3372-3382, July 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14042 2022-11-28 cs.LG cs.AI q-bio.BM 79%

Molecular Joint Representation Learning via Multi-modal Information

Tianyu Wu, Yang Tang, Qiyu Sun, Luolin Xiong

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.13309 2022-11-28 cs.CV cs.LG 79%

How do Cross-View and Cross-Modal Alignment Affect Representations in Contrastive Learning?

Thomas M. Hehn, Julian F. P. Kooij, Dariu M. Gavrila

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.11515 2022-11-24 cs.CY cs.CV 79%

Multimodal Dual Emotion with Fusion of Visual Sentiment for Rumor Detection

Ge Wang, Li Tan, Ziliang Shang, He Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments There is an error in the experimental recording process, and the results reported in the article need to be re-checked

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.06331 2022-11-24 cs.IR cs.CV cs.LG 79%

Multi-modal Embedding Fusion-based Recommender

Anna Wroblewska, Jacek Dabrowski, Michal Pastuszak, Andrzej Michalowski, Michal Daniluk, Barbara Rychalska, Mikolaj Wieczorek, Sylwia Sysko-Romanczuk

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 7 pages, 8 figures

Journal ref revised and improved version: Electronics MDPI - https://www.mdpi.com/2079-9292/11/9/1391

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12021 2022-11-23 cs.CV 79%

ViFi-Loc: Multi-modal Pedestrian Localization using GAN with Camera-Phone Correspondences

Hansi Liu, Kristin Dana, Marco Gruteser, Hongsheng Lu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.06168 2022-11-23 cs.CL 79%

Unimodal and Multimodal Representation Training for Relation Extraction

Ciaran Cooney, Rachel Heyburn, Liam Madigan, Mairead O'Cuinn, Chloe Thompson, Joana Cavadas

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments 11 pages, 3 figures, 30th Irish Conference on Artificial Intelligence and Cognitive Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.09807 2022-11-22 cs.CV 79%

Towards All-in-one Pre-training via Maximizing Multi-modal Mutual Information

Weijie Su, Xizhou Zhu, Chenxin Tao, Lewei Lu, Bin Li, Gao Huang, Yu Qiao, Xiaogang Wang, Jie Zhou, Jifeng Dai

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.00302 2022-11-22 cs.LG cs.MM 79%

Progressive Fusion for Multimodal Integration

Shiv Shankar, Laure Thompson, Madalina Fiterau

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.03390 2022-11-21 cs.LG cs.AI 79%

Geometric Multimodal Contrastive Representation Learning

Petra Poklukar, Miguel Vasco, Hang Yin, Francisco S. Melo, Ana Paiva, Danica Kragic

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments ICML 2022 Camera ready version (update)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.06080 2022-11-14 cs.RO cs.CV 79%

Multi-modal Fusion Technology based on Vehicle Information: A Survey

Yan Gong, Jianli Lu, Jiayi Wu, Wenzhuo Liu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏