arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2012.14740 2022-01-11 cs.CL 79%

LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, Lidong Zhou

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

Comments ACL 2021 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.07800 2022-01-04 cs.CV 79%

Multi-modal Visual Place Recognition in Dynamics-Invariant Perception Space

Lin Wu, Teng Wang, Changyin Sun

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Journal ref IEEE Signal Processing Letters 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.09379 2022-01-04 cs.CV cs.LG 79%

BM-NAS: Bilevel Multimodal Neural Architecture Search

Yihang Yin, Siyu Huang, Xiang Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.12792 2021-12-30 cs.LG cs.MM 79%

Understanding and Measuring Robustness of Multimodal Learning

Nishant Vishwamitra, Hongxin Hu, Ziming Zhao, Long Cheng, Feng Luo

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.12927 2021-12-28 cs.CV 79%

Learning Aligned Cross-Modal Representation for Generalized Zero-Shot Classification

Zhiyu Fang, Xiaobin Zhu, Chun Yang, Zheng Han, Jingyan Qin, Xu-Cheng Yin

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.04537 2021-12-16 eess.IV cs.CV 79%

Multimodal Representation Learning via Maximization of Local Mutual Information

Ruizhi Liao, Daniel Moyer, Miriam Cha, Keegan Quigley, Seth Berkowitz, Steven Horng, Polina Golland, William M. Wells

专题命中 多模态训练与对齐 :multimodal(title);image-text(abstract);分类 cs.CV

Comments In Proceedings of International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI), 2021

Journal ref In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 273-283. Springer, Cham, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.01530 2021-12-13 cs.CV 79%

High-resolution Depth Maps Imaging via Attention-based Hierarchical Multi-modal Fusion

Zhiwei Zhong, Xianming Liu, Junjun Jiang, Debin Zhao, Zhiwen Chen, Xiangyang Ji

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00775 2021-12-03 cs.CV 79%

Routing with Self-Attention for Multimodal Capsule Networks

Kevin Duarte, Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Samuel Thomas, Alexander Liu, David Harwath, James Glass, Hilde Kuehne, Mubarak Shah

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.13361 2021-11-29 cs.LG cs.AI 79%

Geometric Multimodal Deep Learning with Multi-Scaled Graph Wavelet Convolutional Network

Maysam Behmanesh, Peyman Adibi, Mohammad Saeed Ehsani, Jocelyn Chanussot

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.11992 2021-11-29 cs.CV cs.LG 79%

Sparse Fusion for Multimodal Transformers

Yi Ding, Alex Rich, Mason Wang, Noah Stier, Matthew Turk, Pradeep Sen, Tobias Höllerer

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 4 figures, 5 tables, Yi Ding and Alex Rich contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.09624 2021-11-19 cs.CV 79%

IMFNet: Interpretable Multimodal Fusion for Point Cloud Registration

Xiaoshui Huang, Wentao Qu, Yifan Zuo, Yuming Fang, Xiaowei Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.08703 2021-11-18 cs.CV cs.CR 79%

Benchmarking Quality-Dependent and Cost-Sensitive Score-Level Multimodal Biometric Fusion Algorithms

Norman Poh, Thirimachos Bourlai, Josef Kittler, Lorene Allano, Fernando Alonso-Fernandez, Onkar Ambekar, John Baker, Bernadette Dorizzi, Omolara Fatukasi, Julian Fierrez, Harald Ganster, Javier Ortega-Garcia, Donald Maurer, Albert Ali Salah, Tobias Scheidat, Claus Vielhauer

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Published at IEEE Transactions on Information Forensics and Security journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.08456 2021-11-17 cs.LG cs.AI 79%

Trustworthy Multimodal Regression with Mixture of Normal-inverse Gamma Distributions

Huan Ma, Zongbo Han, Changqing Zhang, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.05485 2021-11-11 cs.CV 79%

A Structure Feature Algorithm for Multi-modal Forearm Registration

Jiaxin Li, Yan Ding, Weizhong Zhang, Yifan Zhao, Lingxi Guo, Zhe Yang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.01623 2021-11-10 cs.CV eess.IV 79%

A Tri-attention Fusion Guided Multi-modal Segmentation Network

Tongxue Zhou, Su Ruan, Pierre Vera, Stéphane Canu

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments 33 pages, 11 figures, accepted by Pattern Recognition on 01 November 2021. arXiv admin note: substantial text overlap with arXiv:2102.03111

Journal ref Pattern Recognition 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.03833 2021-11-05 cs.CV 79%

Self-Supervised Model Adaptation for Multimodal Semantic Segmentation

Abhinav Valada, Rohit Mohan, Wolfram Burgard

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments A Live demo is available at http://deepscene.cs.uni-freiburg.de and the code as well as the models are available at https://github.com/DeepSceneSeg

Journal ref International Journal of Computer Vision (IJCV), Special Issue: Deep Learning for Robotic Vision, vol. 128, no. 5, pp. 1239-1285, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.15639 2021-11-01 cs.CV 79%

Multi-Task and Multi-Modal Learning for RGB Dynamic Gesture Recognition

Dinghao Fan, Hengjie Lu, Shugong Xu, Shan Cao

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.04538 2021-10-27 cs.LG cs.AI 79%

What Makes Multi-modal Learning Better than Single (Provably)

Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen, Hang Zhao, Longbo Huang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted to NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.00385 2021-10-05 cs.NE cs.AI cs.LG 79%

Neural Dependency Coding inspired Multimodal Fusion

Shiv Shankar

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.03618 2021-09-27 eess.IV cs.CV 79%

MRI-based Alzheimer's disease prediction via distilling the knowledge in multi-modal data

Hao Guan, Chaoyue Wang, Dacheng Tao

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.04945 2021-09-22 cs.CV 79%

Multi-modal Fusion for Single-Stage Continuous Gesture Recognition

Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted for publication in IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.12712 2021-09-21 cs.CL 79%

Can images help recognize entities? A study of the role of images for Multimodal NER

Shuguang Chen, Gustavo Aguilar, Leonardo Neves, Thamar Solorio

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to W-NUT 2021 at EMNLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.07922 2021-09-17 cs.CV 79%

M2RNet: Multi-modal and Multi-scale Refined Network for RGB-D Salient Object Detection

Xian Fang, Jinchao Zhu, Ruixun Zhang, Xiuli Shao, Hongpeng Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.05766 2021-09-15 cs.CL 79%

Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech Translation

Renjie Zheng, Junkun Chen, Mingbo Ma, Liang Huang

专题命中 多模态训练与对齐 :multimodal(title);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.02804 2021-09-08 cs.CV 79%

Deep Collaborative Multi-Modal Learning for Unsupervised Kinship Estimation

Guan-Nan Dong, Chi-Man Pun, Zheng Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.02271 2021-09-07 cs.SI cs.AI cs.DB 79%

MONITOR: A Multimodal Fusion Framework to Assess Message Veracity in Social Networks

Abderrazek Azri, Cécile Favre, Nouria Harbi, Jérôme Darmont, Camille Noûs

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 25th European Conference on Advances in Databases and Information Systems (ADBIS 2021), Aug 2021, Tartu, Estonia

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.13669 2021-08-31 cs.AI 79%

Bi-Bimodal Modality Fusion for Correlation-Controlled Multimodal Sentiment Analysis

Wei Han, Hui Chen, Alexander Gelbukh, Amir Zadeh, Louis-philippe Morency, Soujanya Poria

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at ICMI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.07229 2021-08-25 cs.IR cs.CL cs.LG 79%

BERTERS: Multimodal Representation Learning for Expert Recommendation System with Transformer

N. Nikzad-Khasmakhi, M. A. Balafar, M. Reza Feizi-Derakhshi, Cina Motamed

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.07422 2021-08-18 cs.CV 79%

Learning by Aligning: Visible-Infrared Person Re-identification using Cross-Modal Correspondences

Hyunjong Park, Sanghoon Lee, Junghyup Lee, Bumsub Ham

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments iccv 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.05009 2021-08-12 cs.CV 79%

Learning Deep Multimodal Feature Representation with Asymmetric Multi-layer Fusion

Yikai Wang, Fuchun Sun, Ming Lu, Anbang Yao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments ACMMM 2020 (2020.3)

详情

展开后加载摘要…

URL PDF HTML 收藏