arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2310.05462 2023-10-25 cs.CV 70%

AdaFuse: Adaptive Medical Image Fusion Based on Spatial-Frequential Cross Attention

Xianming Gu, Lihui Wang, Zeyu Deng, Ying Cao, Xingyu Huang, Yue-min Zhu

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.12964 2023-10-20 cs.CV 70%

Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment

Sarah Ibrahimi, Xiaohang Sun, Pichao Wang, Amanmeet Garg, Ashutosh Sanan, Mohamed Omar

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10913 2023-09-27 cs.CV cs.AI cs.CL cs.LG cs.MM 70%

Weakly-Supervised Visual-Textual Grounding with Semantic Prior Refinement

Davide Rigoni, Luca Parolari, Luciano Serafini, Alessandro Sperduti, Lamberto Ballan

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09329 2023-07-31 cs.CV 70%

Towards a performance analysis on pre-trained Visual Question Answering models for autonomous driving

Kaavya Rekanar, Ciarán Eising, Ganesh Sistu, Martin Hayes

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Journal ref Proceedings of the Irish Machine Vision and Image Processing Conference 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.15932 2023-07-19 cs.CV 70%

Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report Generation

Yaowei Li, Bang Yang, Xuxin Cheng, Zhihong Zhu, Hongxiang Li, Yuexian Zou

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 1)Reassessment of author contributions. 2)Try to solve the problem that Google Scholar does not display the all authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03623 2023-07-10 cs.CV cs.RO 70%

Robust Human Detection under Visual Degradation via Thermal and mmWave Radar Fusion

Kaiwen Cai, Qiyue Xia, Peize Li, John Stankovic, Chris Xiaoxuan Lu

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments To appear at the 2023 International Conference on Embedded Wireless Systems and Networks

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16762 2023-06-30 cs.CL 70%

Unified Language Representation for Question Answering over Text, Tables, and Images

Bowen Yu, Cheng Fu, Haiyang Yu, Fei Huang, Yongbin Li

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CL

Comments Findings of ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.05171 2023-06-14 cs.CV 70%

ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding

Le Xue, Mingfei Gao, Chen Xing, Roberto Martín-Martín, Jiajun Wu, Caiming Xiong, Ran Xu, Juan Carlos Niebles, Silvio Savarese

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01016 2023-06-05 cs.CL cs.AI cs.CV cs.LG cs.MM 70%

PV2TEA: Patching Visual Modality to Textual-Established Information Extraction

Hejie Cui, Rongmei Lin, Nasser Zalmout, Chenwei Zhang, Jingbo Shang, Carl Yang, Xian Li

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ACL 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03277 2023-05-08 cs.CV 70%

FM-ViT: Flexible Modal Vision Transformers for Face Anti-Spoofing

Ajian Liu, Zichang Tan, Zitong Yu, Chenxu Zhao, Jun Wan, Yanyan Liang, Zhen Lei, Du Zhang, Stan Z. Li, Guodong Guo

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.07549 2023-04-18 cs.CV 70%

MA-ViT: Modality-Agnostic Vision Transformers for Face Anti-Spoofing

Ajian Liu, Yanyan Liang

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 7 pages, 4 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08419 2023-03-21 cs.CV 70%

Multi Modal Facial Expression Recognition with Transformer-Based Fusion Networks and Dynamic Sampling

Jun-Hwa Kim, Namho Kim, Chee Sun Won

专题命中 多模态训练与对齐 :multi-modal(abstract);audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.00843 2023-03-08 cs.CV cs.RO 70%

Early or Late Fusion Matters: Efficient RGB-D Fusion in Vision Transformers for 3D Object Recognition

Georgios Tziafas, Hamidreza Kasaei

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Submitted IROS 23. Supplementary video here: https://youtu.be/L2gkDPkHsfo

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10511 2023-02-22 cs.CV 70%

MVFusion: Multi-View 3D Object Detection with Semantic-aligned Radar and Camera Fusion

Zizhang Wu, Guilian Chen, Yuanzhu Gan, Lei Wang, Jian Pu

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by ICRA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.06148 2023-02-14 cs.CV 70%

CoMAE: Single Model Hybrid Pre-training on Small-Scale RGB-D Datasets

Jiange Yang, Sheng Guo, Gangshan Wu, Limin Wang

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 10 pages(including references),5 figures, Accept by Thirty-Seventh AAAI Conference on Artificial Intelligence (AAAI 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.11543 2022-11-23 cs.CV 70%

Middle-level Fusion for Lightweight RGB-D Salient Object Detection

Nianchang Huang, Qiang Zhang, Jungong Han

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.08565 2022-11-17 cs.CV 70%

Using Auxiliary Information for Person Re-Identification -- A Tutorial Overview

Tharindu Fernando, Clinton Fookes, Sridha Sridharan, Dana Michalski

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments Preprint Submitted to Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.01762 2022-08-31 cs.CV 70%

Robust RGB-D Fusion for Saliency Detection

Zongwei Wu, Shriarulmozhivarman Gobichettipalayam, Brahim Tamadazte, Guillaume Allibert, Danda Pani Paudel, Cédric Demonceaux

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to 3DV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.06761 2022-08-16 cs.CV 70%

MAFNet: A Multi-Attention Fusion Network for RGB-T Crowd Counting

Pengyu Chen, Junyu Gao, Yuan Yuan, Qi Wang

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.14395 2022-07-29 cs.CV 70%

Single-Stream Multi-Level Alignment for Vision-Language Pretraining

Zaid Khan, Vijay Kumar BG, Xiang Yu, Samuel Schulter, Manmohan Chandraker, Yun Fu

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.08150 2022-07-19 cs.CV 70%

FashionViL: Fashion-Focused Vision-and-Language Representation Learning

Xiao Han, Licheng Yu, Xiatian Zhu, Li Zhang, Yi-Zhe Song, Tao Xiang

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04625 2022-06-10 cs.LG cs.CV eess.SP 70%

AttX: Attentive Cross-Connections for Fusion of Wearable Signals in Emotion Recognition

Anubhav Bhatti, Behnam Behinaein, Paul Hungler, Ali Etemad

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 13 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12035 2022-04-27 cs.LG cs.CV 70%

Information Fusion: Scaling Subspace-Driven Approaches

Sally Ghanem, Hamid Krim

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.01427 2022-03-16 cs.CV eess.IV 70%

Attention-based Dual Supervised Decoder for RGBD Semantic Segmentation

Yang Zhang, Yang Yang, Chenyun Xiong, Guodong Sun, Yanwen Guo

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.08162 2022-01-11 cs.CV 70%

Specificity-preserving RGB-D Saliency Detection

Tao Zhou, Deng-Ping Fan, Geng Chen, Yi Zhou, Huazhu Fu

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments This is an extensive version and has been accepted by Computational Visual Media

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.11710 2021-12-23 cs.CV 70%

Fusion of medical imaging and electronic health records with attention and multi-head machanisms

Cheng Jiang, Yihao Chen, Jianbo Chang, Ming Feng, Renzhi Wang, Jianhua Yao

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.09129 2021-12-17 cs.CV 70%

Decoupling and Recoupling Spatiotemporal Representation for RGB-D-based Motion Recognition

Benjia Zhou, Pichao Wang, Jun Wan, Yanyan Liang, Fan Wang, Du Zhang, Zhen Lei, Hao Li, Rong Jin

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments open sourced; codes and models are available:https://github.com/damo-cv/MotionRGBD; transformer-based method

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.05144 2021-12-13 cs.CV eess.IV 70%

Edge-aware Guidance Fusion Network for RGB Thermal Scene Parsing

Wujie Zhou, Shaohua Dong, Caie Xu, Yaguan Qian

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by AAAI2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.10747 2021-11-29 cs.CV 70%

MaIL: A Unified Mask-Image-Language Trimodal Network for Referring Image Segmentation

Zizhang Li, Mengmeng Wang, Jianbiao Mei, Yong Liu

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.08495 2021-10-19 cs.CV 70%

Hybrid Mutimodal Fusion for Dimensional Emotion Recognition

Ziyu Ma, Fuyan Ma, Bin Sun, Shutao Li

专题命中 多模态训练与对齐 :multimodal(abstract);audio-visual(abstract);分类 cs.CV

Comments 8 pages, 2 figures, accepted by ACM MM2021

详情

展开后加载摘要…

URL PDF HTML 收藏