arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2206.07981 2022-06-20 cs.CV 79%

Multi-scale Cooperative Multimodal Transformers for Multimodal Sentiment Analysis in Videos

Lianyang Ma, Yu Yao, Tao Liang, Tongliang Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.05515 2022-06-15 cs.CL 79%

CLMLF:A Contrastive Learning and Multi-Layer Fusion Method for Multimodal Sentiment Detection

Zhen Li, Bing Xu, Conghui Zhu, Tiejun Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to Findings of NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04401 2022-06-10 cs.CV 79%

Cross-modal Local Shortest Path and Global Enhancement for Visible-Thermal Person Re-Identification

Xiaohong Wang, Chaoqi Li, Xiangcai Ma

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12396 2022-05-26 cs.LG cs.CL 79%

Recipe2Vec: Multi-modal Recipe Representation Learning with Graph Neural Networks

Yijun Tian, Chuxu Zhang, Zhichun Guo, Yihong Ma, Ronald Metoyer, Nitesh V. Chawla

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted by IJCAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.05545 2022-05-12 eess.IV cs.CV cs.LG 79%

CNN-LSTM Based Multimodal MRI and Clinical Data Fusion for Predicting Functional Outcome in Stroke Patients

Nima Hatami, Tae-Hee Cho, Laura Mechtouff, Omer Faruk Eker, David Rousseau, Carole Frindel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 44th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.04235 2022-05-10 q-bio.NC cs.AI cs.HC 79%

Measuring Cognitive Workload Using Multimodal Sensors

Niraj Hirachan, Anita Mathews, Julio Romero, Raul Fernandez Rojas

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.01380 2022-05-04 cs.CV cs.LG eess.SP 79%

Deep Learning in Multimodal Remote Sensing Data Fusion: A Comprehensive Review

Jiaxin Li, Danfeng Hong, Lianru Gao, Jing Yao, Ke Zheng, Bing Zhang, Jocelyn Chanussot

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.13707 2022-05-02 cs.LG cs.AI 79%

Tag-assisted Multimodal Sentiment Analysis under Uncertain Missing Modalities

Jiandian Zeng, Tianyi Liu, Jiantao Zhou

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by SIGIR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12541 2022-04-28 eess.IV cs.CV cs.LG 79%

Multi stain graph fusion for multimodal integration in pathology

Chaitanya Dwivedi, Shima Nofallah, Maryam Pouryahya, Janani Iyer, Kenneth Leidal, Chuhan Chung, Timothy Watkins, Andrew Billin, Robert Myers, John Abel, Ali Behrooz

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12185 2022-04-27 cs.CV 79%

TranSiam: Fusing Multimodal Visual Features Using Transformer for Medical Image Segmentation

Xuejian Li, Shiqiang Ma, Jijun Tang, Fei Guo

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10760 2022-04-25 cs.CV 79%

iCAR: Bridging Image Classification and Image-text Alignment for Visual Recognition

Yixuan Wei, Yue Cao, Zheng Zhang, Zhuliang Yao, Zhenda Xie, Han Hu, Baining Guo

专题命中 多模态训练与对齐 :image-text(title,abstract);分类 cs.CV

Comments 22 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.06493 2022-04-22 cs.CV 79%

AutoAlign: Pixel-Instance Feature Aggregation for Multi-Modal 3D Object Detection

Zehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang, Qinghong Jiang, Feng Zhao, Bolei Zhou, Hang Zhao

专题命中 多模态训练与对齐 :multi-modal(title);cross-modal(abstract);分类 cs.CV

Comments Accepted to IJCAI2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.12698 2022-04-20 cs.CV 79%

Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling

Dat Huynh, Jason Kuen, Zhe Lin, Jiuxiang Gu, Ehsan Elhamifar

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.12367 2022-04-19 cs.CV 79%

Transformer-based Multimodal Information Fusion for Facial Expression Analysis

Wei Zhang, Feng Qiu, Suzhen Wang, Hao Zeng, Zhimeng Zhang, Rudong An, Bowen Ma, Yu Ding

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.03249 2022-04-13 cs.CV 79%

Multimodal Colored Point Cloud to Image Alignment

Noam Rotstein, Amit Bracha, Ron Kimmel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.05603 2022-04-05 eess.IV cs.CV 79%

Multi-Modal MRI Reconstruction Assisted with Spatial Alignment Network

Kai Xuan, Lei Xiang, Xiaoqian Huang, Lichi Zhang, Shu Liao, Dinggang Shen, Qian Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Final version, IEEE Transactions on Medical Imaging, code available at \url{https://github.com/woxuankai/SpatialAlignmentNetwork}

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.09707 2022-03-29 cs.SE cs.AI 79%

M2TS: Multi-Scale Multi-Modal Approach Based on Transformer for Source Code Summarization

Yuexiu Gao, Chen Lyu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted by ICPC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13411 2022-03-28 cs.RO cs.AI cs.LG cs.SY eess.SY 79%

Reshaping Robot Trajectories Using Natural Language Commands: A Study of Multi-Modal Data Alignment Using Transformers

Arthur Bucker, Luis Figueredo, Sami Haddadin, Ashish Kapoor, Shuang Ma, Rogerio Bonatti

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.11441 2022-03-23 cs.CV 79%

Multi-Modal Learning for AU Detection Based on Multi-Head Fused Transformers

Xiang Zhang, Lijun Yin

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Journal ref FG 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.08195 2022-03-17 cs.CV 79%

DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection

Yingwei Li, Adams Wei Yu, Tianjian Meng, Ben Caine, Jiquan Ngiam, Daiyi Peng, Junyang Shen, Bo Wu, Yifeng Lu, Denny Zhou, Quoc V. Le, Alan Yuille, Mingxing Tan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments CVPR 2022. 1st rank 3D detection method on Waymo Challenge Leaderboard: https://waymo.com/open/challenges/entry/?timestamp=1647356360224524&challenge=DETECTION_3D&emailId=5451f123-a0ea

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.08055 2022-03-16 cs.CL 79%

Modular and Parameter-Efficient Multimodal Fusion with Prompting

Sheng Liang, Mengjie Zhao, Hinrich Schütze

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to Findings of ACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.04586 2022-03-10 eess.IV cs.CV 79%

Multi-modal Brain Tumor Segmentation via Missing Modality Synthesis and Modality-level Attention Fusion

Ziqi Huang, Li Lin, Pujin Cheng, Linkai Peng, Xiaoying Tang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 6 pages, 5 figures, submitted to ICPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.00510 2022-03-03 eess.SP cs.CV cs.HC cs.NI 79%

Multi-Modal Recurrent Fusion for Indoor Localization

Jianyuan Yu, Pu, Wang, Toshiaki Koike-Akino, Philip V. Orlik

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 5 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.08080 2022-02-22 eess.IV cs.CV 79%

Deep multi-modal aggregation network for MR image reconstruction with auxiliary modality

Chun-Mei Feng, Huazhu Fu, Tianfei Zhou, Yong Xu, Ling Shao, David Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.11641 2022-02-15 cs.CV eess.IV 79%

Cross-Modal Distillation for RGB-Depth Person Re-Identification

Frank Hafner, Amran Bhuiyan, Julian F. P. Kooij, Eric Granger

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Journal ref Computer Vision and Image Understanding, 103352 (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.04327 2022-02-10 cs.CV cs.IR 79%

Anchor Graph Structure Fusion Hashing for Cross-Modal Similarity Search

Lu Wang, Jie Yang, Masoumeh Zareapoor, Zhonglong Zheng

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.10274 2022-01-26 cs.CL 79%

Multi-channel Attentive Graph Convolutional Network With Sentiment Fusion For Multimodal Sentiment Analysis

Luwei Xiao, Xingjiao Wu, Wen Wu, Jing Yang, Liang He

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.07520 2022-01-20 cs.CL 79%

CM3: A Causal Masked Multimodal Model of the Internet

Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.05996 2022-01-19 cs.CR cs.CV 79%

Hardware Implementation of Multimodal Biometric using Fingerprint and Iris

Tariq M Khan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.13782 2022-01-19 cs.LG cs.AI 79%

Multimodal Co-learning: Challenges, Applications with Datasets, Recent Advances and Future Directions

Anil Rahate, Rahee Walambe, Sheela Ramanna, Ketan Kotecha

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments This is published in Information Fusion Journal and published copy can be downloaded from https://authors.elsevier.com/c/1eIVS5a7-Gls0Y for 50 days. https://www.sciencedirect.com/science/article/pii/S1566253521002530

Journal ref Information Fusion 81(2022) 203-239

详情

展开后加载摘要…

URL PDF HTML 收藏