arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

1711.08681 2017-11-27 cs.NE cs.CV 79%

Beyond RGB: Very High Resolution Urban Remote Sensing With Multimodal Deep Networks

Nicolas Audebert, Bertrand Le Saux, Sébastien Lefèvre

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments ISPRS Journal of Photogrammetry and Remote Sensing, Elsevier, A Para{î}tre

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.05516 2017-11-23 cs.CL 79%

Investigating Inner Properties of Multimodal Representation and Semantic Compositionality with Brain-based Componential Semantics

Shaonan Wang, Jiajun Zhang, Nan Lin, Chengqing Zong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments To appear in AAAI-18

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.00003 2017-11-02 cs.CV 79%

Common Representation Learning Using Step-based Correlation Multi-Modal CNN

Gaurav Bhatt, Piyush Jha, Balasubramanian Raman

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted in Asian Conference of Pattern Recognition (ACPR-2017)

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.02116 2017-08-09 cs.MM 79%

CCL: Cross-modal Correlation Learning with Multi-grained Fusion by Hierarchical Network

Yuxin Peng, Jinwei Qi, Xin Huang, Yuxin Yuan

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.MM

Comments 16 pages, accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.01471 2017-08-07 cs.CV 79%

Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering

Zhou Yu, Jun Yu, Jianping Fan, Dacheng Tao

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments ICCV 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.07250 2017-07-25 cs.CL 79%

Tensor Fusion Network for Multimodal Sentiment Analysis

Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, Louis-Philippe Morency

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted as full paper in EMNLP 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.04256 2017-06-15 cs.CV 79%

Online Convolutional Dictionary Learning for Multimodal Imaging

Kevin Degraux, Ulugbek S. Kamilov, Petros T. Boufounos, Dehong Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.08624 2017-05-25 cs.CV 79%

VANETs Meet Autonomous Vehicles: A Multimodal 3D Environment Learning Approach

Yassine Maalej, Sameh Sorour, Ahmed Abdel-Rahim, Mohsen Guizani

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 7 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.06676 2017-05-23 cs.CV 79%

MUTAN: Multimodal Tucker Fusion for Visual Question Answering

Hedi Ben-younes, Rémi Cadene, Matthieu Cord, Nicolas Thome

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.07548 2017-04-26 cs.AI cs.LG stat.ML 79%

Semi-supervised Bayesian Deep Multi-modal Emotion Recognition

Changde Du, Changying Du, Jinpeng Li, Wei-long Zheng, Bao-liang Lu, Huiguang He

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.04853 2017-03-16 cs.CV 79%

Face Recognition using Multi-Modal Low-Rank Dictionary Learning

Homa Foroughi, Moein Shakeri, Nilanjan Ray, Hong Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1609.05119 2017-03-13 cs.CV 79%

Deep Impression: Audiovisual Deep Residual Networks for Multimodal Apparent Personality Trait Recognition

Yağmur Güçlütürk, Umut Güçlü, Marcel A. J. van Gerven, Rob van Lier

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.08918 2017-02-01 cs.CV 79%

Feature Selection based on PCA and PSO for Multimodal Medical Image Fusion using DTCWT

Padmavathi K, Mahima Bhat, Maya V Karki

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.06121 2017-01-24 cs.CV 79%

Multimodal Fusion via a Series of Transfers for Noise Removal

Chang-Hwan Son, Xiao-Ping Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1509.00244 2016-11-17 cs.CV 79%

Robust Face Recognition via Multimodal Deep Face Representation

Changxing Ding, Dacheng Tao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments To appear in IEEE Trans. Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.04322 2016-10-17 cs.CV 79%

Learning and Fusing Multimodal Features from and for Multi-task Facial Computing

Wei Li, Zhigang Zhu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments An experiment to feature fusion in deep learning

详情

展开后加载摘要…

URL PDF HTML 收藏
1510.03519 2016-07-04 cs.CL 79%

Bridge Correlational Neural Networks for Multilingual Multimodal Representation Learning

Janarthanan Rajendran, Mitesh M. Khapra, Sarath Chandar, Balaraman Ravindran

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Published at NAACL-HLT 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1604.03443 2016-04-13 cs.CV 79%

Multi-modal Fusion for Diabetes Mellitus and Impaired Glucose Regulation Detection

Jinxing Li, David Zhang, Yongcheng Li, Jian Wu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 9 pages, 8 figures, 30 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
1502.01094 2016-01-20 stat.ML cs.CV cs.LG 79%

Multimodal Task-Driven Dictionary Learning for Image Classification

Soheil Bahrampour, Nasser M. Nasrabadi, Asok Ray, W. Kenneth Jenkins

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments To appear at IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
1502.07432 2015-11-19 cs.CV physics.data-an 79%

Coercive Region-level Registration for Multi-modal Images

Yu-Hui Chen, Dennis Wei, Gregory Newstadt, Jeffrey Simmons, Alfred Hero

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments This work has been accepted to International Conference on Image Processing (ICIP) 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1501.00102 2015-07-21 cs.CV cs.HC cs.LG 79%

ModDrop: adaptive multi-modal gesture recognition

Natalia Neverova, Christian Wolf, Graham W. Taylor, Florian Nebout

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1411.3229 2015-01-14 cs.CV 79%

Multi-modal Image Registration for Correlative Microscopy

Tian Cao, Christopher Zach, Shannon Modla, Debbie Powell, Kirk Czymmek, Marc Niethammer

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1011.6220 2010-11-30 cs.AI 79%

Multimodal Biometric Systems - Study to Improve Accuracy and Performance

K. Sasidhar, Vijaya L Kakulapati, Kolikipogu Ramakrishna, K. KailasaRao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 8 pages,5 figures, published in International Journal of Computer Science & Engineering Survey (IJCSES) Vol.1, No.2, November 2010

详情

展开后加载摘要…

URL PDF HTML 收藏
0909.4280 2009-12-01 cs.CL 79%

Towards Multimodal Content Representation

Harry Bunt, Laurent Romary

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Colloque avec actes et comité de lecture. internationale

Journal ref LREC Workshop on International Standards of Terminology and Language Resources Management, Las Palams : Spain (2002)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15127 2025-11-19 cs.LG 79%

PRIMUS: Pretraining IMU Encoders with Multimodal Self-Supervision

Arnav M. Das, Chi Ian Tang, Fahim Kawsar, Mohammad Malekzadeh

机构 * Nokia Bell Labs Cambridge, UK(诺基亚贝尔实验室(剑桥,英国)) University of Washington, USA(华盛顿大学(美国)) University of Glasgow, UK(格拉斯哥大学(英国))

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Presented at ICASSP 2025. Also presented under the title "PRIMUS: Pretraining IMU Encoders with Multimodal and Self-Supervised Learning" at NeurIPS 2024 TSALM Workshop (Time Series in the Age of Large Models)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13637 2025-11-18 cs.LG 79%

Towards Multimodal Representation Learning in Paediatric Kidney Disease

Ana Durica, John Booth, Ivana Drobnjak

机构 * Institute of Health Informatics(健康信息学研究所) University College London(伦敦大学学院) Data Research, Innovation and Virtual Environments Unit(数据研究、创新与虚拟环境单位) Great Ormond Street Hospital(格雷特奥蒙德医院) Department of Computer Science(计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 4 pages, 3 figures. EurIPS 2025 Multimodal Representation Learning for Healthcare (MMRL4H) workshop paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21971 2025-11-05 cs.LG 79%

GRAM-DTI: adaptive multimodal representation learning for drug target interaction prediction

Feng Jiang, Amina Mollaysa, Hehuan Ma, Tommaso Mansi, Junzhou Huang, Mangal Prakash, Rui Liao

机构 * University of Texas at Arlington(德克萨斯理工大学) Johnson & Johnson Innovative Medicine(强生创新医学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(journal_ref)

Journal ref NeurIPS 2025 2nd Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16026 2025-11-03 cs.LG stat.AP 79%

A tutorial on discovering and quantifying the effect of latent causal sources of multimodal EHR data

Marco Barbero-Mota, Eric V. Strobl, John M. Still, William W. Stead, Thomas A. Lasko

机构 * Department of Biomedical Informatics Vanderbilt University Medical Center(生物医学信息学系范德堡大学医学中心) Department of Biomedical Informatics University of Pittsburgh(生物医学信息学系匹兹堡大学) Departments of Medicine & Biomedical Informatics Vanderbilt University Medical Center(医学与生物医学信息学系范德堡大学医学中心) Departments of Biomedical Informatics & Computer Science Vanderbilt University Medical Center & Vanderbilt University(生物医学信息学与计算机科学系范德堡大学医学中心及范德堡大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted at the 1st Multimodal Representation Learning for Healthcare EurIPS 2025 Workshop (https://multimodal-rep-learning-for-health.github.io/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05993 2025-10-24 cs.IR 79%

Efficient Multimodal Streaming Recommendation via Expandable Side Mixture-of-Experts

Yunke Qu, Liang Qu, Tong Chen, Quoc Viet Hung Nguyen, Hongzhi Yin

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Accepted to CIKM 2025. Code is available at https://github.com/qykcq/Efficient-Multimodal-Streaming-Recommendation-via-Expandable-Side-Mixture-of-Experts

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08936 2025-06-11 cs.LG 79%

BioLangFusion: Multimodal Fusion of DNA, mRNA, and Protein Language Models

Amina Mollaysa, Artem Moskale, Pushpak Pati, Tommaso Mansi, Mangal Prakash, Rui Liao

专题命中 多模态训练与对齐 :multimodal(title);cross-modal(abstract);multi-modal(comments)

Comments Proceedings of ICML 2025 Workshop on Multi-modal Foundation Proceedings of ICML 2025 Workshop on Multi-modal Foundation Proceedings of ICML 2025 Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏