arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6872 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6872 篇

2309.17104 2023-10-04 cs.CV 83%

Prototype-guided Cross-modal Completion and Alignment for Incomplete Text-based Person Re-identification

Tiantian Gong, Guodong Du, Junsheng Wang, Yongkang Ding, Liyan Zhang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Sorry, some collaborators do not agree to publish it on Arxiv, so please withdraw this paper

Journal ref ACM International Conference on Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.00379 2023-10-02 eess.IV cs.CV 83%

Improved Multimodal Fusion for Small Datasets with Auxiliary Supervision

Gregory Holste, Douwe van der Wal, Hans Pinckaers, Rikiya Yamashita, Akinori Mitani, Andre Esteva

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments IEEE ISBI 2023 (see http://2023.biomedicalimaging.org/en/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13266 2023-09-26 cs.RO cs.AI 83%

Robust Navigation with Cross-Modal Fusion and Knowledge Transfer

Wenzhe Cai, Guangran Cheng, Lingyue Kong, Lu Dong, Changyin Sun

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments Accepted by ICRA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12855 2023-09-25 eess.IV cs.CV cs.LG 83%

Cross-Modal Translation and Alignment for Survival Analysis

Fengtao Zhou, Hao Chen

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.02357 2023-09-19 cs.CL cs.AI cs.CV cs.LG cs.MM 83%

Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph Completion

Xiang Chen, Ningyu Zhang, Lei Li, Shumin Deng, Chuanqi Tan, Changliang Xu, Fei Huang, Luo Si, Huajun Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by SIGIR 2022. Fix a severe bug

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09369 2023-08-21 cs.CV 83%

Single Frame Semantic Segmentation Using Multi-Modal Spherical Images

Suresh Guttikonda, Jason Rambach

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted at WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07732 2023-08-16 cs.CV 83%

UniTR: A Unified and Efficient Multi-Modal Transformer for Bird's-Eye-View Representation

Haiyang Wang, Hao Tang, Shaoshuai Shi, Aoxue Li, Zhenguo Li, Bernt Schiele, Liwei Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10469 2023-08-08 cs.CV 83%

Object Segmentation by Mining Cross-Modal Semantics

Zongwei Wu, Jingjing Wang, Zhuyun Zhou, Zhaochong An, Qiuping Jiang, Cédric Demonceaux, Guolei Sun, Radu Timofte

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.00919 2023-08-07 cs.CV 83%

Benchmarking Visual-Inertial Deep Multimodal Fusion for Relative Pose Regression and Odometry-aided Absolute Pose Regression

Felix Ott, Nisha Lakshmana Raichur, David Rügamer, Tobias Feigl, Heiko Neumann, Bernd Bischl, Christopher Mutschler

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10909 2023-07-31 cs.LG cs.AI 83%

Multi-modal Machine Learning in Engineering Design: A Review and Future Directions

Binyang Song, Rui Zhou, Faez Ahmed

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.13205 2023-07-26 cs.MM 83%

Text-oriented Modality Reinforcement Network for Multimodal Sentiment Analysis from Unaligned Multimodal Sequences

Yuxuan Lei, Dingkang Yang, Mingcheng Li, Shunli Wang, Jiawei Chen, Lihua Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

Comments Accepted by CICAI 2023 (Finalist of Best Student Paper Award)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15946 2023-06-29 cs.CV 83%

Knowledge-Enhanced Hierarchical Information Correlation Learning for Multi-Modal Rumor Detection

Jiawei Liu, Jingyi Xie, Fanrui Zhang, Qiang Zhang, Zheng-jun Zha

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.01311 2023-06-29 cs.LG cs.AI cs.CL cs.CV cs.MM 83%

High-Modality Multimodal Transformer: Quantifying Modality & Interaction Heterogeneity for High-Modality Representation Learning

Paul Pu Liang, Yiwei Lyu, Xiang Fan, Jeffrey Tsaw, Yudong Liu, Shentong Mo, Dani Yogatama, Louis-Philippe Morency, Ruslan Salakhutdinov

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments TMLR 2023, Code available at https://github.com/pliang279/HighMMT

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03810 2023-06-07 cs.CV cs.RO 83%

X-Align++: cross-modal cross-view alignment for Bird's-eye-view segmentation

Shubhankar Borse, Senthil Yogamani, Marvin Klingner, Varun Ravi, Hong Cai, Abdulaziz Almuzairee, Fatih Porikli

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted for publication at Springer Machine Vision and Applications Journal. The Version of Record of this article is published in Machine Vision and Applications Journal, and is available online at https://doi.org/10.1007/s00138-023-01400-7. arXiv admin note: substantial text overlap with arXiv:2210.06778

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02050 2023-06-07 cs.LG cs.CV 83%

Provable Dynamic Fusion for Low-Quality Multimodal Data

Qingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu, Huazhu Fu, Joey Tianyi Zhou, Xi Peng

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by ICML 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.06306 2023-05-16 cs.CV 83%

Efficient Multimodal Fusion via Interactive Prompting

Yaowei Li, Ruijie Quan, Linchao Zhu, Yi Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Camera-ready version for CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07713 2023-05-16 cs.CV 83%

Multi-Modal 3D Object Detection by Box Matching

Zhe Liu, Xiaoqing Ye, Zhikang Zou, Xinwei He, Xiao Tan, Errui Ding, Jingdong Wang, Xiang Bai

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06179 2023-05-12 cs.CV cs.RO 83%

A Multi-modal Approach to Single-modal Visual Place Classification

Tomoya Iwasaki, Kanji Tanaka, Kenta Tsukahara

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments 7 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13130 2023-04-27 cs.CV 83%

Hypernymization of named entity-rich captions for grounding-based multi-modal pretraining

Giacomo Nebbia, Adriana Kovashka

专题命中 多模态训练与对齐 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.05699 2023-04-13 cs.IR cs.MM 83%

Multimodal Matching-aware Co-attention Networks with Mutual Knowledge Distillation for Fake News Detection

Linmei Hu, Ziwang Zhao, Weijian Qi, Xuemeng Song, Liqiang Nie

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.00719 2023-04-04 cs.CV 83%

Multi-Modal Representation Learning with Text-Driven Soft Masks

Jaeyoo Park, Bohyung Han

专题命中 多模态训练与对齐 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08600 2023-03-16 cs.CV 83%

MSeg3D: Multi-modal 3D Semantic Segmentation for Autonomous Driving

Jiale Li, Hang Dai, Hao Han, Yong Ding

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to CVPR 2023 (preprint)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.03595 2023-03-15 cs.CV 83%

LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross-Modal Fusion

Xin Li, Tao Ma, Yuenan Hou, Botian Shi, Yuchen Yang, Youquan Liu, Xingjiao Wu, Qin Chen, Yikang Li, Yu Qiao, Liang He

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07027 2023-03-03 eess.IV cs.CV cs.LG 83%

MedFuse: Multi-modal fusion with clinical time-series data and chest X-ray images

Nasir Hayat, Krzysztof J. Geras, Farah E. Shamout

专题命中 多模态训练与对齐 :multi-modal(title,abstract);audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.14264 2023-03-01 cs.RO cs.CV 83%

RGB-D Grasp Detection via Depth Guided Learning with Cross-modal Attention

Ran Qin, Haoxiang Ma, Boyang Gao, Di Huang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted at ICRA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08052 2023-02-17 cs.CV 83%

Hierarchical Cross-modal Transformer for RGB-D Salient Object Detection

Hao Chen, Feihong Shen

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 10 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.02363 2023-02-17 cs.CV 83%

CAVER: Cross-Modal View-Mixed Transformer for Bi-Modal Salient Object Detection

Youwei Pang, Xiaoqi Zhao, Lihe Zhang, Huchuan Lu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted by TIP-2023. Add more details and update the weight illustration

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.09174 2023-01-24 cs.CV cs.HC cs.LG 83%

MATT: Multimodal Attention Level Estimation for e-learning Platforms

Roberto Daza, Luis F. Gomez, Aythami Morales, Julian Fierrez, Ruben Tolosana, Ruth Cobos, Javier Ortega-Garcia

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Preprint of the paper presented to the Workshop on Artificial Intelligence for Education (AI4EDU) of AAAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.04856 2023-01-13 cs.CL cs.LG stat.ML 83%

Multimodal Deep Learning

Cem Akkus, Luyang Chu, Vladana Djakovic, Steffen Jauch-Walser, Philipp Koch, Giacomo Loss, Christopher Marquardt, Marco Moldovan, Nadja Sauter, Maximilian Schneider, Rickmer Schulte, Karol Urbanczyk, Jann Goschenhofer, Christian Heumann, Rasmus Hvingelby, Daniel Schalk, Matthias Aßenmacher

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.02445 2023-01-12 cs.AI cs.LG 83%

IMKGA-SM: Interpretable Multimodal Knowledge Graph Answer Prediction via Sequence Modeling

Yilin Wen, Biao Luo, Yuqian Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments 12pages,10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏