arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6856 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6856 篇

2307.16121 2024-06-18 cs.CV cs.AI 84%

Uncertainty-Encoded Multi-Modal Fusion for Robust Object Detection in Autonomous Driving

Yang Lou, Qun Song, Qian Xu, Rui Tan, Jianping Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments In proceedings of the 26th European Conference on Artificial Intelligence ECAI 2023. 8 pages + 2 appendix pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12416 2024-06-17 cs.CV cs.CL 84%

Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning

Chong Ma, Hanqi Jiang, Wenting Chen, Yiwei Li, Zihao Wu, Xiaowei Yu, Zhengliang Liu, Lei Guo, Dajiang Zhu, Tuo Zhang, Dinggang Shen, Tianming Liu, Xiang Li

专题命中 多模态训练与对齐 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01920 2024-06-05 cs.CV cs.AI 84%

CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models

Junho Kim, Hyunjun Kim, Yeonju Kim, Yong Man Ro

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Project page: https://ivy-lvlm.github.io/CODE/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18751 2024-05-31 cs.CV cs.AI 84%

On the Limits of Multi-modal Meta-Learning with Auxiliary Task Modulation Using Conditional Batch Normalization

Jordi Armengol-Estapé, Vincent Michalski, Ramnath Kumar, Pierre-Luc St-Charles, Doina Precup, Samira Ebrahimi Kahou

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04883 2024-05-13 cs.CV cs.AI cs.LG 84%

FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion

Zehan Wang, Ziang Zhang, Xize Cheng, Rongjie Huang, Luping Liu, Zhenhui Ye, Haifeng Huang, Yang Zhao, Tao Jin, Peng Gao, Zhou Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICML 2024. The code and checkpoints will be released at https://github.com/zehanwang01/FreeBind

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13847 2024-04-23 cs.CV cs.CL 84%

EventLens: Leveraging Event-Aware Pretraining and Cross-modal Linking Enhances Visual Commonsense Reasoning

Mingjie Ma, Zhihuan Yu, Yichao Ma, Guohui Li

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12642 2024-04-22 cs.CL cs.CV 84%

Cooperative Sentiment Agents for Multimodal Sentiment Analysis

Shanmin Wang, Hui Shuai, Qingshan Liu, Fei Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10144 2024-04-11 cs.LG cs.AI cs.CV 84%

Data-Efficient Multimodal Fusion on a Single GPU

Noël Vouitsis, Zhaoyan Liu, Satya Krishna Gorti, Valentin Villecroze, Jesse C. Cresswell, Guangwei Yu, Gabriel Loaiza-Ganem, Maksims Volkovs

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments CVPR 2024 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00968 2024-04-04 cs.CV cs.CL 84%

Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts

Jialin Wu, Xia Hu, Yaqing Wang, Bo Pang, Radu Soricut

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12987 2024-04-02 cs.CL cs.LG cs.SD eess.AS 84%

TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation

Taeyang Yun, Hyunkuk Lim, Jeonghwan Lee, Min Song

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、eess.AS

Comments NAACL 2024 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.20026 2024-04-01 cs.CV cs.CL 84%

FSMR: A Feature Swapping Multi-modal Reasoning Approach with Joint Textual and Visual Clues

Shuang Li, Jiahua Wang, Lijie Wen

专题命中 多模态训练与对齐 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15444 2024-03-26 eess.SP cs.AI cs.CV cs.LG eess.IV 84%

A Survey of IMU Based Cross-Modal Transfer Learning in Human Activity Recognition

Abhi Kamboj, Minh Do

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02328 2024-02-12 cs.MM cs.CL 84%

Enhancing Multimodal Entity and Relation Extraction with Variational Information Bottleneck

Shiyao Cui, Jiangxia Cao, Xin Cong, Jiawei Sheng, Quangang Li, Tingwen Liu, Jinqiao Shi

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL、cs.MM

Journal ref IEEE/ACM Transactions on Audio, Speech and Language Processing, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02137 2024-01-05 cs.CV cs.AI 84%

SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment

Ziping Ma, Furong Xu, Jian Liu, Ming Yang, Qingpei Guo

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.08372 2024-01-05 cs.CL cs.MM 84%

Hierarchical Aligned Multimodal Learning for NER on Tweet Posts

Peipei Liu, Hong Li, Yimo Ren, Jie Liu, Shuaizong Si, Hongsong Zhu, Limin Sun

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11023 2023-12-19 cs.MM cs.AI 84%

Frequency Spectrum is More Effective for Multimodal Representation and Fusion: A Multimodal Spectrum Rumor Detector

An Lao, Qi Zhang, Chongyang Shi, Longbing Cao, Kun Yi, Liang Hu, Duoqian Miao

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM

Comments 12 pages, AAAI-2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10313 2023-12-06 cs.CL cs.AI cs.LG 84%

Investigating the Catastrophic Forgetting in Multimodal Large Language Models

Yuexiang Zhai, Shengbang Tong, Xiao Li, Mu Cai, Qing Qu, Yong Jae Lee, Yi Ma

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.03897 2023-11-08 cs.CV cs.CL cs.IR cs.LG 84%

Geodesic Multi-Modal Mixup for Robust Fine-Tuning

Changdae Oh, Junhyuk So, Hoyoon Byun, YongTaek Lim, Minchul Shin, Jong-June Jeon, Kyungwoo Song

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments To appear at NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09958 2023-09-19 cs.CV cs.CL 84%

An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Yadong Lu, Chunyuan Li, Haotian Liu, Jianwei Yang, Jianfeng Gao, Yelong Shen

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL

Comments Released at LLaVA Model Zoo: https://github.com/haotian-liu/LLaVA/blob/main/docs/MODEL_ZOO.md

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06573 2023-08-15 cs.CV cs.AI 84%

4DRVO-Net: Deep 4D Radar-Visual Odometry Using Multi-Modal and Multi-Scale Adaptive Fusion

Guirong Zhuo, Shouyi Lu, Huanyu Zhou, Lianqing Zheng, Lu Xiong

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 14 pages,12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09721 2023-07-20 cs.AI cs.CV 84%

Multi-Grained Multimodal Interaction Network for Entity Linking

Pengfei Luo, Tong Xu, Shiwei Wu, Chen Zhu, Linli Xu, Enhong Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by KDD 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06476 2023-06-13 cs.CL cs.AI 84%

Modality Influence in Multimodal Machine Learning

Abdelhamid Haouhat, Slimane Bellaouar, Attia Nehar, Hadda Cherroun

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL、cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.00320 2023-05-02 cs.CV cs.AI cs.LG 84%

Fusion for Visual-Infrared Person ReID in Real-World Surveillance Using Corrupted Multimodal Data

Arthur Josi, Mahdi Alehdaghi, Rafael M. O. Cruz, Eric Granger

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 31 pages, 11 figures, First version submitted to IJCV journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14204 2023-04-28 cs.AI cs.CV 84%

Towards Medical Artificial General Intelligence via Knowledge-Enhanced Multimodal Pretraining

Bingqian Lin, Zicong Chen, Mingjie Li, Haokun Lin, Hang Xu, Yi Zhu, Jianzhuang Liu, Wenjia Cai, Lei Yang, Shen Zhao, Chenfei Wu, Ling Chen, Xiaojun Chang, Yi Yang, Lei Xing, Xiaodan Liang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Project page: https://github.com/chenzcv7/MOTOR

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05554 2023-04-13 cs.CV cs.AI 84%

Learning Transferable Pedestrian Representation from Multimodal Information Supervision

Liping Bao, Longhui Wei, Xiaoyu Qiu, Wengang Zhou, Houqiang Li, Qi Tian

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02407 2023-04-06 cs.CV cs.AI cs.LG 84%

Explaining Multimodal Data Fusion: Occlusion Analysis for Wilderness Mapping

Burak Ekim, Michael Schmitt

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.02131 2023-03-16 cs.CV cs.CL cs.LG 84%

Masked Vision and Language Modeling for Multi-modal Representation Learning

Gukyeong Kwon, Zhaowei Cai, Avinash Ravichandran, Erhan Bas, Rahul Bhotika, Stefano Soatto

专题命中 多模态训练与对齐 :multi-modal(title);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments International Conference on Learning Representations (ICLR) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.05952 2023-03-13 cs.LG cs.AI cs.CV 84%

Understanding and Constructing Latent Modality Structures in Multi-modal Representation Learning

Qian Jiang, Changyou Chen, Han Zhao, Liqun Chen, Qing Ping, Son Dinh Tran, Yi Xu, Belinda Zeng, Trishul Chilimbi

专题命中 多模态训练与对齐 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 14 pages, 8 figure, CVPR 2023 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.05543 2023-02-17 cs.CV cs.MM 84%

Towards Reliable Image Outpainting: Learning Structure-Aware Multimodal Fusion with Depth Guidance

Lei Zhang, Kang Liao, Chunyu Lin, Yao Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.15824 2023-01-31 cs.MM cs.AI 84%

Improving the Modality Representation with Multi-View Contrastive Learning for Multimodal Sentiment Analysis

Peipei Liu, Xin Zheng, Hong Li, Jie Liu, Yimo Ren, Hongsong Zhu, Limin Sun

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏