arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2302.12636 2025-01-20 cs.LG cs.AI eess.SP 79%

Streamlining Multimodal Data Fusion in Wireless Communication and Sensor Networks

Mohammud J. Bocus, Xiaoyang Wang, Robert. J. Piechocki

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 10 pages, 12 figures, 3 tables, under review in IEEE Transactions on Cognitive Communications and Networking

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08490 2025-01-16 cs.CV cs.LG 79%

FLAVARS: A Multimodal Foundational Language and Vision Alignment Model for Remote Sensing

Isaac Corley, Simone Fobi Nsutezo, Anthony Ortiz, Caleb Robinson, Rahul Dodhia, Juan M. Lavista Ferres, Peyman Najafirad

专题命中 多模态训练与对齐 :multimodal(title);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07166 2025-01-14 cs.AI 79%

Natural Language-Assisted Multi-modal Medication Recommendation

Jie Tan, Yu Rong, Kangfei Zhao, Tian Bian, Tingyang Xu, Junzhou Huang, Hong Cheng, Helen Meng

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments 10 pages

Journal ref Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, Boise, ID, USA, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05993 2025-01-13 cs.CV 79%

Aria: An Open Multimodal Native Mixture-of-Experts Model

Dongxu Li, Yudong Liu, Haoning Wu, Yue Wang, Zhiqi Shen, Bowen Qu, Xinyao Niu, Fan Zhou, Chengen Huang, Yanpeng Li, Chongyan Zhu, Xiaoyi Ren, Chao Li, Yifan Ye, Peng Liu, Lihuan Zhang, Hanshu Yan, Guoyin Wang, Bei Chen, Junnan Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04373 2025-01-09 cs.CV 79%

FGU3R: Fine-Grained Fusion via Unified 3D Representation for Multimodal 3D Object Detection

Guoxin Zhang, Ziying Song, Lin Liu, Zhonghong Ou

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15209 2025-01-09 cs.CV 79%

MSCoTDet: Language-driven Multi-modal Fusion for Improved Multispectral Pedestrian Detection

Taeheon Kim, Sangyun Chung, Damin Yeom, Youngjoon Yu, Hak Gu Kim, Yong Man Ro

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03332 2025-01-08 cs.CV 79%

CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets

Tanay Agrawal, Mohammed Guermal, Michal Balazia, Francois Bremond

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Preprint. Final paper accepted at the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Tucson, February, 2025. 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05300 2025-01-08 cs.LG cs.AI 79%

Unity by Diversity: Improved Representation Learning in Multimodal VAEs

Thomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard, Norbert Fortin, Julia E. Vogt, Babak Shahbaba, Stephan Mandt

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at Neurips 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.20280 2025-01-03 cs.LG cs.AI 79%

Sparsely Multimodal Data Fusion

Josiah Bjorgaard

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20487 2024-12-31 cs.LG cs.CV cs.IT math.IT 79%

Multimodal Variational Autoencoder: a Barycentric View

Peijie Qiu, Wenhui Zhu, Sayantan Kumar, Xiwen Chen, Xiaotong Sun, Jin Yang, Abolfazl Razi, Yalin Wang, Aristeidis Sotiras

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18962 2024-12-30 cs.IR cs.MM 79%

Don't Lose Yourself: Boosting Multimodal Recommendation via Reducing Node-neighbor Discrepancy in Graph Convolutional Network

Zheyu Chen, Jinfeng Xu, Haibo Hu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02394 2024-12-30 eess.IV cs.CV 79%

Cohort-Individual Cooperative Learning for Multimodal Cancer Survival Analysis

Huajun Zhou, Fengtao Zhou, Hao Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17297 2024-12-24 cs.CV 79%

Revisiting Multimodal Fusion for 3D Anomaly Detection from an Architectural Perspective

Kaifang Long, Guoyang Xie, Lianbo Ma, Jiaqi Liu, Zhichao Lu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15652 2024-12-23 cs.CL 79%

Error-driven Data-efficient Large Multimodal Model Tuning

Barry Menglong Yao, Qifan Wang, Lifu Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14978 2024-12-20 cs.IR cs.MM 79%

Spectrum-based Modality Representation Fusion Graph Convolutional Network for Multimodal Recommendation

Rongqing Kenneth Ong, Andy W. H. Khong

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.MM

Comments Accepted to ACM Web Search and Data Mining (WSDM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01179 2024-12-20 cs.CV 79%

Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information

Yi Chen, Jian Xu, Xu-Yao Zhang, Wen-Zhuo Liu, Yang-Yang Liu, Cheng-Lin Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments AAAI2025 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12545 2024-12-19 cs.CL 79%

Enhancing Knowledge Distillation of Large Language Models through Efficient Multi-Modal Distribution Alignment

Tianyu Peng, Jiajun Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted by COLING 2025, 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13008 2024-12-18 cs.CL 79%

RCLMuFN: Relational Context Learning and Multiplex Fusion Network for Multimodal Sarcasm Detection

Tongguan Wang, Junkai Li, Guixin Su, Yongcheng Zhang, Dongyu Su, Yuxue Hu, Ying Sha

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.02703 2024-12-18 cs.CV 79%

Multi-modal Sensor Fusion for Auto Driving Perception: A Survey

Keli Huang, Botian Shi, Xiang Li, Xin Li, Siyuan Huang, Yikang Li

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 14 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11618 2024-12-17 cs.LG cs.AI 79%

EvoLlama: Enhancing LLMs' Understanding of Proteins via Multimodal Structure and Sequence Representations

Nuowei Liu, Changzhi Sun, Tao Ji, Junfeng Tian, Jianxin Tang, Yuanbin Wu, Man Lan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10650 2024-12-17 cs.CV 79%

DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-Identification

Yuhao Wang, Yang Liu, Aihua Zheng, Pingping Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments This work is accepted by AAAI2025. More motifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10033 2024-12-16 cs.CV 79%

Timealign: A multi-modal object detection method for time misalignment fusing in autonomous driving

Zhihang Song, Lihui Peng, Jianming Hu, Danya Yao, Yi Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05025 2024-12-16 cs.AI 79%

Debiased Multimodal Understanding for Human Language Sequences

Zhi Xu, Dingkang Yang, Mingcheng Li, Yuzheng Wang, Zhaoyu Chen, Jiawei Chen, Jinjie Wei, Lihua Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09511 2024-12-13 cs.CV 79%

GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency

Dongyue Lu, Lingdong Kong, Tianxin Huang, Gim Hee Lee

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments 22 pages, 8 figures, 12 tables; Project Page at https://dylanorange.github.io/projects/geal

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08979 2024-12-13 cs.LG cs.CV 79%

A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter

Zirun Guo, Xize Cheng, Yangyang Wu, Tao Jin

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03681 2024-12-13 cs.CL 79%

Acquired TASTE: Multimodal Stance Detection with Textual and Structural Embeddings

Guy Barel, Oren Tsur, Dan Vilenchik

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments COLING 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17418 2024-12-12 cs.CV 79%

Multimodal Outer Arithmetic Block Dual Fusion of Whole Slide Images and Omics Data for Precision Oncology

Omnia Alwazzan, Amaya Gallagher-Syed, Thomas O. Millner, Sebastian Brandner, Ioannis Patras, Silvia Marino, Gregory Slabaugh

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Revised to 10 pages, with corrected typos, updated references (some added, others removed), improved figure quality, modified text for better method validation, added one more co-author, and identified the IEEE member

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07475 2024-12-11 cs.CV 79%

Progressive Multi-Modal Fusion for Robust 3D Object Detection

Rohit Mohan, Daniele Cattaneo, Florian Drews, Abhinav Valada

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Journal ref 8th Annual Conference on Robot Learning, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06212 2024-12-10 cs.LG cs.AI 79%

A Self-guided Multimodal Approach to Enhancing Graph Representation Learning for Alzheimer's Diseases

Zhepeng Wang, Runxue Bao, Yawen Wu, Guodong Liu, Lei Yang, Liang Zhan, Feng Zheng, Weiwen Jiang, Yanfu Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15896 2024-12-09 cs.CV 79%

Multimodal Instruction Tuning with Conditional Mixture of LoRA

Ying Shen, Zhiyang Xu, Qifan Wang, Yu Cheng, Wenpeng Yin, Lifu Huang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 12 pages, 7 figures, ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏