arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2508.15852 2025-08-25 cs.LG cs.CL 79%

PGF-Net: A Progressive Gated-Fusion Framework for Efficient Multimodal Sentiment Analysis

Bin Wen, Tien-Ping Tan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15505 2025-08-22 cs.CV 79%

Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion

Mengyu Wang, Zhenyu Liu, Kun Li, Yu Wang, Yuwei Wang, Yanyan Wei, Fei Wang

机构 * Key Laboratory of Opto-Electronic Information Science and Technology of Jiangxi Province, Nanchang Hangkong University(江西省光电信息科学与技术重点实验室,南昌航空大学) ReLER, CCAI, Zhejiang University(ReLER、CCAI、浙江大学) College of Engineering, Anhui Agricultural University(安徽农业大学工程学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13792 2025-08-22 cs.LG cs.AI cs.RO 79%

Continual Learning for Multimodal Data Fusion of a Soft Gripper

Nilay Kushawaha, Egidio Falotico

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted in Wiley Advanced Robotics Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12992 2025-08-20 cs.MM 79%

MAGNeT: Multimodal Adaptive Gaussian Networks for Intent Inference in Moving Target Selection across Complex Scenarios

Xiangxian Li, Yawen Zheng, Baiqiao Zhang, Yijia Ma, Xianhui Cao, Juan Liu, Yulong Bian, Jin Huang, Chenglei Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13196 2025-08-20 cs.LG cs.AI cs.IR 79%

Contextual Attention-Based Multimodal Fusion of LLM and CNN for Sentiment Analysis

Meriem Zerkouk, Miloud Mihoubi, Belkacem Chikhaoui

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments The 38th Canadian Conference on Artificial Intelligence ( 2025 )

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11350 2025-08-18 cs.CV 79%

HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model

Zhenhao Zhang, Hanqing Wang, Xiangyu Zeng, Ziyu Cheng, Jiaxin Liu, Haoyu Yan, Zhirui Liu, Kaiyang Ji, Tianxiang Gui, Ke Hu, Kangyi Chen, Yahao Fan, Mokai Pan

专题命中 多模态训练与对齐 :multimodal(title);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09717 2025-08-14 cs.CV cs.LG 79%

Multimodal Sheaf-based Network for Glioblastoma Molecular Subtype Prediction

Shekhnaz Idrissova, Islem Rekik

机构 * BASIRA Lab, Imperial-X(BASIRA实验室、Imperial-X) Department of Computing, Imperial College London, United Kingdom(计算系、帝国理工学院伦敦分校,英国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09182 2025-08-14 eess.IV cs.CV 79%

MedPatch: Confidence-Guided Multi-Stage Fusion for Multimodal Clinical Data

Baraa Al Jorf, Farah Shamout

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08925 2025-08-13 eess.AS cs.SD 79%

LPGNet: A Lightweight Network with Parallel Attention and Gated Fusion for Multimodal Emotion Recognition

Zhining He, Yang Xiao

机构 * Guangzhou University(广州大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 eess.AS

Comments Under peering review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01057 2025-08-13 cs.AI cs.RO 79%

Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance

Fengze Yang, Bo Yu, Yang Zhou, Xuewen Luo, Zhengzhong Tu, Chenxi Liu

机构 * Department of Civil & Environmental Engineering University of Utah(土木与环境工程系 犹他大学) Zachry Department of Civil and Environmental Engineering Texas A&M University(扎克里系 土木与环境工程系 德克萨斯农工大学) Department of Computer Science & Engineering Texas A&M University(计算机科学与工程系 德克萨斯农工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 24 pages, 6 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08279 2025-08-13 cs.LG cs.AI 79%

XFMNet: Decoding Cross-Site and Nonstationary Water Patterns via Stepwise Multimodal Fusion for Long-Term Water Quality Forecasting

Ziqi Wang, Hailiang Zhao, Cheng Bao, Wenzhuo Qian, Yuhao Yang, Xueqiang Sun, Shuiguang Deng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04770 2025-08-12 cs.LG cs.AI q-bio.MN 79%

Bidirectional Hierarchical Protein Multi-Modal Representation Learning

Xuefeng Liu, Songhao Jiang, Chih-chan Tien, Jinbo Xu, Rick Stevens

机构 * Argonne National Laboratory(阿贡国家实验室)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06496 2025-08-12 cs.CV cs.MA 79%

Med-GRIM: Enhanced Zero-Shot Medical VQA using prompt-embedded Multimodal Graph RAG

Rakesh Raj Madavan, Akshat Kaimal, Hashim Faisal, Chandrakala S

机构 * Shiv Nadar University Chennai(施瓦斯纳大学钦奈)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07951 2025-08-12 cs.CV 79%

Scaling Laws for Native Multimodal Models

Mustafa Shukor, Enrico Fini, Victor Guilherme Turrisi da Costa, Matthieu Cord, Joshua Susskind, Alaaeldin El-Nouby

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2025 (Oral). 28 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05934 2025-08-11 cs.HC cs.AI cs.LG 79%

ASLSL: Adaptive shared latent structure learning with incomplete multi-modal physiological data for multi-dimensional emotional feature selection

Xueyuan Xu, Tianze Yu, Wenjia Dong, Fulin Wei, Li Zhuo

机构 * School of Information Science and Technology, Beijing University of Technology, Beijing 100124, China(信息科学与技术学院,北京理工大学,北京) School of Artificial Intelligence, Anhui University, Beijing 100124, China(人工智能学院,安徽大学,北京)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03155 2025-08-11 cs.LG cs.AI 79%

Fusing Cross-Domain Knowledge from Multimodal Data to Solve Problems in the Physical World

Yu Zheng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06836 2025-08-11 cs.LG cond-mat.mtrl-sci cs.AI 79%

CAST: Cross Attention based multimodal fusion of Structure and Text for materials property prediction

Jaewan Lee, Changyoung Park, Hongjun Yang, Sungbin Lim, Woohyung Lim, Sehui Han

机构 * LG AI Research(LG人工智能研究)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04205 2025-08-07 cs.CV 79%

Small Lesions-aware Bidirectional Multimodal Multiscale Fusion Network for Lung Disease Classification

Jianxun Yu, Ruiquan Ge, Zhipeng Wang, Cheng Yang, Chenyu Lin, Xianjun Fu, Jikui Liu, Ahmed Elazab, Changmiao Wang

机构 * Xidian University(西安电子科技大学) Hangzhou Dianzi University(杭州电子科技大学) Zhejiang College of Security Technology, School of Artificial Intelligence(浙江安全技术学院人工智能学院) Shenzhen Polytechnic University(深圳职业技术学院) Shenzhen University(深圳大学) Shenzhen Research Institute of Big Data(深圳大数据研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08344 2025-08-06 cs.CV 79%

MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion

Jihao Gu, Fei Wang, Kun Li, Yanyan Wei, Zhiliang Wu, Dan Guo

机构 * University College London (UCL)(伦敦大学学院) School of Computer Science(计算机科学学院) Information Engineering, School of Artificial Intelligence, Hefei University of Technology (HFUT)(信息工程学院,人工智能学院,合肥工业大学) ReLER, CCAI, Zhejiang University, China(ReLER、CCAI、浙江大学,中国) Key Laboratory of Knowledge Engineering with Big Data (HFUT), Ministry of Education(大数据知识工程重点实验室(HFUT),教育部) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China(人工智能研究所,合肥综合性国家科学中心,中国) Xinsight Lab, Research Institute, Hefei Zhongjuyuan Intelligent Technology Co., Ltd., China(Xinsight实验室,研究院,合肥中睿智能科技有限公司,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 1st Place in Micro-gesture Classification sub-challenge in 3rd MiGA at IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01316 2025-08-05 cs.CV cs.HC 79%

Multimodal Attention-Aware Fusion for Diagnosing Distal Myopathy: Evaluating Model Interpretability and Clinician Trust

Mohsen Abbaspour Onari, Lucie Charlotte Magister, Yaoxin Wu, Amalia Lupi, Dario Creazzo, Mattia Tordin, Luigi Di Donatantonio, Emilio Quaia, Chao Zhang, Isel Grau, Marco S. Nobile, Yingqian Zhang, Pietro Liò

机构 * Information Systems Group, Eindhoven University of Technology, The Netherlands(埃因霍温技术大学信息系统组) Eindhoven Artificial Intelligence Systems Institute, The Netherlands(埃因霍温人工智能系统研究所) Department of Computer Science and Technology, University of Cambridge, United Kingdom(剑桥大学计算机科学与技术系) Department of Medicine - DIMED, Padua University Hospital, Italy(帕多瓦大学医院医学部-DIMED) Department of Environmental Sciences, Informatics, and Statistics, Ca’ Foscari University of Venice, Italy(威尼斯卡弗里大学环境科学、信息学与统计学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00963 2025-08-05 cs.LG cs.AI 79%

Rethinking Multimodality: Optimizing Multimodal Deep Learning for Biomedical Signal Classification

Timothy Oladunni, Alex Wong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00830 2025-08-05 cs.CV 79%

Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection

Yang Cao, Yihan Zeng, Hang Xu, Dan Xu

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科技大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Code Page: NeurIPS2023" target="_blank" rel="noopener">https://github.com/yangcaoai/CoDA_NeurIPS2023 This paper is accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.07348 2025-08-05 cs.CV 79%

Juggling With Representations: On the Information Transfer Between Imagery, Point Clouds, and Meshes for Multi-Modal Semantics

Dominik Laupheimer, Norbert Haala

机构 * Institute for Photogrammetry, University of Stuttgart, Germany(摄影测量研究所,斯图加特大学,德国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00447 2025-08-04 cs.CV cs.LG 79%

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text

Anju Rani, Daniel Ortiz-Arroyo, Petar Durdevic

机构 * Department of Energy Technology(能源技术系) Aalborg University(奥尔堡大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22268 2025-08-04 cs.IR cs.AI 79%

Multi-modal Relational Item Representation Learning for Inferring Substitutable and Complementary Items

Junting Wang, Chenghuan Guo, Jiao Yang, Yanhui Guo, Yan Gao, Hari Sundaram

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19651 2025-08-04 cs.LG cs.CL 79%

Unlocking Multi-Modal Potentials for Link Prediction on Dynamic Text-Attributed Graphs

Yuanyuan Xu, Wenjie Zhang, Ying Zhang, Xuemin Lin, Xiwei Xu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22037 2025-07-30 cs.CR cs.AI 79%

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security

Muzhi Dai, Shixuan Liu, Zhiyuan Zhao, Junyu Gao, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI),中国电信,中国) Northwestern Polytechnical University(西北工业大学) China Telecom, China(中国电信,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10774 2025-07-30 cs.LG cs.AI 79%

Context-Aware Probabilistic Modeling with LLM for Multimodal Time Series Forecasting

Yueyang Yao, Jiajun Li, Xingyuan Dai, MengMeng Zhang, Xiaoyan Gong, Fei-Yue Wang, Yisheng Lv

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20714 2025-07-29 cs.LG cs.AI q-bio.QM stat.AP 79%

Prostate Cancer Classification Using Multimodal Feature Fusion and Explainable AI

Asma Sadia Khan, Fariba Tasnia Khan, Tanjim Mahmud, Salman Karim Khan, Rishita Chakma, Nahed Sharmen, Mohammad Shahadat Hossain, Karl Andersson

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14683 2025-07-29 cs.CV 79%

Emerging Properties in Unified Multimodal Pretraining

Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, Guang Shi, Haoqi Fan

机构 * ByteDance Seed(字节跳动种子) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) Monash University(墨尔本大学) Hong Kong University of Science and Technology(香港科学与技术大学) UC Santa Cruz(加州大学圣克ruz分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 37 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏