arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6872 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6872 篇

2507.16158 2025-07-23 cs.CV 83%

AMMNet: An Asymmetric Multi-Modal Network for Remote Sensing Semantic Segmentation

Hui Ye, Haodong Chen, Zeke Zexi Hu, Xiaoming Chen, Yuk Ying Chung

机构 * School of Computer Science, The University of Sydney(计算机科学学院,悉尼大学) School of Computer and Artificial Intelligence, Beijing Technology and Business University(计算机与人工智能学院,北京技术与商业大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15253 2025-07-22 cs.AI cs.LG cs.SI 83%

Disentangling Homophily and Heterophily in Multimodal Graph Clustering

Zhaochen Guo, Zhixiang Shen, Xuanting Xie, Liangjian Wen, Zhao Kang

机构 * University of Electronic Science and Technology of China(电子科技大学) Southwestern University of Finance and Economics(西南财经大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments Appear in ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14935 2025-07-22 cs.CV 83%

Open-set Cross Modal Generalization via Multimodal Unified Representation

Hai Huang, Yan Xia, Shulei Wang, Hanting Wang, Minghui Fang, Shengpeng Ji, Sashuai Zhou, Tao Jin, Zhou Zhao

机构 * Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07886 2025-07-22 cs.CV 83%

EgoM2P: Egocentric Multimodal Multitask Pretraining

Gen Li, Yutong Chen, Yiqian Wu, Kaifeng Zhao, Marc Pollefeys, Siyu Tang

机构 * ETH Zürich(苏黎世联邦理工学院) Zhejiang University(浙江大学) Microsoft(微软公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17066 2025-07-17 cs.CV cs.LG 83%

DUNIA: Pixel-Sized Embeddings via Cross-Modal Alignment for Earth Observation Applications

Ibrahim Fayad, Max Zimmer, Martin Schwartz, Fabian Gieseke, Philippe Ciais, Gabriel Belouze, Sarah Brood, Aurelien De Truchis, Alexandre d'Aspremont

机构 * Laboratoire des Sciences du Climat et de l’Environnement, LSCE/IPSL, France Department for AI in Society, Science Technology, Zuse Institute Berlin, Germany Department of Information Systems, University of Münster, Germany Department of Computer Science, CNRS, INRIA \& École Normale Supérieure, Paris 75230, France

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments 26 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10213 2025-07-15 cs.CV 83%

Boosting Multimodal Learning via Disentangled Gradient Learning

Shicai Wei, Chunbo Luo, Yang Luo

机构 * The Laboratory of Intelligent Collaborative Computing of UESTC(UESTC智能协同计算实验室) The School of Information and Communication Engineering of UESTC(UESTC信息与通信工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08855 2025-07-15 eess.IV cs.CV cs.LG 83%

Multi-omic Prognosis of Alzheimer's Disease with Asymmetric Cross-Modal Cross-Attention Network

Yang Ming, Jiang Shi Zhong, Zhou Su Juan

机构 * College of Medical Information Engineering, Guangdong Pharmaceutical University, Guangzhou, Guangdong 510006, China(医学信息工程学院,广东药科大学,广州,广东510006,中国)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07108 2025-07-11 cs.CV cs.AI cs.CL cs.LG cs.MM 83%

Multi-level Mixture of Experts for Multimodal Entity Linking

Zhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li, Jeff Z. Pan

机构 * School of Computer Information Technology Shanxi University Taiyuan China School of Computer Science Informatics Cardiff University Cardiff UK ILCC, School of Informatics University of Edinburgh Edinburgh UK Information Technology Shanxi University Informatics Cardiff University ILCC, School of Informatics University of Edinburgh

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at KDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12151 2025-07-08 cs.LG cs.AI 83%

Towards Explainable Fusion and Balanced Learning in Multimodal Sentiment Analysis

Miaosen Luo, Yuncheng Jiang, Sijie Mai

机构 * School of Computer Science, South China Normal University(华南师范大学计算机学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04664 2025-07-08 cs.CV 83%

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs

Tao Zhang, Shiqing Wei, Shihao Chen, Wenling Yu, Muying Luo, Shunping Ji

机构 * School of Remote Sensing and Information Engineering(遥感与信息工程学院) College of Oceanography and Space Informatics(海洋学与空间信息学院) School of Surveying and Geoinformation Engineering(测绘与地理信息工程学院)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04635 2025-07-08 cs.CV 83%

MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding

Zhicheng Zhang, Wuyou Xia, Chenxi Zhao, Zhou Yan, Xiaoqiang Liu, Yongjie Zhu, Wenyu Qin, Pengfei Wan, Di Zhang, Jufeng Yang

机构 * VCIP \& TMCC \& DISSec, College of Computer Science, Nankai University Pengcheng Laboratory

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments ICML 2025 (Spotlight, Top 2.6%)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03019 2025-07-08 cs.CV cs.LG 83%

Look-Back: Implicit Visual Re-focusing in MLLM Reasoning

Shuo Yang, Yuwei Niu, Yuyang Liu, Yang Ye, Bin Lin, Li Yuan

机构 * Peking University(北京大学) Shenzhen Graduate School(深圳研究生院) Peng Cheng Laboratory(鹏城实验室)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02859 2025-07-04 cs.CV 83%

Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation

Jiaer Xia, Bingkui Tong, Yuhang Zang, Rui Shao, Kaiyang Zhou

机构 * Hong Kong Baptist University(香港 Baptist 大学) Sichuan University(四川大学) Shanghai AI Lab(上海人工智能实验室) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16236 2025-07-04 cs.CV 83%

LLaVA-KD: A Framework of Distilling Multimodal Large Language Models

Yuxuan Cai, Jiangning Zhang, Haoyang He, Xinwei He, Ao Tong, Zhenye Gan, Chengjie Wang, Zhucun Xue, Yong Liu, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Zhejiang University(浙江大学) Youtu Lab, Tencent(腾讯优图实验室) Huazhong Agricultural University(华中农业大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments ICCV'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02080 2025-07-04 cs.MM cs.SD 83%

TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation

Yubeen Lee, Sangeun Lee, Chaewon Park, Junyeop Cha, Eunil Park

机构 * Sungkyunkwan University(全北大学) Electronics and Telecommunications Research Institute(电子电信研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

Comments 9 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21096 2025-07-02 cs.CL 83%

DALR: Dual-level Alignment Learning for Multimodal Sentence Representation Learning

Kang He, Yuzhe Ding, Haining Wang, Fei Li, Chong Teng, Donghong Ji

机构 * Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University(航天信息安全部门与可信计算重点实验室,教育部,网络安全科学与工程学院,武汉大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted by ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23462 2025-07-01 cs.LG cs.AI 83%

Can We Predict the Unpredictable? Leveraging DisasterNet-LLM for Multimodal Disaster Classification

Manaswi Kulahara, Gautam Siddharth Kashyap, Nipun Joshi, Arpita Soni

机构 * TERI School Of Advanced Studies(TERI高级研究学院) Macquarie University(麦考瑞大学) Cornell University(康奈尔大学) Eudoxia Research University(欧多西亚研究大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments Accepted in the 2025 IEEE International Geoscience and Remote Sensing Symposium (IGARSS 2025), scheduled for 3 - 8 August 2025 in Brisbane, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22736 2025-07-01 cs.CV 83%

UniFuse: A Unified All-in-One Framework for Multi-Modal Medical Image Fusion Under Diverse Degradations and Misalignments

Dayong Su, Yafei Zhang, Huafeng Li, Jinxing Li, Yu Liu

机构 * Kunming University of Science and Technology(昆明理工大学) Harbin Institute of Technology at Shenzhen(哈尔滨工业大学深圳研究院) Hefei University of Technology(合肥工业大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19326 2025-07-01 cs.CV 83%

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment

Ziang Yan, Zhilin Li, Yinan He, Chenting Wang, Kunchang Li, Xinhao Li, Xiangyu Zeng, Zilei Wang, Yali Wang, Yu Qiao, Limin Wang, Yi Wang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) Nanjing University(南京大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21536 2025-07-01 cs.CL 83%

Tracing Intricate Cues in Dialogue: Joint Graph Structure and Sentiment Dynamics for Multimodal Emotion Recognition

Jiang Li, Xiaoping Wang, Zhigang Zeng

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) Institute of Artificial Intelligence, Huazhong University of Science and Technology(华中科技大学人工智能研究院) Hubei Key Laboratory of Brain-Inspired Intelligent Systems, Huazhong University of Science and Technology(华中科技大学脑启发智能系统省重点实验室) Key Laboratory of Image Processing and Intelligent Control (Huazhong University of Science and Technology), Ministry of Education(图像处理与智能控制重点实验室(华中科技大学))

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22446 2025-07-01 cs.LG cs.AI 83%

EAGLE: Efficient Alignment of Generalized Latent Embeddings for Multimodal Survival Prediction with Interpretable Attribution Analysis

Aakash Tripathi, Asim Waqas, Matthew B. Schabath, Yasin Yilmaz, Ghulam Rasool

机构 * Dept. of Machine Learning Moffitt Cancer Center(机器学习系莫菲特癌症中心) Dept. of Cancer Epidemiology Moffitt Cancer Center(癌症流行病学系莫菲特癌症中心) Dept. of Electrical Engineering University of South Florida(电气工程系佛罗里达州立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13980 2025-06-25 cs.CV 83%

FusionSAM: Visual Multi-Modal Learning with Segment Anything

Daixun Li, Weiying Xie, Mingxiang Cao, Yunke Wang, Yusi Zhang, Leyuan Fang, Yunsong Li, Chang Xu

机构 * Xidian University(西安电子科技大学) University of Sydney(悉尼大学) Hunan University(湖南大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16804 2025-06-25 cs.LG cs.AI cs.CY cs.ET 83%

Multimodal Machine Learning in Mental Health: A Survey of Data, Algorithms, and Challenges

Zahraa Al Sahili, Ioannis Patras, Matthew Purver

机构 * Queen Mary University of London United Kingdom Queen Mary University of London \& Jo z ef Stefan Institute United Kingdom \& Slovenia Queen Mary University of London Queen Mary University of London \& Jo z ef Stefan Institute

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18042 2025-06-24 cs.CV 83%

CmFNet: Cross-modal Fusion Network for Weakly-supervised Segmentation of Medical Images

Dongdong Meng, Sheng Li, Hao Wu, Suqing Tian, Wenjun Ma, Guoping Wang, Xueqing Yan

机构 * School of Physics, Peking University(北京大学物理学院) School of Computer Science, Peking University(北京大学计算机学院) Department of Radiotherapy, Peking University Cancer Hospital(北京大学肿瘤医院放疗科) Department of Radiation Oncology, Peking University Third Hospital(北京大学第三医院放疗科)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11976 2025-06-24 cs.CV cs.LG 83%

How Visual Representations Map to Language Feature Space in Multimodal LLMs

Constantin Venhoff, Ashkan Khakzar, Sonia Joseph, Philip Torr, Neel Nanda

机构 * University of Oxford(牛津大学) McGill University(麦吉尔大学) Meta

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08454 2025-06-24 cs.CL 83%

Alignment Helps Make the Most of Multimodal Data

Christian Arnold, Andreas Küpfer

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Working Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16737 2025-06-23 cs.CV 83%

Cross-modal Offset-guided Dynamic Alignment and Fusion for Weakly Aligned UAV Object Detection

Liu Zongzhen, Luo Hui, Wang Zhixing, Wei Yuxing, Zuo Haorui, Zhang Jianlin

机构 * State Key Laboratory of Optical Field Manipulation Science and Technology, Chinese Academy of Sciences(光学场操控科学与技术国家重点实验室,中国科学院) Institute of Optics and Electronics, Chinese Academy of Sciences(中国科学院光电研究所)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07895 2025-06-23 cs.LG cs.AI 83%

Representation Learning with Mutual Influence of Modalities for Node Classification in Multi-Modal Heterogeneous Networks

Jiafan Li, Jiaqi Zhu, Liang Chang, Yilin Li, Miaomiao Li, Yang Wang, Hongan Wang

机构 * Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学) School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院) Binzhou Institute of Technology(滨州职业技术学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14532 2025-06-18 cs.CL 83%

M2BeamLLM: Multimodal Sensing-empowered mmWave Beam Prediction with Large Language Models

Can Zheng, Jiguang He, Chung G. Kang, Guofa Cai, Zitong Yu, Merouane Debbah

机构 * Department of Electrical and Computer Engineering, Korea University(韩国大学电子与计算机工程系) School of Computing and Information Technology, Great Bay University(大湾大学计算机与信息科技学院) Dongguan Key Laboratory for Intelligence and Information Technology(东莞智能与信息科技重点实验室) Great Bay Institute for Advanced Study (GBIAS)(大湾先进研究学院) School of Information Engineering, Guangdong University of Technology(广东工业大学信息工程学院) Center for 6G Technology, Khalifa University of Science and Technology(哈里发科技大学6G技术中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

Comments 13 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07950 2025-06-17 cs.CV 83%

Decoupled Cross-Modal Alignment Network for Text-RGBT Person Retrieval and A High-Quality Benchmark

Yifei Deng, Chenglong Li, Zhenyu Chen, Zihen Xu, Jin Tang

机构 * School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院) School of Artificial Intelligence, Anhui University(安徽大学人工智能学院) National Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology(光电信息采集与防护技术国家重点实验室) Anhui Provincial Key Laboratory of Multimodal Cognitive Computation(安徽省多模态认知计算重点实验室)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏