arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2505.02486 2025-05-06 cs.LG cs.AI 79%

SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning

Jinpeng Chen, Runmin Cong, Yuzhi Zhao, Hongzheng Yang, Guangneng Hu, Horace Ho Shing Ip, Sam Kwong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09256 2025-05-06 cs.SD eess.AS 79%

Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation

Yifei Xin, Zhihong Zhu, Xuxin Cheng, Xusheng Yang, Yuexian Zou

机构 * YifeiXin(西菲·欣) ZhihongZhu(朱之宏) XuxinCheng(程旭鑫) XushengYang(杨旭生) YuexianZou(邹岳先)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 eess.AS

Comments Accepted by Interspeech2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01766 2025-05-06 cs.CV cs.RO 79%

Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement

Long Bai, Boyi Ma, Ruohan Wang, Guankun Wang, Beilei Cui, Zhongliang Jiang, Mobarakol Islam, Zhe Min, Jiewen Lai, Nassir Navab, Hongliang Ren

机构 * Department of Electronic Engineering, The Chinese University of Hong Kong(香港中文大学电子工程系) Chair for Computer Aided Medical Procedures, Technical University of Munich(慕尼黑技术大学计算机辅助医疗程序主席职位) Department of Biomedical Engineering, University of Toronto(多伦多大学生物医学工程系) Center for Computational and Molecular Biology, Brown University(布朗大学计算与分子生物学中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21281 2025-05-01 cs.CV 79%

Mamba Based Feature Extraction And Adaptive Multilevel Feature Fusion For 3D Tumor Segmentation From Multi-modal Medical Image

Zexin Ji, Beiji Zou, Xiaoyan Kui, Hua Li, Pierre Vera, Su Ruan

机构 * School of Computer Science and Engineering, Central South University, Changsha, 410083, China(计算机科学与工程学院,中南大学,长沙,410083,中国) Hunan Engineering Research Center of Machine Vision and Intelligent Medicine, Central South University, Changsha, 410083, China(机器视觉与智能医学工程研究中心,中南大学,长沙,410083,中国) Department of Radiation Oncology, Washington University in St. Louis, USA(放射肿瘤科,华盛顿大学圣路易斯分校,美国) Department of Nuclear Medicine, Henri Becquerel Cancer Center, Rouen, France(核医学科,亨利·贝克勒尔癌症中心,鲁昂,法国) University of Rouen-Normandy, AMIS - QuantIF UR 4108, F-76000, Rouen, France(鲁昂-诺曼底大学,AMIS - QuantIF UR 4108,法国,F-76000,鲁昂,法国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20362 2025-04-30 cs.CV 79%

TTTFusion: A Test-Time Training-Based Strategy for Multimodal Medical Image Fusion in Surgical Robots

Qinhua Xie, Hao Tang

机构 * School of Data Science and Engineering, East China Normal University(东华大学数据科学与工程学院) School of Computer Science, Peking University(北京大学计算机学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20178 2025-04-30 cs.CV cs.LG 79%

A Transformer-based Multimodal Fusion Model for Efficient Crowd Counting Using Visual and Wireless Signals

Zhe Cui, Yuli Li, Le-Nam Tran

机构 * School of Electrical and Electronic Engineering, University College Dublin(都柏林大学电子与电气工程学院) College of Electrical Engineering and Automation, Shandong University of Science and Technology(山东科技大学电气工程与自动化学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments This paper was accepted at IEEE WCNC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19839 2025-04-29 cs.CV 79%

SRMF: A Data Augmentation and Multimodal Fusion Approach for Long-Tail UHR Satellite Image Segmentation

Yulong Guo, Zilun Zhang, Yongheng Shang, Tiancheng Zhao, Shuiguang Deng, Yingchun Yang, Jianwei Yin

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Haina Institute of Zhejiang University, Advanced Technology Institute, Zhejiang University(浙江大学海纳研究院、先进技术研究院) Binjiang Research Institute of Zhejiang University(浙江大学滨江研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments None

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18961 2025-04-29 cs.IR cs.AI 79%

Feature Fusion Revisited: Multimodal CTR Prediction for MMCTR Challenge

Junjie Zhou

机构 * National Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家实验室,南京大学,中国) School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments A technical report for the MMCTR Challenge held by EReL@MIR Workshop at WWW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17223 2025-04-25 cs.CV 79%

Towards Generalizable Deepfake Detection with Spatial-Frequency Collaborative Learning and Hierarchical Cross-Modal Fusion

Mengyu Qiao, Runze Tian, Yang Wang

机构 * North China University of Technology(华北理工大学) Ultramain Systems, Inc.(Ultramain系统公司)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15848 2025-04-23 cs.CL 79%

Exploring Cognitive and Aesthetic Causality for Multimodal Aspect-Based Sentiment Analysis

Luwei Xiao, Rui Mao, Shuai Zhao, Qika Lin, Yanhao Jia, Liang He, Erik Cambria

机构 * School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) Saw Swee Hock School of Public Health, National University of Singapore(新加坡国立大学 Saw Swee Hock 公共卫生学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by TAFFC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15619 2025-04-23 cs.CV 79%

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization

Jinda Lu, Jinghan Li, Yuan Gao, Junkang Wu, Jiancan Wu, Xiang Wang, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10150 2025-04-23 cs.IR cs.MM 79%

HistLLM: A Unified Framework for LLM-Based Multimodal Recommendation with User History Encoding and Compression

Chen Zhang, Bo Hu, Weidong Chen, Zhendong Mao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments We want to withdraw this paper and revise its experimental details. The revised version will be uploaded after further verification

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07677 2025-04-23 cs.RO cs.CV 79%

Localization Meets Uncertainty: Uncertainty-Aware Multi-Modal Localization

Hye-Min Won, Jieun Lee, Jiyong Oh

机构 * Polaris3D

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14847 2025-04-22 cs.CV 79%

Reliable Multi-Modal Object Re-Identification via Modality-Aware Graph Reasoning

Xixi Wan, Aihua Zheng, Zi Wang, Bo Jiang, Jin Tang, Jixin Ma

机构 * Information Materials and Intelligent Sensing Laboratory of Anhui Province(安徽省信息材料与智能传感实验室) Anhui Provincial Key Laboratory of Security Artificial Intelligence(安徽省安全人工智能省级重点实验室) School(学校)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13319 2025-04-22 cs.CV cs.LG eess.IV 79%

HyperFusion: A Hypernetwork Approach to Multimodal Integration of Tabular and Medical Imaging Data for Predictive Modeling

Daniel Duenias, Brennan Nichyporuk, Tal Arbel, Tammy Riklin Raviv

机构 * Ben Gurion University of the Negev(本· Gurion 大学) Centre for Intelligent Machines, McGill University(智能机器中心,麦吉尔大学) Mila - Quebec AI Institute(魁北克人工智能研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 20 pages, 11 figures

Journal ref Medical Image Analysis, Volume 102, May 2025, 103503

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15379 2025-04-22 cs.CL 79%

DualKanbaFormer: An Efficient Selective Sparse Framework for Multimodal Aspect-based Sentiment Analysis

Adamu Lawan, Juhua Pu, Haruna Yunusa, Muhammad Lawan, Aliyu Umar, Adamu Sani Yahya, Mahmoud Basi

机构 * School of Computer Science and Technology, Beihang University, Beijing, China(北京航空航天大学计算机科学与技术学院) School of Automation Science and Electrical Engineering, Beihang University, Beijing, China(北京航空航天大学自动化科学与电气工程学院) School of Computing, University of Portsmouth, Portsmouth, UK(普利茅斯大学计算机学院) Department of Information and Communication Technology, Federal University, Gusau, Nigeria(加兹阿大学信息与通信技术系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments 12 pages, 2 figures, and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19950 2025-04-17 cs.LG cs.AI 79%

Multimodal Lego: Model Merging and Fine-Tuning Across Topologies and Modalities in Biomedicine

Konstantin Hemker, Nikola Simidjievski, Mateja Jamnik

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11262 2025-04-16 cs.CV 79%

Enhanced Small Target Detection via Multi-Modal Fusion and Attention Mechanisms: A YOLOv5 Approach

Xiaoxiao Ma, Junxiong Tong

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ATC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10985 2025-04-16 cs.CV 79%

DMPT: Decoupled Modality-aware Prompt Tuning for Multi-modal Object Re-identification

Minghui Lin, Shu Wang, Xiang Wang, Jianhua Tang, Longbin Fu, Zhengrong Zuo, Nong Sang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09197 2025-04-15 cs.AI 79%

Graph Learning-Driven Multi-Vessel Association: Fusing Multimodal Data for Maritime Intelligence

Yuxu Lu, Kaisen Yang, Dong Yang, Haifeng Ding, Jinxian Weng, Ryan Wen Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09106 2025-04-15 cs.CV 79%

Multi-modal and Multi-view Fundus Image Fusion for Retinopathy Diagnosis via Multi-scale Cross-attention and Shifted Window Self-attention

Yonghao Huang, Leiting Chen, Chuan Zhou

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14496 2025-04-15 cs.CV 79%

WikiStyle+: A Multimodal Approach to Content-Style Representation Disentanglement for Artistic Image Stylization

Ma Zhuoqi, Zhang Yixuan, You Zejun, Tian Long, Liu Xiyang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.07300 2025-04-09 cs.LG cs.CL 79%

CALF: Aligning LLMs for Time Series Forecasting via Cross-modal Fine-Tuning

Peiyuan Liu, Hang Guo, Tao Dai, Naiqi Li, Jigang Bao, Xudong Ren, Yong Jiang, Shu-Tao Xia

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13659 2025-04-08 cs.CV 79%

LMFNet: An Efficient Multimodal Fusion Approach for Semantic Segmentation in High-Resolution Remote Sensing

Tong Wang, Guanzhou Chen, Xiaodong Zhang, Chenxi Liu, Xiaoliang Tan, Jiaqi Wang, Chanjuan He, Wenlin Zhou

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11223 2025-04-07 cs.CV 79%

Multimodal Attention-Enhanced Feature Fusion-based Weekly Supervised Anomaly Violence Detection

Yuta Kaneko, Abu Saleh Musa Miah, Najmul Hassan, Hyoun-Sup Lee, Si-Woong Jang, Jungpil Shin

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Journal ref IEEE Open Journal of the Computer Society, vol. 6, pp. 129-140, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19776 2025-04-01 cs.CV 79%

Resilient Sensor Fusion under Adverse Sensor Failures via Multi-Modal Expert Fusion

Konyul Park, Yecheol Kim, Daehun Kim, Jun Won Choi

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17426 2025-04-01 cs.LG cs.AI cs.CR cs.ET 79%

Enhanced Smart Contract Reputability Analysis using Multimodal Data Fusion on Ethereum

Cyrus Malik, Josef Bajada, Joshua Ellul

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21367 2025-03-28 cs.CV 79%

Multimodal surface defect detection from wooden logs for sawing optimization

Bořek Reich, Matej Kunda, Fedor Zolotarev, Tuomas Eerola, Pavel Zemčík, Tomi Kauppi

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21124 2025-03-28 cs.CV 79%

AdaMHF: Adaptive Multimodal Hierarchical Fusion for Survival Prediction

Shuaiyu Zhang, Xun Lin, Rongxiang Zhang, Yu Bai, Yong Xu, Tao Tan, Xunbin Zheng, Zitong Yu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20190 2025-03-27 cs.CV 79%

Cross-Modal Prototype Allocation: Unsupervised Slide Representation Learning via Patch-Text Contrast in Computational Pathology

Yuxuan Chen, Jiawen Li, Jiali Hu, Xitong Ling, Tian Guan, Anjia Han, Yonghong He

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments 11pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏