arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2507.04638 2025-11-20 cs.CV 79%

UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-Identification

Xixi Wan, Aihua Zheng, Bo Jiang, Beibei Wang, Chenglong Li, Jin Tang

机构 * Information Materials and Intelligent Sensing Laboratory of Anhui Province, School of Artificial Intelligence, Anhui University(安徽省信息材料与智能感知实验室,人工智能学院,安徽大学) Anhui Provincial Key Laboratory of Multimodal Cognitive Computation, School of Computer Science and Technology, Anhui University(安徽省多模态认知计算重点实验室,计算机科学与技术学院,安徽大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14693 2025-11-19 cs.CL 79%

Talk, Snap, Complain: Validation-Aware Multimodal Expert Framework for Fine-Grained Customer Grievances

Rishu Kumar Singh, Navneet Shreya, Sarmistha Das, Apoorva Singh, Sriparna Saha

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments To be published in the Proceedings of the 40th Annual AAAI Conference on Artificial Intelligence (AAAI 2026 Special Track on AI for Social Impact )

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14157 2025-11-19 cs.CV 79%

Learning Representation and Synergy Invariances: A Povable Framework for Generalized Multimodal Face Anti-Spoofing

Xun Lin, Shuai Wang, Yi Yu, Zitong Yu, Jiale Zhou, Yizhong Liu, Xiaochun Cao, Alex Kot, Yefeng Zheng

机构 * Beihang University(北京航空航天大学) Nanyang Technological University(南洋理工大学) Westlake University(西湖大学) Great Bay University(大亚湾大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10218 2025-11-18 cs.AI 79%

MTP: Exploring Multimodal Urban Traffic Profiling with Modality Augmentation and Spectrum Fusion

Haolong Xiang, Peisi Wang, Xiaolong Xu, Kun Yi, Xuyun Zhang, Quanzheng Sheng, Amin Beheshti, Wei Fan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11452 2025-11-17 q-bio.QM cs.CV cs.LG eess.IV 79%

Synergy vs. Noise: Performance-Guided Multimodal Fusion For Biochemical Recurrence-Free Survival in Prostate Cancer

Seth Alain Chang, Muhammad Mueez Amjad, Noorul Wahab, Ethar Alzaid, Nasir Rajpoot, Adam Shephard

机构 * Tissue Image Analytics Centre, Department of Computer Science, University of Warwick, UK(沃里克大学计算机科学系组织图像分析中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 5 pages, 1 figure, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26289 2025-11-17 cs.MM 79%

Contribution-Guided Asymmetric Learning for Robust Multimodal Fusion under Imbalance and Noise

Zijing Xu, Yunfeng Kou, Kunming Wu, Hong Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10035 2025-11-14 cs.CV 79%

DGFusion: Dual-guided Fusion for Robust Multi-Modal 3D Object Detection

Feiyang Jia, Caiyan Jia, Ailin Liu, Shaoqing Xu, Qiming Xia, Lin Liu, Lei Yang, Yan Gong, Ziying Song

机构 * School of Computer Science and Technology, Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, Beijing Jiaotong University(计算机科学与技术学院、交通数据挖掘与具身智能北京市重点实验室、北京交通大学) State Key Laboratory of Internet of Things for Smart City and Department of Electrome chanical Engineering, University of Macau(智能城市物联网国家重点实验室、澳门大学机电工程系) Fujian Key Laboratory of Sensing and Computing for Smart Cities, Xiamen University(智能城市感知与计算福建省重点实验室、厦门大学) School of Mechanical and Aerospace Engineering, Nanyang Technological University(机械与航空航天工程学院、南洋理工大学) State Key Laboratory of Robotics and System, Harbin Institute of Technology(机器人系统国家重点实验室、哈尔滨工业大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12174 2025-11-14 cs.CV cs.RO 79%

UniGS: Unified Geometry-Aware Gaussian Splatting for Multimodal Rendering

Yusen Xie, Zhenmin Huang, Jianhao Jiao, Dimitrios Kanoulas, Jun Ma

机构 * HKUST (GZ)(香港科技大学(广州)) HKUST(香港科技大学) UCL(伦敦大学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09286 2025-11-13 cs.CV 79%

Enriching Knowledge Distillation with Cross-Modal Teacher Fusion

Amir M. Mansourian, Amir Mohammad Babaei, Shohreh Kasaei

机构 * Image Processing Lab, Sharif University of Technology(沙斐大学技术实验室)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments 11 pages, 5 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00452 2025-11-13 cs.IR cs.AI 79%

M^2VAE: Multi-Modal Multi-View Variational Autoencoder for Cold-start Item Recommendation

Chuan He, Yongchao Liu, Qiang Li, Wenliang Zhong, Chuntao Hong, Xinwei Yao

机构 * Ant Group(蚂蚁集团)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08152 2025-11-12 cs.CV cs.LG 79%

Boomda: Balanced Multi-objective Optimization for Multimodal Domain Adaptation

Jun Sun, Xinxin Zhang, Simin Hong, Jian Zhu, Xiang Gao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06805 2025-11-11 cs.AI cs.LG 79%

MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning

Jinhao Chen, Zhen Yang, Jianxin Shi, Tianyu Wo, Jie Tang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 19 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06593 2025-11-11 cs.CV 79%

Spatial-Frequency Enhanced Mamba for Multi-Modal Image Fusion

Hui Sun, Long Lv, Pingping Zhang, Tongdan Tang, Feng Tian, Weibing Sun, Huchuan Lu

机构 * School of Future Technology, School of Artificial Intelligence, Dalian University of Technology(大连理工大学未来技术学院、人工智能学院) Affiliated Zhongshan Hospital of Dalian University(大连大学附属中山医院) Central Hospital of Dalian University of Technology(大连理工大学中心医院) School of Information and Communication Engineering, Dalian University of Technology(大连理工大学信息与通信工程学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments This work is accepted by IEEE Transactions on Image Processing. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05474 2025-11-10 cs.CV 79%

Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection

Xian-Hong Huang, Hui-Kai Su, Chi-Chia Sun, Jun-Wei Hsieh

机构 * Department of Electrical Engineering, National Formosa University, Taiwan(台湾国立Formosa大学电子工程系) Department of Electrical Engineering, National Taipei University, Taiwan(台湾国立台北大学电子工程系) College of Artificial Intelligence and Green Energy, National Yang Ming Chiao Tung University, Taiwan(台湾国立阳明交通大学人工智能与再生能源学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04789 2025-11-06 cs.CV 79%

Object-X: Learning to Reconstruct Multi-Modal 3D Object Representations

Gaia Di Lorenzo, Federico Tombari, Marc Pollefeys, Daniel Barath

机构 * ETH Zurich(苏黎世联邦理工学院) Google(谷歌) Microsoft(微软)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18174 2025-11-05 eess.SP cs.AI cs.LG 79%

NMCSE: Noise-Robust Multi-Modal Coupling Signal Estimation Method via Optimal Transport for Cardiovascular Disease Detection

Peihong Zhang, Zhixin Li, Rui Sang, Yuxuan Liu, Yiqiang Cai, Yizhou Tan, Shengchen Li

机构 * School of Advanced Technology, Xi’an Jiaotong-Liverpool University(先进技术学院,西安交通大学利物浦大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15991 2025-11-05 cs.CV 79%

CrossRay3D: Geometry and Distribution Guidance for Efficient Multimodal 3D Detection

Huiming Yang, Wenzhuo Liu, Yicheng Qiao, Lei Yang, Xianzhu Zeng, Li Wang, Zhiwei Li, Zijian Zeng, Zhiying Jiang, Huaping Liu, Kunfeng Wang

机构 * Beijing University of Chemical Technology(北京化工大学) Division of Energy-Mobility Convergence, Beijing Institute of Technology(北京理工大学能源-交通融合学院) School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动学院) School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与航空航天工程学院) School of Mechanical Engineering, Beijing Institute of Technology(北京理工大学机械工程学院) Institute of Computer Science and Digital Innovation, UCSI University(UCSI大学计算机科学与数字创新学院) State Key Laboratory of Intelligent Technology and Systems and Department of Computer Science and Technology, Tsinghua University(清华大学智能技术与系统国家重点实验室与计算机科学与技术系)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01435 2025-11-04 cs.CV 79%

Contrast-Guided Cross-Modal Distillation for Thermal Object Detection

SiWoo Kim, JhongHyun An

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00859 2025-11-04 cs.CV 79%

Layer-Wise Modality Decomposition for Interpretable Multimodal Sensor Fusion

Jaehyun Park, Konyul Park, Daehun Kim, Junseo Park, Jun Won Choi

机构 * Seoul National University(首尔国立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19769 2025-11-04 cs.CV 79%

AIM: Adaptive Intra-Network Modulation for Balanced Multimodal Learning

Shu Shen, C. L. Philip Chen, Tong Zhang

机构 * Guangdong Provincial Key Laboratory of Computational AI Models and Cognitive Intelligence(广东省计算人工智能模型与认知智能重点实验室) School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) Pazhou Lab(琶洲实验室) Engineering Research Center of the Ministry of Education on Health Intelligent Perception and Paralleled Digital-Human(教育部健康智能感知与并行数字人工程研究中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 13pages,7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27166 2025-11-03 cs.CV 79%

M^3Detection: Multi-Frame Multi-Level Feature Fusion for Multi-Modal 3D Object Detection with Camera and 4D Imaging Radar

Xiaozhi Li, Huijun Di, Jian Li, Feng Liu, Wei Liang

机构 * Radar Technology Research Institute, School of Information and Electronics, Beijing Institute of Technology(雷达技术研究所,信息电子学院,北京理工大学) Key Laboratory of Electronic and Information Technology in Satellite Navigation, Ministry of Education(卫星导航电子信息技术重点实验室,教育部) School of Computer Science, Beijing Institute of Technology(计算机学院,北京理工大学) Beijing Racobit Electronic Information Technology Co., Ltd.(北京Racobit电子信息技术有限公司)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24919 2025-10-30 cs.CV cs.LG 79%

Modality-Aware SAM: Sharpness-Aware-Minimization Driven Gradient Modulation for Harmonized Multimodal Learning

Hossein R. Nowdeh, Jie Ji, Xiaolong Ma, Fatemeh Afghah

机构 * Holcombe Department of ECE(霍尔科姆电气与计算机工程系) Clemson University(克莱姆森大学) University of Arizona(亚利桑那大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03318 2025-10-30 cs.CV 79%

Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning

Yibin Wang, Zhimin Li, Yuhang Zang, Chunyu Wang, Qinglin Lu, Cheng Jin, Jiaqi Wang

机构 * College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) Shanghai Innovation Institute(上海创新研究院) Shanghai AI Lab(上海人工智能实验室) Hunyuan, Tencent(腾讯 Hunyuan)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments [NeurIPS2025] Project Page: https://codegoat24.github.io/UnifiedReward/think

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06456 2025-10-29 cs.CV 79%

DynCIM: Dynamic Curriculum for Imbalanced Multimodal Learning

Chengxuan Qian, Kai Han, Jiaxin Liu, Zhenlong Yuan, Zhengzhong Zhu, Jingchao Wang, Chongwen Lyu, Jun Chen, Zhe Liu

机构 * Jiangsu University(江苏大学) UIUC(伊利诺伊大学香槟分校) Alibaba(阿里巴巴) UCAS(中国科学院大学) Sichuan University(四川大学) Peking University(北京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22964 2025-10-28 cs.CV 79%

Survey of Multimodal Geospatial Foundation Models: Techniques, Applications, and Challenges

Liling Yang, Ning Chen, Jun Yue, Yidan Liu, Jiayi Ma, Pedram Ghamisi, Antonio Plaza, Leyuan Fang

机构 * School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) Institute of Remote Sensing and Geographic Information System, Peking University(北京大学遥感与地理信息系统研究所) School of Automation, Central South University(中南大学自动化学院) Electronic Information School, Wuhan University(武汉大学电子信息学院) Helmholtz-Zentrum Dresden-Rossendorf(德累斯顿-罗斯托克亥姆霍尔茨中心) Lancaster Environment Centre, Lancaster University(兰卡斯特大学环境研究中心) Hyperspectral Computing Laboratory, Department of Technology of Computers and Communications, Escuela Politécnica, University of Extremadura(埃斯特雷马杜拉大学技术计算机与通讯系超光谱计算实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11128 2025-10-27 cs.LG cs.CV 79%

Lightweight Facial Landmark Detection in Thermal Images via Multi-Level Cross-Modal Knowledge Transfer

Qiyi Tong, Olivia Nocentini, Marta Lagomarsino, Kuanqi Cai, Marta Lorenzini, Arash Ajoudani

机构 * Human-Robot Interfaces and Interaction Laboratory, Istituto Italiano di Tecnologia, Genoa, Italy(人机交互实验室,意大利技术研究院,热那亚,意大利) Ph.D. Program of National Interest in Robotics and Intelligent Machines (DRIM), Università di Genova(机器人与智能机器国家利益博士项目,热那亚大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19201 2025-10-24 cs.CL 79%

DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding

Yunhai Hu, Tianhua Xia, Zining Liu, Rahul Raman, Xingyu Liu, Bo Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang

机构 * Courant Institute of Mathematical Sciences, New York University(纽约大学数学科学学院) Tandon School of Engineering, New York University(纽约大学工程学院) Cerebras Systems Inc.(Cerebras Systems公司) University of Pennsylvania(宾夕法尼亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18321 2025-10-22 cs.CV 79%

Beyond Single Models: Mitigating Multimodal Hallucinations via Adaptive Token Ensemble Decoding

Jinlin Li, Yuran Wang, Yifei Yuan, Xiao Zhou, Yingying Zhang, Xixian Yong, Yefeng Zheng, Xian Wu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学 Gallup 学院) Department of Electrical and Computer Engineering, McGill University(麦吉尔大学电气与计算机工程系) School of Statistics, Renmin University of China(中国人民大学统计学院) Tencent Jarvis Lab(腾讯 Jarvis 实验室) Medical Artificial Intelligence Lab, Westlake University(西湖大学医学人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18244 2025-10-22 cs.CV 79%

BlendCLIP: Bridging Synthetic and Real Domains for Zero-Shot 3D Object Classification with Multimodal Pretraining

Ajinkya Khoche, Gergő László Nagy, Maciej Wozniak, Thomas Gustafsson, Patric Jensfelt

机构 * KTH Royal Institute of Technology(皇家理工学院) Scania CV AB(斯堪尼亚公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12445 2025-10-22 cs.MM 79%

M3ST-DTI: A multi-task learning model for drug-target interactions based on multi-modal features and multi-stage alignment

Xiangyu Li, Ran Su, Liangliang Liu

专题命中 多模态训练与对齐 :multi-modal(title);cross-modal(abstract);分类 cs.MM

Comments This paper accepted by IEEE BIBM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏