arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6878 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6878 篇

2603.21562 2026-03-24 cs.CV 79%

Exploring Multimodal Prompts For Unsupervised Continuous Anomaly Detection

探索多模态提示用于无监督连续异常检测

Mingle Zhou, Jiahui Liu, Jin Wan, Gang Li, Min Li

机构 * Key Laboratory of Computing Power Network(计算能力网络重点实验室) Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(信息安全,教育部,山东计算机科学中心(济南国家超算中心),齐鲁工业大学(山东科学院)) Shandong Provincial Key Laboratory of Computing Power Internet(山东省计算能力互联网重点实验室) Service Computing, Shandong Fundamental Research Center for Computer Science(服务计算,山东省计算机基础研究中心) Faculty of Data Science, City University of Macau(数据科学学院,澳门城市大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于多模态提示的无监督连续异常检测框架,通过持续多模态提示记忆库和缺陷语义引导自适应融合机制提升检测精度与鲁棒性,实验表明在MVTec AD和VisA数据集上取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21108 2026-03-24 cs.LG cs.AI 79%

DMMRL: Disentangled Multi-Modal Representation Learning via Variational Autoencoders for Molecular Property Prediction

DMMRL:通过变分自编码器进行解耦多模态表示学习以进行分子性质预测

Long Xu, Junping Guo, Jianbo Zhao, Jianbo Lu, Yuzhong Peng

机构 * Guangxi Key Lab of Human-machine Interaction and Intelligent Decision(广西人机交互与智能决策重点实验室) Nanning Normal University(南宁师范大学) College of Big Data and Software Engineering(大数据与软件工程学院) Zhejiang Wanli University(浙江万里大学)

专题命中 多模态训练与对齐 :multi-modal(title);cross-modal(abstract);分类 cs.AI

AI总结 本文提出DMMRL,通过变分自编码器解耦分子表示为共享和私有潜在空间,提升可解释性和预测性能,实验验证其在七个基准数据集上的优越表现。

Comments 9 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19718 2026-03-23 cs.CV 79%

BALM: A Model-Agnostic Framework for Balanced Multimodal Learning under Imbalanced Missing Rates

BALM:一种面向不平衡缺失率下平衡多模态学习的模型无关框架

Phuong-Anh Nguyen, Tien Anh Pham, Duc-Trong Le, Cam-Van Thi Nguyen

机构 * VNU University of Engineering and Technology(越南工程与技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出BALM框架,通过特征校准模块和梯度重平衡模块解决多模态学习中因不平衡缺失率导致的不平衡问题,提升模型鲁棒性和性能。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19681 2026-03-23 cs.CV 79%

Unbiased Dynamic Multimodal Fusion

无偏动态多模态融合

Shicai Wei, Kaijie Zhang, Luyi Chen, Tao He, Guiduo Duan

机构 * University of Electronic Science and Technology of China(电子科学与技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出UDML框架,通过引入噪声感知不确定性估计器和模态dropout量化模态依赖偏置,解决动态多模态融合中模态质量评估不准确和模态贡献分配不合理的问题。

Comments CVPR2026 Findings, 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15164 2026-03-23 cs.CV 79%

Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance

多模态持续指令微调与动态梯度指导

Songze Li, Mingyu Gao, Tonghua Su, Xu-Yao Zhang, Zhongjie Wang

机构 * Harbin Institute of Technology(哈尔滨工业大学) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济发展实验室) Chongqing Research Institute of HIT(重庆哈工大研究院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出通过几何参数空间方向向量近似缺失梯度,结合有限回放缓冲区和伯努利采样策略,有效缓解多模态持续指令微调中的灾难性遗忘问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18660 2026-03-20 cs.CV 79%

Multimodal Model for Computational Pathology:Representation Learning and Image Compression

多模态模型用于计算病理学:表示学习与图像压缩

Peihang Wu, Zehong Chen, Lijian Xu

机构 * Shenzhen University of Advanced Technology(深圳大学先进技术学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文探讨了多模态计算病理学中的表示学习和图像压缩技术,分析了自监督学习、多模态数据生成、参数高效适应及多代理协作推理等方向,旨在提升病理诊断的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13032 2026-03-20 cs.CV 79%

Multimodal OCR: Parse Anything from Documents

多模态OCR:从文档中解析一切

Handong Zheng, Yumeng Li, Kaile Zhang, Liang Xin, Guangwei Zhao, Hao Liu, Jiayu Chen, Jie Lou, Qi Fu, Rui Yang, Shuo Jiang, Weijian Luo, Weijie Su, Weijun Zhang, Xingyu Zhu, Yabin Li, Yiwei ma, Yu Chen, Yuqiu Ji, Zhaohui Yu, Guang Yang, Colin Zhang, Lei Zhang, Yuliang Liu, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) hi lab, Xiaohongshu Inc(小红书实验室,小红书公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出多模态OCR(MOCR),通过联合解析文本和图形,生成统一的文本表示。方法将图表、表格等视觉元素作为解析目标,提升文档重建的准确性,并通过端到端训练实现跨模态监督,最终在文档和结构化图形解析任务中取得优异表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17753 2026-03-19 cs.CV 79%

PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and Segmentation

PC-CrossDiff:点-簇双级跨模态微分注意力用于统一的3D指称与分割

Wenbin Tan, Jiawen Lin, Fangyong Wang, Yuan Xie, Yong Xie, Yachao Zhang, Yanyun Qu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出PC-CrossDiff框架,通过双级跨模态微分注意力解决复杂多物体场景中指称理解和分割的挑战,提升3D视觉定位的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17347 2026-03-19 cs.MM 79%

Beyond Forced Modality Balance: Intrinsic Information Budgets for Multimodal Learning

超越强制模态平衡:多模态学习中的内在信息预算

Zechang Xiong, Da Li, Kexin Tang, Pengyuan Li, Wenkang Kong, Yulan Hu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

AI总结 本文提出IIBalance框架,通过内在信息预算对齐模态贡献,解决多模态学习中的模态不平衡问题,实验表明其优于现有方法。

Comments 6 pages, 4 figures, paper accepted by ICME 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16259 2026-03-18 cs.MM 79%

Hyperbolic Multimodal Generative Representation Learning for Generalized Zero-Shot Multimodal Information Extraction

双曲多模态生成表示学习用于广义零样本多模态信息提取

Baohang Zhou, Kehui Song, Rize Jin, Yu Zhao, Xuhui Sui, Xinying Qian, Xingyue Guo, Ying Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

AI总结 本文提出双曲多模态生成表示学习框架HMGRL,解决零样本多模态信息提取中见与不见类别共存的问题,通过双曲空间建模多级语义关联,提升模型泛化能力。

Comments Accepted by WWW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16143 2026-03-18 eess.SP cs.AI 79%

Structure-Aware Multimodal LLM Framework for Trustworthy Near-Field Beam Prediction

面向可信近场波束预测的结构感知多模态大语言模型框架

Mengyuan Li, Qianfan Lu, Jiachen Tian, Hongjun Hu, Yu Han, Xiao Li, Chao-kai Wen, Shi Jin

机构 * School of Information Science and Engineering, Southeast University(信息科学与工程学院,东南大学) Institute of Communications Engineering, National Sun Yat-sen University(通讯工程学院,国立中山大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出基于大语言模型的多模态框架,融合历史GPS数据、RGB图像、LiDAR数据及任务特定文本提示,以提升复杂低空环境中的近场波束对齐能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15100 2026-03-17 cs.CV 79%

Learning from Limited and Incomplete Data: A Multimodal Framework for Predicting Pathological Response in NSCLC

从有限和不完整数据中学习:一种多模态框架用于预测非小细胞肺癌的病理反应

Alice Natalina Caragliano, Giulia Farina, Fatih Aksu, Camillo Maria Caruso, Claudia Tacconi, Carlo Greco, Lorenzo Nibid, Edy Ippolito, Michele Fiore, Giuseppe Perrone, Sara Ramella, Paolo Soda, Valerio Guarrasi

机构 * Department of Medicine and Surgery(医学与外科系) Department of Diagnostics and Intervention, Radiation Physics, Biomedical Engineering(诊断与介入系,放射物理,生物医学工程) Umeå University, Umeå, Sweden(乌梅大学,瑞典)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出一种多模态深度学习框架,整合基础模型的CT特征提取与缺失感知架构,以在有限数据和不完整临床资料下准确预测NSCLC的病理反应。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13298 2026-03-17 cs.LG cs.AI 79%

FusionCast: Enhancing Precipitation Nowcasting with Asymmetric Cross-Modal Fusion and Future Radar Priors

FusionCast: 通过不对称跨模态融合和未来雷达先验增强降水现在预报

Henan Wang, Shengwu Xiong, Yifang Zhang, Wenjie Yin, Chen Zhou, Yuqiang Zhang, Pengfei Duan

机构 * School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院) School of Earth and Space Science and Technology, Wuhan University(武汉大学地球和空间科学与技术学院)

专题命中 多模态训练与对齐 :cross-modal(title);multimodal(abstract);分类 cs.AI

AI总结 本文提出FusionCast框架,结合历史雷达QPE、PWV数据和预测雷达QPE,通过不对称融合和未来先验提升降水现在预报精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13291 2026-03-17 cs.LG cs.AI 79%

FedUAF: Uncertainty-Aware Fusion with Reliability-Guided Aggregation for Multimodal Federated Sentiment Analysis

FedUAF: 基于不确定性融合与可靠性引导聚合的多模态联邦情感分析

Xianxun Zhu, Zezhong Sun, Imad Rida, Erik Cambria, Junqi Su, Rui Wang, Hui Chen

机构 * Shanghai University(上海大学) North China Electric Power University(华北电力大学) Université de Technologie de Compiègne(法国图卢兹国立理工学院) Nanyang Technological University(南洋理工大学) City University of Hong Kong(香港城市大学) Macquarie University(麦考瑞大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出FedUAF框架,通过不确定性融合和可靠性引导聚合解决联邦学习中多模态数据缺失、分布异质和客户端更新不可靠的问题,实验证明其在CMU-MOSI和CMU-MOSEI数据集上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12762 2026-03-16 cs.CV cs.LG 79%

TerraFlow: Multimodal, Multitemporal Representation Learning for Earth Observation

TerraFlow:用于地球观测的多模态、多时间序列表示学习

Nazar Puriy, Johannes Jakubik, Benedikt Blumenstiel, Konrad Schindler

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 TerraFlow提出了一种新的多模态、多时间序列学习方法,适用于地球观测,能够处理变长输入,优于现有模型,在GEO-Bench-2基准测试中表现优异,并在自然灾害风险地图预测中取得进展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12166 2026-03-13 cs.CV 79%

LatentGeo: Learnable Auxiliary Constructions in Latent Space for Multimodal Geometric Reasoning

LatentGeo: 在潜在空间中学习可学习的辅助构造以进行多模态几何推理

Haiying Xu, Zihan Wang, Song Dai, Zhengxuan Zhang, Kairan Dou, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Nankai University(南开大学) Communication University of China(中国传媒大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 LatentGeo通过学习潜在空间中的连续视觉表示,解决多模态几何推理中辅助构造的表示问题,采用三阶段课程和强化学习方法提升几何推理任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02686 2026-03-13 cs.LG cs.AI 79%

A Systematic Review of Intermediate Fusion in Multimodal Deep Learning for Biomedical Applications

多模态深度学习在生物医学应用中的中间融合系统综述

Valerio Guarrasi, Fatih Aksu, Camillo Maria Caruso, Francesco Di Feola, Aurora Rofena, Filippo Ruffini, Paolo Soda

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文系统综述了多模态深度学习在生物医学应用中的中间融合方法,分析了现有技术、挑战及未来方向,并提出结构化符号以促进方法的广泛应用。

Journal ref Image and Vision Computing 158 (2025) 105509

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09706 2026-03-11 cs.AI 79%

OOD-MMSafe: Advancing MLLM Safety from Harmful Intent to Hidden Consequences

OOD-MMSafe: 推进多模态大语言模型安全性从有害意图到隐藏后果

Ming Wen, Kun Yang, Jingyu Zhang, Yuxuan Liu, shiwen cui, Shouling Ji, Xingjun Ma

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.AI

AI总结 OOD-MMSafe通过CASPO框架提升多模态大语言模型对隐藏后果的识别能力,显著降低风险识别失败率。

Comments 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09258 2026-03-11 cs.CV 79%

Multimodal Graph Representation Learning with Dynamic Information Pathways

多模态图表示学习中的动态信息路径

Xiaobin Hong, Mingkai Lin, Xiaoli Wang, Chaoqun Wang, Wenzhong Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出DiP框架,通过动态信息路径实现多模态图表示学习,提升跨模态消息传播的适应性和表达性。

Comments 12 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13015 2026-03-11 cs.CV 79%

Multimodal Classification via Total Correlation Maximization

通过总相关性最大化实现多模态分类

Feng Yu, Xiangyu Wu, Yang Yang, Jianfeng Lu

机构 * Nanjing University of Science and Technology(南京理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出TCMax方法,通过最大化多模态特征与标签间的总相关性,缓解模态竞争并提升多模态分类性能。

Comments Accepted for publication at ICLR 2026; 19 pages; 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21064 2026-03-11 cs.CV 79%

Multimodal Skeleton-Based Action Representation Learning via Decomposition and Composition

多模态骨骼基动作表示学习通过分解与组合

Hongsong Wang, Heng Fei, Bingxuan Dai, Jie Gui

机构 * School of Computer Science and Engineering, Southeast University, Nanjing 210096, China(东南大学计算机科学与工程学院,南京210096,中国) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用关键实验室(东南大学),中华人民共和国教育部,中国) School of Cyber Science and Engineering, Southeast University, Nanjing 210096, China(东南大学网络安全科学与工程学院,南京210096,中国) Engineering Research Center of Blockchain Application, Supervision And Management (Southeast University), Ministry of Education, China(区块链应用、监督与管理工程研究中心(东南大学),中华人民共和国教育部,中国) Purple Mountain Laboratories, Nanjing 210000, China(紫金山实验室,南京210000,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出一种自监督的多模态骨骼基动作表示学习框架,通过分解与组合策略平衡效率与效果,提升动作识别性能。

Comments Accepted by Machine Intelligence Research (Journal Impact Factor 8.7, 2024)

Journal ref Machine Intelligence Research, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07874 2026-03-10 cs.CV cs.LG 79%

Toward Unified Multimodal Representation Learning for Autonomous Driving

迈向自动驾驶的统一多模态表示学习

Ximeng Tao, Dimitar Filev, Gaurav Pandey

机构 * J. Mike Walker ’66 Department of Mechanical Engineering, Texas A&M University, College Station, TX 77843, USA(德克萨斯大学机械工程系,德克萨斯农工大学,学院站,德克萨斯,77843,美国) The Department of Engineering Technology and Industrial Distribution Texas A&M University, College Station, TX 77843, USA(工程技术与工业分布系,德克萨斯农工大学,学院站,德克萨斯,77843,美国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出CTP框架,通过统一多模态张量对齐提升自动驾驶性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10328 2026-03-09 cs.CV 79%

Fuse4Seg: Image Fusion for Multi-Modal Medical Segmentation via Bi-level Optimization

Fuse4Seg: 多模态医学分割的图像融合 via 两级优化

Yuchen Guo, Junli Gong, Hongmin Cai, Yiu-ming Cheung, Weifeng Su

机构 * Northwestern University(西北大学) Northeastern University(东北大学) South China University of Technology(华南理工大学) Hong Kong Baptist University(香港 Baptist大学) Beijing Normal - Hong Kong Baptist University(北京师范大学-香港 Baptist大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 Fuse4Seg通过两级优化实现多模态医学图像融合,解决视觉与语义间的差距问题,提升分割任务的准确性和临床可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04887 2026-03-06 cs.CV 79%

Federated Modality-specific Encoders and Partially Personalized Fusion Decoder for Multimodal Brain Tumor Segmentation

联邦模态特定编码器和部分个性化融合解码器用于多模态脑肿瘤分割

Hong Liu, Dong Wei, Qian Dai, Xian Wu, Yefeng Zheng, Liansheng Wang

机构 * National Institute for Data Science in Health and Medicine(国家医学数据科学研究院) Department of Computer Science at School of Informatics(信息学院计算机科学系) Jarvis Research Center(Jarvis研究中心) Medical Artificial Intelligence Lab(医学人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出FedMEPD框架,通过联邦模态特定编码器和部分个性化融合解码器,解决多模态医学图像分析中的模态间异质性和个性化需求问题。

Comments Medical Image Analysis 2025. arXiv admin note: substantial text overlap with arXiv:2403.11803

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04562 2026-03-06 cs.CV cs.LG 79%

Fusion and Grouping Strategies in Deep Learning for Local Climate Zone Classification of Multimodal Remote Sensing Data

深度学习中多模态遥感数据局部气候区分类的融合与分组策略

Ancymol Thomas, Jaya Sreevalsan-Nair

机构 * Graphics-Visualization-Computing Lab, International Institute of Information Technology Bangalore, Karnataka 560100, India(图形可视化计算实验室,国际信息学院班加罗尔,卡纳塔克邦560100,印度)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究提出了一种基于深度学习的多模态遥感数据局部气候区分类方法,通过融合与分组策略提升分类准确率,最终达到76.6%的整体准确率。

Comments 25 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02629 2026-03-04 cs.CV 79%

Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspective

迈向增量统一多模态异常检测:从信息瓶颈视角增强多模态去噪

Kaifang Long, Lianbo Ma, Jiaqi Liu, Liming Liu, Guoyang Xie

机构 * Software College, Northeastern University, China(东北大学软件学院) CATL, China(宁德时代)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出IB-IUMAD框架,通过Mamba解码器和信息瓶颈融合模块解决多模态异常检测中的灾难性遗忘问题,提升模型对新兴对象的适应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02162 2026-03-03 cs.CV 79%

Bridging the gap between Performance and Interpretability: An Explainable Disentangled Multimodal Framework for Cancer Survival Prediction

弥合性能与可解释性之间的鸿沟:一种可解释的解耦多模态框架用于癌症生存预测

Aniek Eijpe, Soufyan Lakbir, Melis Erdal Cesur, Sara P. Oliveira, Angelos Chatzimparmpas, Sanne Abeln, Wilson Silva

机构 * AI Technology for Life(人工智能技术与生命科学) Department of Information and Computing Sciences(信息与计算科学系) Department of Biology(生物学系) Utrecht University(乌得勒支大学) Department of Metabolic Diseases(代谢疾病部门) Wilhelmina Children’s Hospital(维廉明娜儿童医院) University Medical Center Utrecht(乌得勒支大学医学中心) Regenerative Medicine Center Utrecht(乌得勒支再生医学中心) Computational Pathology(计算病理学) Department of Pathology(病理学系) The Netherlands Cancer Institute(荷兰癌症研究所) Visualization and Graphics(可视化与图形学) The Netherlands Cancer Institute, Amsterdam(荷兰癌症研究所,阿姆斯特丹)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 DIMAFx通过解耦多模态表示提升癌症生存预测的性能与可解释性,揭示了多模态交互和生物学信息。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01758 2026-03-03 cs.CV 79%

Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining

通过语言枢轴预训练统一异构多模态遥感检测

Yuxuan Li, Yuming Chen, Yunheng Li, Ming-Ming Cheng, Xiang Li, Jian Yang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 BabelRS通过语言枢轴预训练框架统一异构多模态遥感检测,解耦模态对齐与任务学习,提升训练稳定性与检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09285 2026-03-03 cs.CV 79%

Spotlight on Token Perception for Multimodal Reinforcement Learning

多模态强化学习中的token感知聚焦

Siyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo, Zefeng He, Daizong Liu, Yu Cheng

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学) Nanjing University(南京大学) Wuhan University(武汉大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出VPPO算法,通过token感知优化提升多模态强化学习的视觉推理能力。

Comments Accepted by ICLR 2026, project page: https://github.com/huaixuheqing/VPPO-RL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03214 2026-03-03 cs.CV 79%

RTGMFF: Enhanced fMRI-based Brain Disorder Diagnosis via ROI-driven Text Generation and Multimodal Feature Fusion

RTGMFF:基于ROI驱动文本生成和多模态特征融合的增强型fMRI脑部疾病诊断

Junhao Jia, Yifei Sun, Yunyou Liu, Cheng Yang, Changmiao Wang, Feiwei Qin, Yong Peng, Wenwen Min

机构 * Hangzhou Dianzi University(杭州电子科技大学) Zhejiang University(浙江大学) Shenzhen Research Institute of Big Data(深圳大数据研究院) Yunnan University(云南大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 RTGMFF通过结合ROI驱动文本生成和多模态特征融合,提升fMRI在脑部疾病诊断中的准确性。

Comments The paper has been accepted by BIBM 2025

Journal ref 2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). IEEE, 2025, pp. 2301-2308

详情

展开后加载摘要…

URL PDF HTML 收藏