A Multi-View Dynamic Fusion Framework: How to Improve the Multimodal Brain Tumor Segmentation from Multi-Views?
专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV
Comments 21 pages, 6 figures
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV
Comments 21 pages, 6 figures
专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV
Comments 2020 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments NeurIPS 2019 Workshops
专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV
Comments 10 pages, 3 figures, submitted to BIBM2020
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV
Comments 2018 IEEE International Conference on Acoustics, Speech and Signal Processing
专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV
Comments CVPR 2019, Jun 2019, Long Beach, United States http://cvpr2019.thecvf.com/
专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV
Comments Zhe Guo and Xiang Li contribute equally to this work
Journal ref 2018 IEEE 15th International Symposium on Biomedical Imaging (ISBI 2018), Washington, DC, 2018, pp. 903-907
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Some error corrections in Sect.2.2 and Table 5, Machine Translation, 2017
RMS@CC-MMD 2026:通过几何交互和多视图共识进行多模态厌女症检测
机构 * Chittagong University of Engineering & Technology (CUET)(吉大港工程技术大学)
专题命中 多模态训练与对齐 :multimodal(title,comments);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 研究针对多模态厌女症检测问题,提出GeoMVC方法,通过几何交互层建模跨模态对齐,用多视图共识策略减轻分布偏移,在相关挑战中取得一定排名成绩,凸显特定文化背景下建模的挑战。
Comments Accepted for the CC-MMD Grand Challenge at the 28th ACM International Conference on Multimodal Interaction (ICMI 2026)
专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract)
Comments An Empirical Study of Training ID-Agnostic Multi-modal Sequential Recommenders
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM;multi-modal(comments)
Comments Accepted by 6th Multi-Modal Learning and Applications Workshop (MULA), CVPR 2023. Code available at: https://github.com/zihuixue/DynMM
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)
Comments This paper is accepted by IEEE Transactions on Multimedia. This version addresses some mistakes and typos in the original paper. The appendix is available at https://github.com/TmacMai/Multimodal-Information-Bottleneck/blob/main/appendix.pdf
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM
Journal ref European Conference on Computer Vision Workshops: Multimodal Learning and Applications, Sep 2018, Munich, Germany. https://mula2018.github.io/
TerraMind:面向地球观测的大规模生成式多模态模型
机构 * IBM Research – Europe(IBM欧洲研究院) ; ETH Zurich(苏黎世联邦理工学院) ; Forschungszentrum Jülich(尤利希研究中心) ; European Space Agency(欧洲航天局) ; Φ \Phi -Lab(Φ实验室) ; NASA IMPACT ; University of Iceland(爱沙尼亚大学)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);any-to-any(abstract);multimodal foundation model(abstract)
AI总结 提出首个任意到任意生成式多模态基础模型TerraMind,通过双尺度表示(token级和像素级)预训练,实现零样本/少样本应用,并引入“模态思考”能力,在PANGAEA等基准上达到领先性能。
Comments Accepted at ICCV'25
专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL、cs.AI;multimodal(comments)
Comments Our original goal was to use Commander-GPT: Dividing and Routing for Multimodal Sarcasm Detection (arXiv:2506.19420) to replace Commander-GPT: Fully Unleashing the Sarcasm Detection Capability of Multi-Modal Large Language Models (arXiv:2503.18681). Due to various reasons, both versions were released, so we would like to withdraw the latter
面向基于大语言模型(LLM)的多模态情感分析的多粒度情感集成
机构 * Fuzhou University(福州大学) ; Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) ; Jiangxia University(江夏大学)
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)
AI总结 该研究提出MGSI多粒度情感集成框架,通过多尺度编码、文本引导对齐等优化,提升基于LLM的多模态情感分析性能,在四个公开基准上效果优于冻结LLM基线。
Comments Accepted to NLPCC 2026
双曲多模态持续学习
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)
AI总结 本研究针对双曲多模态持续学习的遗忘问题,建立理论基础推导了保留几何结构的持续学习框架,经实验验证其有效性。
Comments ICML 2026. 33 pages, 11 figures. Code: ICML" target="_blank" rel="noopener">https://github.com/HUBERILT/HMCL_ICML
多模态机器人演示中指令-轨迹不匹配的审计
机构 * AI Robot Association (AIRoA)(人工智能机器人协会(AIRoA)) ; The University of Tokyo(东京大学)
专题命中 多模态训练与对齐 :multimodal(title,abstract)
AI总结 针对多模态机器人演示中指令-轨迹不匹配问题,提出无需训练的MMPF审计框架,在LIBERO基准及真实机器人数据上实现最优ITM检测与标签修正,可提升下游策略学习性能并展示过滤演示的权衡。
Comments Accepted for publication in IEEE Robotics and Automation Letters (RA-L). 8 pages, 3 figures, 7 tables
LF²AR:考虑分层动态以改进语言模型的多模态适配
机构 * Université de Toulon, Aix-Marseille Université, CNRS, LIS, France(法国图卢兹大学、马赛大学、CNRS、LIS) ; Department of Engineering, University of Cambridge, UK(剑桥大学工程系) ; Mitsubishi Electric Research Laboratories (MERL), Cambridge, MA, USA(三菱电机研究实验室(MERL)) ; LIA, Avignon Université, France(法国阿维尼翁大学LIA) ; LIUM, Le Mans Université, France(法国勒芒大学LIUM) ; Zenidoc, Marseille, France(法国马赛Zenidoc)
专题命中 多模态训练与对齐 :multimodal(title);cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS
AI总结 本研究提出LF²AR架构,通过分层抽象-细化动态设计适配机制,在文本转图像、语音模态上提升语言模型性能,支持1.9倍生成加速。
Comments Published as a conference paper at COLM 2026
对比表示学习的几何力学:对齐势、熵分散和跨模态散度
机构 * University of Science and Technology of China(中国科学技术大学)
专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract)
AI总结 本文通过测度论框架,在大批量极限下证明InfoNCE目标与确定性能量景观的等价性,揭示单模态与对称多模态之间的几何分岔,并指出跨模态散度项导致模态间隙。
Comments 54 Pages, ICML 2026 (Refined document aesthetics for clearer reading)
Q-CueGraph:用于多模态推理的查询条件化视觉证据图
专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 Q-CueGraph是一种用于多模态推理的查询条件化视觉证据图,它为冻结读取器生成受预算约束的坐标级观测,在多个基准测试中显著提升了推理性能,尤其适用于证据可定位、问题能区分位置且分辨率受限的场景。
基于联合核熵 gromov-wasserstein 最优传输的多模态对齐
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)
AI总结 针对跨模态配对数据稀缺的场景,提出 JK-EGW 框架实现多模态对齐,理论样本复杂度匹配标准最优传输,实验中在数据稀缺的预训练编码器嵌入对齐任务上性能优于基线。
DAIF:一种基于数据的近似消息传递多模态监督学习中间融合框架
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)
AI总结 本研究提出DAIF数据自适应中间融合框架,结合随机矩阵理论与非参数依赖度量,通过近似消息传递生成去噪特征,在模拟及两个多模态数据集上的预测任务中表现优于或媲美现有方法。
标签偏移与分块模态缺失下的多模态域适应
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)
AI总结 针对标签偏移与分块模态缺失的多模态域适应问题,提出参考锚定方法,结合代理标签辅助策略,在模拟与RCC应用中实现了分布偏移下的稳定预测。
Comments 49 pages, 3 figures, 15 tables; includes supplementary material
GALA:淘宝商购推荐系统中用于自适应多模态表示的生成对齐学习
专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract)
AI总结 本文针对外卖推荐系统多模态融合难、语义与行为对齐不足的问题,提出三阶段流程GALA,通过生成式RL对齐阶段弥合预训练-微调差距,在淘宝商购部署后提升了订单量与AUC等指标。
Comments 13 pages, 12 figures, 5 tables. Accepted at the 2026 IEEE International Conference on Data Engineering (ICDE 2026), Industry and Applications Track
TIER-MoE:用于生物医学分类多模态融合的、基于条件模态风险的信任感知专家路由
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)
AI总结 该研究提出TIER-MoE风险引导子空间混合专家模型,用于生物医学分类多模态融合,可提升预测性能与概率校准,在多数据集上优于现有最优方法且具备强零样本泛化能力。
多模态持续学习中模态贡献漂移的正则化
机构 * School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机与人工智能学院)
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)
AI总结 针对多模态持续学习中的模态贡献漂移问题,提出含基于重放和无重放版本的CMCDR方法,经实验验证其通用性与有效性。
通过标签条件对比对齐实现成人到儿科心电图转换的知识引导跨模态融合
机构 * School of Instrument Science and Engineering, Southeast University(东南大学仪器科学与工程学院) ; Nanjing Medical University(南京医科大学) ; Zhengzhou University(郑州大学)
专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract)
AI总结 研究针对成人与儿科心电图转换问题,提出知识引导的跨模态融合框架PEACE,通过标签条件对比对齐等方法,在有限监督下实现更好的儿科心电图解释,消融实验证明标签条件知识对齐是关键驱动因素。
Comments This article was accidentally submitted as a new arXiv paper instead of a replacement of arXiv:2605.00647. Please refer to arXiv:2605.00647 for the correct and updated version