arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2507.12108 2026-02-16 cs.SI cs.AI cs.CY cs.HC cs.LG 79%

Multimodal Coordinated Online Behavior: Trade-offs and Strategies

多模态协调在线行为:权衡与策略

Lorenzo Mannocci, Stefano Cresci, Matteo Magnani, Anna Monreale, Maurizio Tesconi

机构 * University of Pisa(比萨大学) Institute for Informatics and Telematics, National Research Council (IIT-CNR)(信息学与电信学研究院,国家研究理事会(IIT-CNR)) InfoLab, Department of Information Technology, Uppsala University(信息实验室,信息科技系,乌普萨拉大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本研究探讨多模态协调在线行为的权衡与策略,通过比较不同方法揭示多模态分析在捕捉协调行为结构方面的优势。

Comments Postprint of the article published in the Information Sciences journal. Please, cite accordingly

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09934 2026-02-11 cs.CV 79%

VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization

VersaViT: 通过任务引导优化增强MLLM视觉骨干

Yikun Liu, Yuan Liu, Shangzhe Di, Haicheng Wang, Zhongyin Zhao, Le Tian, Xiao Zhou, Jie Zhou, Jiangchao Yao, Yanfeng Wang, Weidi Xie

机构 * School of Artificial Intelligence, Shanghai Jiao Tong University, China(上海交通大学人工智能学院) CMIC, Shanghai Jiao Tong University, China(上海交通大学计算机学院) WeChat AI, Tencent Inc., China(腾讯公司)

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.CV

AI总结 VersaViT通过任务引导优化增强MLLM视觉骨干,解决其在密集预测任务中的性能问题,提升视觉任务的适应性与表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09485 2026-02-11 cs.AI 79%

Bridging Efficiency and Transparency: Explainable CoT Compression in Multimodal Large Reasoning Models

连接效率与透明性:可解释的CoT压缩在多模态大推理模型中

Yizhi Wang, Linan Yue, Min-Ling Zhang

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of Computer Network and Information Integration (SEU), Ministry of Education, China(计算机网络与信息集成重点实验室(SEU))

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 XMCC通过强化学习优化的顺序决策过程,实现多模态大推理模型中可解释的CoT压缩,同时保持推理正确性和生成可解释的压缩解释。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04121 2026-02-11 cs.CV 79%

Graph-Based Multimodal and Multi-view Alignment for Keystep Recognition

基于图的多模态和多视角对齐用于按键识别

Julia Lee Romero, Kyle Min, Subarna Tripathi, Morteza Karimzadeh

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于图的多模态和多视角对齐框架,用于提高egocentric视频中按键识别的准确率,通过构建稀疏图结构并利用多模态特征提升性能。

Comments We expanded the paper and resubmitted as a separate submission to arXiv. This submission is outdated and readers can refer to arXiv:2506.01102

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08713 2026-02-10 cs.CV cs.LG 79%

Towards Understanding Multimodal Fine-Tuning: Spatial Features

迈向多模态微调的理解:空间特征

Lachin Naghashyar, Hunar Batra, Ashkan Khakzar, Philip Torr, Ronald Clark, Christian Schroeder de Witt, Constantin Venhoff

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文通过分阶段模型差分技术,揭示了多模态微调过程中视觉特征如何形成及空间关系编码,提升了多模态训练的可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08077 2026-02-10 cs.LG cs.AI 79%

Multimodal normative modeling in Alzheimers Disease with introspective variational autoencoders

在阿尔茨海默病中使用反思变分自编码器的多模态规范建模

Sayantan Kumar, Peijie Qiu, Aristeidis Sotiras

机构 * Washington University in St Louis(华盛顿大学圣路易斯分校) Washington University in St Louis School of Medicine(华盛顿大学圣路易斯医学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出mmSIVAE,通过结合MOPOE聚合提升多模态数据的规范建模效果,提高参考分布保真度和多模态整合能力,用于阿尔茨海默病的偏差分析。

Comments Conference on Health, Inference, and Learning (CHIL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07163 2026-02-06 cs.CV 79%

Test-time Adaptive Hierarchical Co-enhanced Denoising Network for Reliable Multimodal Classification

测试时自适应层次联合增强去噪网络用于可靠的多模态分类

Shu Shen, C. L. Philip Chen, Tong Zhang

机构 * The Guangdong Provincial Key Laboratory of Computational Intelligence and Cyberspace Information, the School of Computer Science and Engineering, South China University of Technology(广东省计算智能与网络信息重点实验室、计算机科学与工程学院、华南理工大学) The Pazhou Laboratory(琶洲实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出TAHCD网络,通过自适应稳定子空间对齐和样本自适应置信度对齐,有效去除多模态噪声,提升多模态分类的鲁棒性和泛化能力。

Comments 14 pages,9 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04512 2026-02-05 q-bio.NC cs.AI 79%

BrainVista: Modeling Naturalistic Brain Dynamics as Multimodal Next-Token Prediction

BrainVista: 以多模态下一项令牌预测建模自然主义脑动力学

Xuanhua Yin, Runkai Zhao, Lina Yao, Weidong Cai

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 BrainVista通过多模态自回归框架建模自然主义脑动力学,采用网络级令牌化器和空间混合头以解构系统特定动态并捕捉跨网络信息流,通过S2B掩码机制实现严格因果条件,提升fMRI编码性能。

Comments 17 pages, 7 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04670 2026-02-05 cs.AI 79%

Improving Multimodal Brain Encoding Model with Dynamic Subject-awareness Routing

改进多模态脑编码模型的动态主体感知路由

Xuanhua Yin, Runkai Zhao, Weidong Cai

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 AFIRE和MIND通过动态主体感知路由提升多模态脑编码模型的性能,增强跨受试者泛化能力并实现可解释的专家模式。

Comments 7 pages, 4 figures, accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20977 2026-02-05 cs.CL 79%

Evaluating and Steering Modality Preferences in Multimodal Large Language Model

评估和引导多模态大语言模型中的模态偏好

Yu Zhang, Jinlong Ma, Yongshuai Hou, Xuefeng Bai, Kehai Chen, Yang Xiang, Jun Yu, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CL

AI总结 本文提出MC²基准测试,通过受控证据冲突场景评估和引导多模态大语言模型的模态偏好,揭示了其可通过指令引导和潜在表示控制,并展示了通过表示工程方法提升多模态任务性能的潜力。

Comments Modality Preference

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03815 2026-02-04 cs.CV cs.LG 79%

Fast-Slow Efficient Training for Multimodal Large Language Models via Visual Token Pruning

通过视觉标记修剪实现多模态大语言模型的快慢高效训练

Dingkun Zhang, Shuhan Qi, Yulin Wu, Xinyu Xiao, Xuan Wang, Long Chen

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 DualSpeed通过快慢双模式实现多模态大语言模型的高效训练,提升训练速度并保持性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03665 2026-02-04 cs.CV cs.HC 79%

MM-SCALE: Grounded Multimodal Moral Reasoning via Scalar Judgment and Listwise Alignment

MM-SCALE: 基于标量判断和列表对齐的 grounded 多模态道德推理

Eunkyu Park, Wesley Hanwen Deng, Cheyon Jin, Matheus Kunzler Maldaner, Jordan Wheeler, Jason I. Hong, Hong Shen, Adam Perer, Ken Holstein, Motahhare Eslami, Gunhee Kim

机构 * Seoul National University(首尔国立大学) Carnegie Mellon University(卡内基梅隆大学) University of Florida(佛罗里达大学) Epic Games

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 MM-SCALE通过5点标量评分和显式模态接地,提升多模态模型在道德推理任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11442 2026-02-04 cs.CV 79%

MultiMAE for Brain MRIs: Robustness to Missing Inputs Using Multi-Modal Masked Autoencoder

MultiMAE用于脑部MRI:通过多模态掩码自编码器提高对缺失输入的鲁棒性

Ayhan Can Erdur, Christian Beischl, Daniel Scholz, Jiazhen Pan, Benedikt Wiestler, Daniel Rueckert, Jan C Peeken

机构 * Department of Radiation Oncology, TUM University Hospital, Munich, Germany(辐射肿瘤科,技术大学慕尼黑医院,慕尼黑,德国) Chair for AI in Healthcare and Medicine, Technical University of Munich (TUM)(医疗与医学人工智能教研室,技术大学慕尼黑(TUM)) TUM University Hospital, Munich, Germany(技术大学慕尼黑医院,慕尼黑,德国) Chair for AI for Image-Guided Diagnosis and Therapy, Technical University of Munich (TUM)(图像引导诊断与治疗人工智能教研室,技术大学慕尼黑(TUM)) Munich Center for Machine Learning (MCML), Munich, Germany(慕尼黑机器学习中心(MCML),慕尼黑,德国) Department of Computing, Imperial College London, London, UK(计算系,伦敦帝国学院,伦敦,英国) Deutsches Konsortium für Translationale Krebsforschung (DKTK), Partner Site Munich, Munich, Germany(德国转化癌症研究联盟(DKTK),慕尼黑合作伙伴站点,慕尼黑,德国) Institute of Radiation Medicine (IRM), Department of Radiation Sciences (DRS), Helmholtz Center Munich, Munich, Germany(辐射医学研究所(IRM),辐射科学部门(DRS),慕尼黑海德堡中心,慕尼黑,德国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 MultiMAE通过多模态掩码自编码器提升脑部MRI在缺失输入下的鲁棒性,实现分割和分类任务的性能提升。

Comments Official implementation: https://github.com/chris-beischl/multimae-for-brain-mri

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01954 2026-02-03 cs.CV 79%

Beyond Open Vocabulary: Multimodal Prompting for Object Detection in Remote Sensing Images

超越开放词汇:面向遥感图像的目标检测多模态提示

Shuai Yang, Ziyue Huang, Jiaxin Chen, Qingjie Liu, Yunhong Wang

机构 * School of Computer Science and Engineering, Beihang University, Beijing, China(计算机科学与工程学院,北京航空航天大学,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 RS-MPOD提出了一种多模态开放词汇检测框架,通过整合视觉和文本提示提升遥感图像中目标检测的稳定性与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01833 2026-02-03 cs.MM 79%

Mixture of Disentangled Experts with Missing Modalities for Robust Multimodal Sentiment Analysis

多模态情感分析中缺失模态的解耦专家混合方法

Xiang Li, Xiaoming Zhang, Dezhuang Miao, Xianfu Cheng, Dawei Li, Honggui Han, Zhoujun Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

AI总结 DERL通过解耦专家和多层次重建策略,提升多模态情感分析在缺失模态下的鲁棒性和表现

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00946 2026-02-03 cs.CV 79%

ConsensusDrop: Fusing Visual and Cross-Modal Saliency for Efficient Vision Language Models

ConsensusDrop:融合视觉与跨模态显著性以提高视觉语言模型的效率

Dhruv Parikh, Haoyang Fan, Rajgopal Kannan, Viktor Prasanna

机构 * University of Southern California(南加州大学) DEVCOM Army Research Office(陆军研究办公室)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 ConsensusDrop通过融合视觉与跨模态显著性,提高视觉语言模型的效率和准确性,优于现有token修剪方法。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22729 2026-02-02 cs.CV 79%

GaussianOcc3D: A Gaussian-Based Adaptive Multi-modal 3D Occupancy Prediction

GaussianOcc3D: 一种基于高斯的自适应多模态3D占用预测

A. Enes Doruk, Hasan F. Ates

机构 * Department of Artificial Intelligence and Data Engineering, Ozyegin University(人工智能与数据工程系,奥祖根大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 GaussianOcc3D通过高斯表示实现多模态3D占用预测,提升自动驾驶环境感知的鲁棒性和精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17564 2026-02-02 eess.IV cs.CV cs.LG 79%

ModalTune: Fine-Tuning Slide-Level Foundation Models with Multi-Modal Information for Multi-task Learning in Digital Pathology

ModalTune: 通过多模态信息细调滑片级基础模型以实现数字病理学中的多任务学习

Vishwesh Ramanathan, Tony Xu, Pushpak Pati, Faruk Ahmed, Maged Goubran, Anne L. Martel

机构 * Sunnybrook Research Institute(辛普森布鲁斯研究所在) University of Toronto(多伦多大学) Google Research(谷歌研究)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 ModalTune通过引入多模态信息和大型语言模型,实现数字病理学中多任务学习的统一细调框架,提升癌症生存和亚型预测性能。

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21673 2026-01-30 cs.CV 79%

Multimodal Visual Surrogate Compression for Alzheimer's Disease Classification

多模态视觉代理压缩用于阿尔茨海默病分类

Dexuan Ding, Ciyuan Peng, Endrowednes Kuantama, Jingcai Guo, Jia Wu, Jian Yang, Amin Beheshti, Ming-Hsuan Yang, Yuankai Qi

机构 * Macquarie University(麦考瑞大学) Federation University Australia(联邦大学澳大利亚) The Hong Kong Polytechnic University(香港理工大学) University of California at Merced(加州大学默塞德分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 MVSC通过压缩和适应大尺寸3D sMRI体积为紧凑的2D视觉代理,提升阿尔茨海默病分类性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21453 2026-01-30 cs.AI 79%

LION: A Clifford Neural Paradigm for Multimodal-Attributed Graph Learning

LION: 一种基于克莱因代数的多模态属性图学习神经范式

Xunkai Li, Zhengyu Wu, Zekai Chen, Henan Sun, Daohan Su, Guang Zeng, Hongchao Qin, Rong-Hua Li, Guoren Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 LION基于克莱因代数和解耦图神经范式,通过几何诱导的高阶图传播实现模态交互,结合自适应全息聚合模块提升多模态属性图学习性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21426 2026-01-30 cs.CV 79%

MultiModal Fine-tuning with Synthetic Captions

多模态微调与合成描述

Shohei Enomoto, Shin'ya Yamaguchi

机构 * NTT Tokyo(日本NTT东京)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出通过多模态大语言模型生成合成描述,提升多模态微调效果,尤其在少样本学习中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20305 2026-01-29 cs.AI 79%

Endogenous Reprompting: Self-Evolving Cognitive Alignment for Unified Multimodal Models

内生再提示:面向统一多模态模型的自进化认知对齐

Zhenchen Tang, Songlin Yang, Zichuan Wang, Bo Peng, Yang Li, Beibei Dong, Jing Dong

机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家) School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 内生再提示通过自进化框架提升统一多模态模型的认知对齐能力,有效解决生成过程引导问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20206 2026-01-29 cs.AI 79%

Towards Intelligent Urban Park Development Monitoring: LLM Agents for Multi-Modal Information Fusion and Analysis

面向智能城市公园发展监测:LLM代理用于多模态信息融合与分析

Zixuan Xiao, Chunguang Hu, Jun Ma

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 本研究提出基于LLM的多模态代理框架,用于提升城市公园发展监测中多模态数据的融合与分析能力,解决传统方法在高级分析和灵活适应性方面的不足。

Journal ref IEEE International Geoscience and Remote Sensing Symposium (IGARSS) 2025, Aug 3-8 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11740 2026-01-27 cs.LG cs.CV 79%

Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descent

通过模态解耦梯度下降缓解MLLM指令微调中的视觉知识遗忘

Junda Wu, Yuxin Xiong, Xintong Li, Yu Xia, Ruoyu Wang, Yu Wang, Tong Yu, Sungchul Kim, Ryan A. Rossi, Lina Yao, Jingbo Shang, Julian McAuley

机构 * UC San Diego(加州大学圣迭戈分校) University of New South Wales(新南威尔士大学) Adobe Research(Adobe研究)

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.CV

AI总结 本文提出模态解耦梯度下降方法,通过有效秩量化视觉退化并缓解过度压缩,保留预训练视觉知识并提升任务适应能力。

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05855 2026-01-26 cs.CV 79%

Decoupling Multi-Contrast Super-Resolution: Self-Supervised Implicit Re-Representation for Unpaired Cross-Modal Synthesis

解耦多对比超分辨率:自监督隐式重表示用于无配对跨模态合成

Yinzhe Wu, Hongyu Rui, Fanwen Wang, Jiahao Huang, Zhenxuan Zhang, Haosen Zhang, Zi Wang, Guang Yang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出了解耦多对比超分辨率框架,通过自监督隐式重表示实现无配对跨模态合成,提升极端尺度下的超分辨率性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15734 2026-01-23 cs.CV 79%

Sub-Region-Aware Modality Fusion and Adaptive Prompting for Multi-Modal Brain Tumor Segmentation

子区域感知模态融合与自适应提示的多模态脑肿瘤分割

Shadi Alijani, Fereshteh Aghaee Meibodi, Homayoun Najjaran

机构 * University of Victoria, BC, Canada(维多利亚大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出子区域感知模态融合与自适应提示方法,提升多模态脑肿瘤分割的准确性与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14448 2026-01-22 cs.CV 79%

Gaussian Based Adaptive Multi-Modal 3D Semantic Occupancy Prediction

基于高斯的自适应多模态3D语义占位预测

A. Enes Doruk

机构 * Abdullah Enes Doruk (2025)(Abdullah Enes Doruk)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

AI总结 本文提出一种基于高斯的自适应多模态3D语义占位预测模型,通过高效3D高斯模型融合相机与激光雷达数据,提升自动驾驶安全性能。

Comments Master Thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13331 2026-01-21 cs.CV cs.LG 79%

MultiST: A Cross-Attention-Based Multimodal Model for Spatial Transcriptomic

MultiST: 一种基于交叉注意力的多模态模型用于空间转录组

Wei Wang, Quoc-Toan Ly, Chong Yu, Jun Bai

机构 * Department of Computer Science, University of Cincinnati(计算机科学系,辛辛那提大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 MultiST通过基于交叉注意力的多模态融合方法,实现了对空间转录组数据中空间拓扑、基因表达和组织形态的联合建模,提升了空间域边界解析的准确性和生物解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12994 2026-01-21 cs.CV 79%

AsyncBEV: Cross-modal Flow Alignment in Asynchronous 3D Object Detection

AsyncBEV: 无感异步3D目标检测中的跨模态流对齐

Shiming Wang, Holger Caesar, Liangliang Nan, Julian F. P. Kooij

专题命中 多模态训练与对齐 :cross-modal(title);multi-modal(abstract);分类 cs.CV

AI总结 AsyncBEV通过跨模态流对齐提升3D目标检测在传感器异步情况下的鲁棒性,尤其在动态物体识别中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11885 2026-01-21 cs.AI 79%

MyGram: Modality-aware Graph Transformer with Global Distribution for Multi-modal Entity Alignment

MyGram: 多模态实体对齐的模态感知图变换器与全局分布

Zhifei Li, Ziyue Qin, Xiangyu Luo, Xiaoju Hou, Yue Zhao, Miao Zhang, Zhifang Huang, Kui Xiao, Bing Yang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 MyGram通过模态感知图变换器与全局分布机制,提升多模态实体对齐的性能。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏