arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6856 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6856 篇

2606.23885 2026-06-24 cs.CV cs.AI cs.CL cs.MM 新提交 87%

Mind the Heads: Topological Representation Alignment for Multimodal LLMs

注意头:多模态大语言模型的拓扑表示对齐

Davide Caffagni, Alberto Compagnoni, Federico Melis, Sara Sarto, Pier Luigi Dovesi, Mark Granroth-Wilding, Marcella Cornia, Lorenzo Baraldi

机构 * University of Modena and Reggio Emilia(摩德纳和雷焦艾米利亚大学) University of Pisa(比萨大学) AMD Silo AI

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract_cn);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出头级表示对齐(HeRA)方法,在注意力头级别强制跨模态对齐,通过对比目标匹配局部拓扑结构,选择对齐最差的头进行训练,有效提升视觉任务性能并减少幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08113 2026-06-11 cs.CL 87%

Multimodal LLMs Do Not Compose Skills Optimally Across Modalities

多模态大语言模型在跨模态间无法最优组合技能

Paula Ontalvilla, Aitor Ormazabal, Gorka Azkune

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);cross-modal(abstract);分类 cs.CL

AI总结 本文研究多模态大语言模型跨模态技能组合能力,发现其存在显著差距,并提出链式推理和微调策略以改善,但效果有限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09859 2026-06-10 cs.LG cs.AI 新提交 87%

Mitigating Manifold Departure: Uncertainty-Aware Subspace Rectification for Trustworthy MLLM Decoding

缓解流形偏离:面向可信MLLM解码的不确定性感知子空间校正

Yingxuan Zhuang, Jingxiao Yang, Miao Pan, Cheng Tan, Yuxiang Cai, Siwei Tan, Chen Zhi, Xuhong Zhang, Jianwei Yin, Jintao Chen

机构 * Nanyang Technological University(南洋理工大学)

专题命中 多模态训练与对齐 :MLLM(title,title_cn);multimodal(abstract);分类 cs.AI

AI总结 提出MGAP方法,通过SVD构建语言先验子空间并自适应衰减投影分量,在抑制幻觉的同时保持语义结构,优于现有解码基线。

Comments ICML 2026 regular

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21670 2026-05-26 cs.CV cs.LG 87%

Diverse via bounded Agreement: Geometric Regularization for Multimodal Fusion

通过有界一致性实现多样性:多模态融合的几何正则化

Zixuan Xia, Hao Wang, Pengcheng Weng, Yanyu Qian, Yangxin Xu, William Dan, Fei Wang

机构 * Department of Informatics University of Bern(伯尔尼大学信息学院) College of Computing and Data Science Nanyang Technological University(南洋理工大学计算机与数据科学学院) School of Software Engineering Xi’an Jiaotong University(西安交通大学软件工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);audio-visual(abstract)

AI总结 提出一种轻量级即插即用的几何正则化框架,通过有界一致性原则在保持模态特异多样性的同时约束跨模态漂移,提升多模态融合性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18041 2026-05-19 cs.CV 87%

OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models

OmniSelect: 动态模态感知的令牌压缩用于高效多模态大语言模型

Morunliu Yang, Ruotao Xu, Le Li, Yue Wang, Jianxin Zhang, Juntao Li, Yihang Lou, Siwei Feng, Peifeng Li

机构 * Soochow University(苏州大学) Peking University(北京大学)

专题命中 多模态训练与对齐 :omni-modal(title);multimodal(abstract,abstract_cn);cross-modal(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出OmniSelect,一种无需训练的模态自适应令牌剪枝框架,通过动态选择压缩策略来提高多模态大语言模型的效率,通过轻量级AudioCLIP模型估计跨模态相关性,并根据相关性得分在不同时间组中进行细粒度令牌剪枝,从而在不增加训练成本的情况下实现高效的多模态令牌压缩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17497 2026-05-12 cs.LG cs.AI 87%

Multimodal Representation Learning Conditioned on Semantic Relations

基于语义关系的多模态表示学习

Yang Qiao, Yuntong Hu, Bowen Zhu, Hasibul Haque, Liang Zhao

机构 * Emory University(埃默里大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.AI

AI总结 本文提出RCML框架,通过将语义关系作为显式条件,使多模态表示学习能根据不同的关系上下文生成不同表示,提升检索和分类任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05522 2026-04-08 cs.CL 87%

Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs

跨模态指代对齐:在全模态大语言模型中实现可靠信息传输

Hongcheng Liu, Yuhao Wang, Zhe Chen, Pingjie Wang, Zhiyuan Zhu, Yixuan Hou, Yanfeng Wang, Yu Wang

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);omni-modal(abstract)

AI总结 研究提出跨模态指代对齐问题,通过CrossOmni数据集和两种方法提升全模态推理能力,揭示指代意识缺失是跨模态推理弱化的主要原因。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06965 2026-03-13 cs.CV 87%

MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images

MedMO:为医学图像构建和理解多模态大语言模型

Ankan Deria, Komal Kumar, Adinath Madhavrao Dukre, Eran Segal, Salman Khan, Imran Razzak

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);multimodal foundation model(abstract)

AI总结 MedMO是一种基于通用MLLM架构构建的医学多模态基础模型,通过多阶段训练提升跨模态和任务的性能,超越现有开源基线,在医学图像识别和报告生成中取得显著提升。

Comments 21 pages, 6 figures and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00969 2025-12-17 cs.LG cs.AI 87%

Masked Omics Modeling for Multimodal Representation Learning across Histopathology and Molecular Profiles

掩码组学建模用于病理学与分子特征的多模态表示学习

Lucas Robinet, Ahmad Berjaoui, Elizabeth Cohen-Jonathan Moyal

机构 * Oncopole(奥恩波尔) IRT Saint Exupéry(国际研究与技术圣埃克苏佩里) INSERM Cancer Research Center of Toulouse(里沃利癌症研究中心) Toulouse(图卢兹)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);any-to-any(abstract);multimodal foundation model(abstract)

AI总结 MORPHEUS通过整合病理学图像和多组学数据,提出了一种多模态预训练策略,以提升癌症研究中的跨模态表示学习能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11273 2024-05-21 cs.AI cs.CL cs.CV cs.MM 87%

Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts

Yunxin Li, Shenyuan Jiang, Baotian Hu, Longyue Wang, Wanqi Zhong, Wenhan Luo, Lin Ma, Min Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 22 pages, 13 figures. Project Website: https://uni-moe.github.io/. Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03170 2024-03-12 cs.MM cs.AI cs.CL cs.CV cs.CY 87%

SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection

Peng Qi, Zehong Yan, Wynne Hsu, Mong Li Lee

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments To appear in CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12595 2026-06-12 cs.LG cs.AI cs.CV 新提交 87%

Emerging Flexible Designs for Geospatial Multimodal Foundation Models

地理空间多模态基础模型的新兴灵活设计

Philipe Dias, Waqwoya Abebe, Abhishek Potnis, Aristeidis Tsaris, Dan Lu, Xiao Wang, Dalton Lunga

机构 * Oak Ridge National Laboratory(橡树岭国家实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.CV、cs.AI

AI总结 本文系统比较了不同架构的地理空间基础模型,在统一设置下评估其灵活性与性能,为多模态推理提供设计指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00496 2025-12-02 cs.CL cs.AI 87%

CACARA: Cross-Modal Alignment Leveraging a Text-Centric Approach for Cost-Effective Multimodal and Multilingual Learning

CACARA:基于文本中心方法的跨模态对齐,用于高效多模态和多语言学习

Diego A. B. Moreira, Alef I. Ferreira, Jhessica Silva, Gabriel O. dos Santos, Gustavo Bonil, João Gondim, Marina dos Santos, Helena Maia, Simone Hashiguti, Nádia da Silva, Carolina Scarton, Helio Pedrini, Sandra Avila

机构 * Instituto de Computação, Universidade Estadual de Campinas (UNICAMP), Brasil(计算机学院,Campinas州立大学(UNICAMP)) Instituto de Estudos da Linguagem, Universidade Estadual de Campinas (UNICAMP), Brasil(语言研究学院,Campinas州立大学(UNICAMP)) Instituto de Informática, Universidade Federal de Goiás (UFG), Goiás, Brasil(信息学院,戈亚斯联邦大学(UFG)) Department of Computer Science, University of Sheffield, Sheffield, United Kingdom(计算机科学系,谢菲尔德大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title);分类 cs.CL、cs.AI

AI总结 CACARA通过文本中心方法实现多模态和多语言学习,无需重新训练即可支持100多种语言,提升音频到文本检索性能达14.24个百分点。

Comments 25 pages, 12 tables, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05991 2025-08-11 cs.CV cs.AI cs.CY 87%

ECMF: Enhanced Cross-Modal Fusion for Multimodal Emotion Recognition in MER-SEMI Challenge

Juewen Hu, Yexin Li, Jiulin Li, Shuo Chen, Pring Wong

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(人工智能通用基础理论国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14717 2025-03-11 cs.LG cs.CL cs.CV 87%

FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data

Binqian Xu, Xiangbo Shu, Haiyang Mei, Guosen Xie, Basura Fernando, Jinhui Tang

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16033 2024-04-25 cs.CV cs.CL 87%

Cantor: Inspiring Multimodal Chain-of-Thought of MLLM

Timin Gao, Peixian Chen, Mengdan Zhang, Chaoyou Fu, Yunhang Shen, Yan Zhang, Shengchuan Zhang, Xiawu Zheng, Xing Sun, Liujuan Cao, Rongrong Ji

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(title);分类 cs.CV、cs.CL

Comments The project page is available at https://ggg0919.github.io/cantor/

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19654 2024-04-03 cs.CV cs.AI 87%

MCAD: Multi-teacher Cross-modal Alignment Distillation for efficient image-text retrieval

Youbo Lei, Feifei He, Chen Chen, Yingbin Mo, Si Jia Li, Defeng Xie, Haonan Lu

专题命中 多模态训练与对齐 :image-text(title,abstract);cross-modal(title);分类 cs.CV、cs.AI

Comments Accepted by NAACL 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05909 2026-08-07 cs.CR 新提交 87%

MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration

MMAligner:通过表示校准保护多模态大语言模型

Shenyi Zhang, Keyan Guo, Zihao Wang, Xuebin Li, Lingchen Zhao, Hongxin Hu, Chao Shen, Qian Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn)

AI总结 MMAligner通过校准多模态大语言模型的表示,将不安全多模态输入的拒绝率提升至99%,仅造成不足2%的效用下降,显著优化了安全与效用的权衡。

Comments To Appear in the Proceedings of The ACM Conference on Computer and Communications Security (CCS), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04949 2026-08-06 cs.CV cs.CL cs.IT cs.MM math.IT 新提交 87%

UG-UMRE: Uncertainty-Guided Modality Augmentation and Distributional Calibration for Unified Multimodal Relation Extraction

UG-UMRE:面向统一多模态关系抽取的不确定性引导模态增强与分布校准

Bo Kong, Liruiz Jia, Yi Liang, Chao Liu, Dongfang Han, Tianwei Yan, Yuan Liu, Shengquan Liu

机构 * Xinjiang University(新疆大学) Chongqing Jiaotong University(重庆交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 针对统一多模态关系抽取的噪声传播与模态分布异质性问题,提出含UDUA和JAUA模块的UG-UMRE网络,在UMRE等三个基准数据集上实现最优性能,模块可插拔且有效。

Comments Accepted at ACM MM2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08277 2026-05-21 q-bio.NC cs.AI cs.CL cs.CV cs.LG 87%

Task-conditioned probing of instruction-tuned multimodal LLMs: Region-specific brain alignment patterns under naturalistic stimuli

基于任务的指令调制多模态大语言模型探测:在自然主义刺激下的区域特定大脑对齐模式

Subba Reddy Oota, Khushbu Pahwa, Prachi Jindal, Satya Sai Srinath Namburi, Maneesh Singh, Tanmoy Chakraborty, Bapi S. Raju, Manish Gupta

机构 * Technische Universität Berlin(柏林技术大学) Rice University(Rice 大学) AWS AI Labs, Amazon(Amazon 人工智能实验室) IIT Delhi(德里理工学院) University of Wisconsin - Madison(威斯康星大学麦迪逊分校) Spector Inc(Spector 公司) IIIT-Hyderabad(海得拉巴理工学院) Microsoft(微软)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.CL、cs.AI

AI总结 本研究探讨了指令调制多模态大语言模型在自然主义刺激下的大脑对齐模式,通过比较不同模型在视频和音频任务中的表现,揭示了指令调制对模型表示能力的影响。

Comments 57 pages, 39 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11301 2026-05-13 cs.AI cs.CL cs.CV 87%

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer?

LatentRouter: 在看到答案之前,我们能否选择合适的多模态模型?

Xueqi Cheng, Yushun Dong

机构 * Department of Computer Science(计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.CL、cs.AI

AI总结 LatentRouter通过多模态效用预测实现多模态大语言模型的路由,通过隐式通信和胶囊修正提升模型选择准确性,在MMR-Bench和VL-RouterBench实验中优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06728 2026-04-09 cs.CV cs.AI cs.MM 87%

URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection

URMF:面向多模态讽刺检测的不确定性感知鲁棒多模态融合

Zhenyu Wang, Weichen Cheng, Weijia Li, Junjie Mou, Zongyou Zhao, Guoying Zhang

机构 * School of Artificial Intelligence, China University of Mining and Technology-Beijing(中国矿业大学(北京)人工智能学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文提出URMF框架,通过建模模态可靠性提升多模态讽刺检测的准确性和鲁棒性,采用多头交叉注意力和自注意力机制,并结合不确定性建模和联合训练目标,在公开基准上优于现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14129 2024-11-26 cs.CL cs.AI cs.CV 87%

AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability

Fei Zhao, Taotian Pang, Chunhui Li, Zhen Wu, Junjie Guo, Shangyu Xing, Xinyu Dai

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);cross-modal(abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17854 2024-07-26 cs.AI cs.CL cs.MM 87%

Shapley Value-based Contrastive Alignment for Multimodal Information Extraction

Wen Luo, Yu Xia, Shen Tianshu, Sujian Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CL、cs.AI、cs.MM

Comments Accepted at ACM Multimedia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02570 2023-08-08 cs.LG cs.AI cs.CL cs.CV 87%

Learning Implicit Entity-object Relations by Bidirectional Generative Alignment for Multimodal NER

Feng Chen, Jiajia Liu, Kaixiang Ji, Wang Ren, Jian Wang, Jingdong Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04921 2025-07-08 cs.CV cs.CL 87%

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Yunxin Li, Zhenyu Liu, Zitao Li, Xuanyu Zhang, Zhenran Xu, Xinyu Chen, Haoyuan Shi, Shenyuan Jiang, Xintong Wang, Jifang Wang, Shouzheng Huang, Xinping Zhao, Borui Jiang, Lanqing Hong, Longyue Wang, Zhuotao Tian, Baoxing Huai, Wenhan Luo, Weihua Luo, Zheng Zhang, Baotian Hu, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);omni-modal(abstract);分类 cs.CV、cs.CL

Comments v2, 91 Pages, 10 figures; Project: https://github.com/HITsz-TMG/Awesome-Large-Multimodal-Reasoning-Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10412 2026-06-10 cs.AI 新提交 86%

A Unified Multi-Modal Framework for Intelligent Financial Systems: Integrating Reinforcement Learning, High-Frequency Trading, and Game-Theoretic Approaches with Cross-Modal Sentiment Analysis

面向智能金融系统的统一多模态框架:整合强化学习、高频交易和博弈论方法与跨模态情感分析

Fanrong Liu, Zhang Yuwei, Mingni Luo

机构 * Henan University, International Eurasia College(河南大学,国际欧亚学院) City University of Hong Kong, College of Business(香港城市大学,商学院) Northeastern University, School of Electronic and Information Engineering(东北大学,电子与信息工程学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(title);分类 cs.AI

AI总结 提出统一框架整合PPO、高频预测、上下文学习、博弈论和跨模态情感分析,在多个金融任务上平均提升20%以上性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22251 2026-05-14 cs.LG cond-mat.mtrl-sci cs.AI 86%

Zatom-1: Towards a Multimodal Foundation Model for 3D Molecules and Materials

Zatom-1:迈向3D分子和材料的多模态基础模型

Alex Morehead, Miruna Cretu, Antonia Panescu, Rishabh Anand, Maurice Weiler, Tynan Perez, Samuel Blau, Steven Farrell, Wahid Bhimji, Anubhav Jain, Hrushikesh Sahasrabuddhe, Pietro Lio, Tommi Jaakkola, Rafael Gomez-Bombarelli, Rex Ying, N. Benjamin Erichson, Michael W. Mahoney

机构 * LBNL(劳伦斯伯克利国家实验室) ICSI(国际计算机科学研究所) University of Cambridge(剑桥大学) Yale University(耶鲁大学) MIT(麻省理工学院) UC Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.AI

AI总结 Zatom-1是一种跨领域、通用的模型架构,统一了3D分子和材料的生成与预测学习,通过多模态流匹配目标实现高效的预训练和稳定采样,提升了生成推理速度和跨领域预测性能。

Comments 38 pages, 10 figures, 15 tables. ICLR 2026 FM4Science. Code, data, and model weights are available at https://github.com/Zatom-AI/zatom

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17980 2026-05-08 cs.CV 86%

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding

感知空间:面向高效准确3D场景理解的自我运动感知视频表示

Shuyao Shi, Kang G. Shin

机构 * Department of Computer Science(计算机科学系) University of Michigan(密歇根大学) University of Michigan Ann Arbor(密歇根大学安阿伯分校)

专题命中 多模态训练与对齐 :MLLM(summary_cn,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出Motion-MLLM框架,结合IMU数据与视觉特征,通过运动-视觉关键帧过滤模块和异构跨模态融合模块,提升3D场景理解与空间推理的效率和准确性。

Comments 22 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20032 2026-04-30 cs.CV cs.LG cs.RO 86%

ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers

ViTaPEs: 视觉触觉位置编码用于多模态转换器中的跨模态对齐

Fotios Lygerakis, Ozan Özdenizci, Elmar Rückert

机构 * Chair of Cyber-Physical Systems(感知系统教授席位) Institute of Machine Learning and Neural Computation(机器学习与神经计算研究所) Technical University of Leoben(莱比锡技术大学) Graz University of Technology(格拉茨技术大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(title);分类 cs.CV

AI总结 ViTaPEs通过两阶段位置注入学习任务无关的视觉触觉表示,提升多模态对齐性能,实验证明其在多种任务和领域中的优越表现。

详情

展开后加载摘要…

URL PDF HTML 收藏