arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6856 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6856 篇

2308.13437 2023-09-15 cs.CV 85%

Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models

Chi Chen, Ruoyu Qin, Fuwen Luo, Xiaoyue Mi, Peng Li, Maosong Sun, Yang Liu

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.14865 2023-03-28 cs.CV 85%

Revisiting Multimodal Representation in Contrastive Learning: From Patch and Token Embeddings to Finite Discrete Tokens

Yuxiao Chen, Jianbo Yuan, Yu Tian, Shijie Geng, Xinyu Li, Ding Zhou, Dimitris N. Metaxas, Hongxia Yang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted to CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12644 2023-01-31 cs.CV 85%

Tagging before Alignment: Integrating Multi-Modal Tags for Video-Text Retrieval

Yizhen Chen, Jie Wang, Lijian Lin, Zhongang Qi, Jin Ma, Ying Shan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted to AAAI 2023 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.00332 2021-04-02 cs.CV 85%

UC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-training

Mingyang Zhou, Luowei Zhou, Shuohang Wang, Yu Cheng, Linjie Li, Zhou Yu, Jingjing Liu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.06238 2016-03-03 cs.LG cs.CV stat.ML 85%

Multimodal sparse representation learning and applications

Miriam Cha, Youngjune Gwon, H. T. Kung

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10884 2024-11-06 cs.CL cs.AI cs.CV cs.LG 85%

Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models

Shengzhi Li, Rongyu Lin, Shichao Pei

专题命中 多模态训练与对齐 :multi-modal(title,abstract);MLLM(abstract,comments);分类 cs.CV、cs.CL、cs.AI

Comments Project code, model and data: https://github.com/findalexli/mllm-dpo

Journal ref Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14188-14200, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01473 2026-08-04 cs.CV cs.CL cs.LG 新提交 85%

Slot2Text: Object-Centric Visual Tokenization for Efficient and Spatially Traceable Surgical MLLMs

Slot2Text:面向高效且空间可追溯的手术多模态大语言模型的以对象为中心的视觉分词

Guiqiu Liao, Matjaz Jogan, Daniel A. Hashimoto

专题命中 多模态训练与对齐 :MLLM(summary_cn,abstract);multimodal(abstract);分类 cs.CV、cs.CL

AI总结 Slot2Text将视觉输入转为槽潜变量,推出双模式手术MLLM,在多基准上实现高效推理,Slot2Text-Fast大幅降本,Slot2Text-Reason支持可追溯空间推理。

Comments 17 pages, 8 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21420 2026-04-30 cs.CV cs.CL 85%

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs

ReGATE:通过减少令牌数在MLLMs中实现更快更好的学习

Chaoyu Li, Yogesh Kulkarni, Pooyan Fazli

机构 * Arizona State University(亚利桑那州立大学)

专题命中 多模态训练与对齐 :MLLM(summary_cn,abstract);multimodal(abstract);分类 cs.CV、cs.CL

AI总结 ReGATE通过参考引导的自适应令牌消除方法加速MLLM训练,减少计算量同时保持模型准确性,在多个基准测试中表现出色。

Comments ACL 2026. Project page: https://people-robots.github.io/regate

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09270 2026-08-11 cs.CV cs.AI cs.IR cs.MM 新提交 85%

GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views

GRASP:面向无人机视角细粒度跨模态理解的粒度感知区域对齐与语义原型学习

Jiahui Cui, Yan Zhao, Kan Wei, Enze Zhu, Peirong Zhang, Lei Wang, Yiru Wang

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院空天信息创新研究院) University of Chinese Academy of Sciences(中国科学院大学) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 针对无人机视角细粒度跨模态理解的背景杂波干扰与视觉同构歧义问题,提出GRASP框架,通过RFA和SPEM策略提升性能,在相关基准数据集上验证了有效性。

Comments Accepted at the 34th ACM International Conference on Multimedia (ACM Multimedia 2026, MM '26). 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07279 2026-08-10 eess.SP 新提交 85%

Token Communication for Multimodal Large Language Model

多模态大语言模型的令牌通信

Jingkai Ying, Zhijin Qin, Yuan Shen, Khaled B. Letaief

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn)

AI总结 本文针对多模态大语言模型(MLLMs)的令牌传输问题,提出适配MLLMs的令牌通信框架,通过集成神经编解码器、两阶段语义对齐训练及自适应适配器,在相同传输数据量下实现更优任务性能。

Comments 13 pages, 11 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04054 2026-08-06 cs.MM cs.AI cs.CL cs.LG 新提交 85%

Modality Agreement- and Conflict-Aware Prototype Hypergraph Learning for Multimodal Intent Understanding

面向多模态意图理解的模态一致与冲突感知原型超图学习

Mohnish Raj, Suraj Kumar, Soumi Chattopadhayay, Chandranath Adak, Ayan Dutta

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 针对多模态意图理解中分歧信息被多数融合方法忽略的问题,提出MACH框架,通过分层一致与冲突原型超图及自适应仲裁机制,在基准数据集上验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03475 2026-08-05 cs.MM cs.AI cs.CL 新提交 85%

Adaptive Modality Reliability Diagnosis and Restoration for Robust Multimodal Intent Recognition

用于鲁棒多模态意图识别的自适应模态可靠性诊断与恢复

Suraj Kumar, Mohnish Raj, Soumi Chattopadhayay, Chandranath Adak, Ayan Dutta

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 提出PRIME闭环可靠性引导框架,通过诊断恢复多模态意图识别的不可靠模态,在保持干净数据性能的同时提升了多模态数据各类退化场景下的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02916 2026-07-31 cs.IR 版本更新 85%

Towards Transfer-Efficient Multi-modal Sequential Recommendation with State Space Duality

面向高效多模态序列推荐的态空间双重视角

Hao Fan, Qingyang Liu, Hongjiu Liu, Yanrong Hu, Kai Fang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract)

AI总结 本文提出MMM4Rec模型,通过结合态空间双重视角和全局时间建模,提升多模态序列推荐的迁移效率与准确性,实验表明其在大规模下游数据集上收敛速度提升10倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01390 2026-07-03 cs.CV cs.AI cs.MM 版本更新 85%

SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment

SEPS:面向细粒度跨模态对齐的语义增强补丁精简框架

Xinyu Mao, Junsi Li, Haoji Zhang, Yu Liang, Ming Sun

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 提出SEPS框架,通过两阶段机制整合稠密与稀疏文本语义,识别显著视觉补丁,并利用相关性感知选择与均值计算突出关键补丁-词对应,提升跨模态相似度评估,在Flickr30K和MS-COCO上rSum提升23%-86%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24596 2026-06-15 eess.AS cs.AI cs.CL 版本更新 85%

X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs

X-OPD:面向语音大语言模型能力对齐的跨模态在策略蒸馏

Di Cao, Dongjie Fu, Hai Yu, Siqi Zheng, Xu Tan, Tao Jin

机构 * Tencent Hunyuan(腾讯文心) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 提出X-OPD框架,通过跨模态在策略蒸馏对齐语音LLM与文本LLM的能力,利用文本教师模型评估语音模型的轨迹并提供令牌级反馈,显著缩小复杂任务性能差距。

Comments Accepted by Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07026 2026-06-08 cs.CV cs.AI cs.MM 版本更新 85%

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

模态间隙驱动的子空间对齐训练范式用于多模态大语言模型

Xiaomin Yu, Yi Xin, Yuhui Zhang, Wenjie Zhang, Chonghan Liu, Hanzhen Zhao, Chen Liu, Xiaoxing Hu, Ziyue Qiao, Hao Tang, Xiaobin Hu, Chengwei Qin, Hui Xiong, Yu Qiao, Shuicheng Yan

机构 * HKUST(GZ)(香港科技大学(广州)) NUS(新加坡国立大学) sh AILab SII Stanford(斯坦福大学) UCLA(加州大学洛杉矶分校) Yale(耶鲁大学) SJTU(上海交通大学) GBU(国防大学) PKU(北京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 针对多模态对比学习中的模态间隙问题,提出固定帧模态间隙理论,并基于该理论设计无训练的对齐策略ReAlign和可扩展训练范式ReVision,利用无配对数据实现视觉与语言表示的高效对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04448 2026-06-04 cs.IR 85%

Bridging Short Videos and Live Streams: Reasoning-Guided Multimodal LLMs for Cross-Domain Representation Learning

连接短视频与直播:推理引导的多模态大语言模型用于跨域表示学习

Le Zhang, Xiaolan Zhu, Yuchen Wang, Shilong Kang, Jiaqi Xue, Xiaoyu Zhang, Xiang Chen, Yalong Guan, Xiangyu Wu, Shijun Wang, Lantao Hu, Kun Gai

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn)

AI总结 针对短视频到直播的跨域推荐冷启动问题,提出推理引导的跨域表示学习框架RGCD-Rep,利用多模态大语言模型生成结构化推理知识并蒸馏到轻量模型,通过两阶段训练学习可迁移的物品表示,在快手部署后显著提升核心业务指标。

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18257 2026-05-19 cs.CV cs.AI cs.CL 85%

CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook

CodeBind: 一种用于多模态对齐的解耦表示学习框架

Zeyu Chen, Jie Li, Kai Han

机构 * Visual AI Lab, The University of Hong Kong(视觉人工智能实验室,香港大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 CodeBind通过统一的组合代码本设计优化多模态表示空间,解决了传统方法在跨模态信息差异和数据稀缺导致的对齐空间不足问题,实现了多模态分类和检索任务中的最佳性能。

Comments ACL 2026 Findings; Project page: https://visual-ai.github.io/codebind

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12258 2026-05-13 cs.LG 85%

Instruction Lens Score: Your Instruction Contributes a Powerful Object Hallucination Detector for Multimodal Large Language Models

指令透镜分数:您的指令为多模态大语言模型提供了一个强大的对象幻觉检测器

Runhe Lai, Xinhua Lu, Yanqi Wu, Jinlun Ye, Weijiang Yu, Ruixuan Wang

机构 * School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China(中山大学计算机科学与工程学院,广州,中国) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国) Key Laboratory of Machine Intelligence and Advanced Computing, MOE, Guangzhou, China(机器智能与高级计算关键实验室,教育部,广州,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn)

AI总结 本文提出InsLen,通过结合校准局部分数和上下文一致性分数,有效检测多模态大语言模型中的对象幻觉,无需额外训练或辅助模型。

Comments Accepted by ICML-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12067 2026-04-14 cs.LG cs.AI cs.CL cs.CV 85%

MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

MM-LIMA:多模态数据集对齐中的‘少即是多’

Lai Wei, Xiaozhe Li, Zihao Jiang, Weiran Huang, Lichao Sun

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学与工程学院) Shanghai Innovation Institute(上海创新研究院) Lehigh University(里海大学) Tongji University(同济大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出MM-LIMA,仅用200个示例(约6%的指令数据)训练,通过数据选择器过滤低质量数据,使模型在多项评估中超越MiniGPT-4,证明高质量少量指令数据的有效性。

Comments Published at Artificial Intelligence for Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.27693 2026-03-31 cs.CV cs.AI cs.LG cs.MA cs.MM 85%

LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation

LVRPO:基于GRPO的语言-视觉对齐用于多模态理解和生成

Shentong Mo, Sukmin Yun

机构 * Department of Machine Learning, CMU, USA(卡内基梅隆大学机器学习系,美国) Department of Machine Learning, MBZUAI, UAE(穆罕默德·本·扎耶德人工智能大学机器学习系,阿联酋) Department of Artificial Intelligence, Hanyang University ERICA, South Korea(汉阳大学ERICA校区人工智能系,韩国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文提出LVRPO框架,通过GRPO显式对齐语言和视觉表示,提升多模态理解和生成性能,无需辅助编码器或人工目标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13879 2026-03-12 cs.MM cs.CL cs.CV 85%

Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring

链式推理压缩不应盲目:通过双路径锚定实现高效的多模态推理的V-Skip

Dongxu Zhang, Yiding Sun, Cheng Tan, Wenbiao Yan, Ning Yang, Jihua Zhu, Haijun Zhang

机构 * School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Harbin Institute of Technology, Shenzhen(深圳哈尔滨工业大学) University of Science and Technology Beijing(北京科技大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 V-Skip通过双路径锚定机制解决多模态推理中令牌剪枝的盲目性问题,实现高效的推理速度提升与精度保持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06403 2026-03-09 cs.LG 85%

Adapter-Augmented Bandits for Online Multi-Constrained Multi-Modal Inference Scheduling

适配器增强的带状机用于在线多约束多模态推断调度

Xianzhi Zhang, Yue Xu, Yinlin Zhu, Di Wu, Yipeng Zhou, Miao Hu, Guocong Quan

机构 * School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, 510006, China.(中山大学计算机科学与工程学院) School of Computing, Macquarie University, NSW 2109, Australia(麦考瑞大学计算机学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);MLLM(abstract)

AI总结 M-CMAB通过多适配器增强框架,实现多模态任务调度中的多约束优化,提升推断效率和预算利用效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19605 2026-02-24 cs.CV cs.AI cs.MM 85%

CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning

CLCR:多模态学习中的跨层语义协作表示

Chunlei Meng, Guanhong Huang, Rong Fu, Runmin Jian, Zhongxue Gan, Chun Ouyang

机构 * Fudan University(复旦大学) Shantou University(汕头大学) University of Macau(澳门大学) Guangzhou Huashang College(广州华商学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 CLCR通过跨层语义协同表示方法,有效解决多模态数据中的语义错位问题,提升多模态学习的表示质量与任务泛化能力。

Comments This study has been Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10138 2026-02-12 cs.CV cs.AI cs.CL cs.LG 85%

Multimodal Information Fusion for Chart Understanding: A Survey of MLLMs -- Evolution, Limitations, and Cognitive Enhancement

多模态信息融合用于图表理解:MLLMs的综述——演变、局限与认知增强

Zhihang Yi, Jian Zhao, Jiancheng Lv, Tao Wang

机构 * College of Computer Science, Sichuan University(四川大学计算机科学学院) Engineering Research Center of Machine Learning(机器学习工程研究中心) China Telecom Institute of AI(中国电信人工智能研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文综述了多模态大语言模型在图表理解中的应用,分析了其演变、局限及未来发展方向,旨在推动更稳健的系统发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25801 2026-02-02 cs.LG cs.AI cs.CL cs.CV 85%

Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start

Metis-SPECS: 通过基于偏好自我蒸馏的冷启动解耦多模态学习

Kun Chen, Peng Shi, Haibo Qiu, Zhixiong Zeng, Siqi Yang, Wenji Mao, Lin Ma

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) Meituan(美团)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 Metis-SPECS通过基于偏好的自我蒸馏冷启动框架解耦多模态学习,提升泛化能力和下游RL表现。

Comments Published as a conference paper at ICLR 2026!

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17986 2026-01-27 cs.LG 85%

Federated learning for unpaired multimodal data through a homogeneous transformer model

通过同质Transformer模型实现无配对多模态数据的联邦学习

Anders Eklund

机构 * Department of Biomedical Engineering, Linköping University, Sweden(_linköping大学生物医学工程系) Department of Computer and Information Science, Linköping University, Sweden(_linköping大学计算机与信息科学系) Center for Medical Image Science and Visualization (CMIV), Linköping University, Sweden(_linköping大学医学影像科学与可视化中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);multimodal foundation model(abstract)

AI总结 本文提出通过同质Transformer模型实现联邦学习,解决无配对多模态数据的训练问题,通过公共锚点对齐和子空间稳定化微调方法,在不传输私有数据的情况下实现全局模型统一表示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13659 2026-01-21 cs.CL cs.AI cs.MM 85%

Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis

在动作前解耦时序与空间:用于多模态情感分析的解耦表示学习

Chunlei Meng, Ziyang Zhou, Lucas He, Xiaojing Du, Chun Ouyang, Zhongxue Gan

机构 * Fudan University(复旦大学) Shantou University(汕尾大学) University College London(伦敦大学学院) University of South Australia(澳大利亚南澳大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 本文提出TSDA方法,在动作前解耦时序与空间信息,通过因子一致对齐和门控重联模块提升多模态情感分析性能。

Comments This study has been accepted by IEEE ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22829 2025-10-28 cs.CV cs.AI cs.MM 85%

LLM-based Fusion of Multi-modal Features for Commercial Memorability Prediction

Aleksandar Pramov

机构 * Georgia Institute of Technology, USA(佐治亚理工学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17391 2025-10-01 cs.CV cs.AI cs.CL 85%

LFTR: Learning-Free Token Reduction for Multimodal Large Language Models

Zihui Zhao, Yingxin Li, Yang Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏