Query-based Cross-Modal Projector Bolstering Mamba Multimodal LLM
基于查询的跨模态投影器增强Mamba多模态大语言模型
机构 * Korea Advanced Institute of Science and Technology / Korea, Republic of(韩国科学技术院) ; University of Illinois in Urbana-Champaign / United States of America(伊利诺伊大学厄巴纳-香槟分校) ; Korea University / Korea, Republic of(韩国大学)
专题命中 VLM训练与架构 :vision-language model(abstract)
AI总结 提出基于查询的跨模态投影器,通过交叉注意力压缩视觉令牌,消除手动设计2D扫描顺序的需求,提升Mamba多模态LLM的性能和吞吐量。
Comments Accepted to EMNLP 2024 Findings