ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
ABMAMBA: 多模态大语言模型中的对齐层次双向扫描用于高效的视频描述
机构 * Keio University(庆应义塾大学) ; National Institute of Informatics(国立信息学研究所) ; National Institute of Informatics Research and Development Center for Large Language Models(国立信息学研究所大型语言模型研发中心)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
AI总结 本文提出ABMamba,一种具有线性计算复杂度的多模态大语言模型,通过替代二次注意力机制,实现视频序列的高效处理,在视频描述任务中表现出色。
Comments Accepted to ICPR 2026