MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models
MAIL++: 视觉语言模型的多模态双向智能体层
机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) ; Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用国家重点实验室(东南大学),中华人民共和国教育部,中国)
专题命中 图文多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV
AI总结 提出MAIL/MAIL++方法,通过将跨模态耦合嵌入VLM内在计算模块并引入双向桥接,实现参数高效微调,在少样本分类和跨域检索中超越现有方法。