CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models
CoVFT:面向多模态大语言模型的上下文感知视觉微调
机构 * State Key Laboratory of Complex and Critical Software Environment, Beihang University(北京航空航天大学复杂关键软件环境国家重点实验室) ; School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV
AI总结 本文提出CoVFT框架,通过整合上下文向量提取和上下文混合专家模块,解决多模态任务中视觉微调的不稳定性问题,实现更稳定的视觉更新。
Comments Accepted by CVPR 2026