Make Your LVLM KV Cache More Lightweight
使大型视觉-语言模型的KV缓存更轻量
机构 * Integrative Sciences and Engineering Programme, National University of Singapore(国立新加坡大学整合科学与工程学程) ; School of Computing, National University of Singapore(国立新加坡大学计算机学院)
AI总结 本文提出LightKV方法,通过利用视觉token嵌入的冗余性,减少KV缓存大小,提升解码效率并降低GPU内存消耗。
Comments Accepted to Transactions on Machine Learning Research (TMLR), 2026