CentroidKV: Efficient Long-Context LLM Inference via KV Cache Clustering
CentroidKV: 通过KV缓存聚类实现高效的长上下文LLM推理
机构 * Peking University(北京大学) ; Huawei Technologies(华为技术有限公司) ; University of Science and Technology of China(中国科学技术大学)
AI总结 提出CentroidKV框架,通过在线聚类KV缓存减少内存占用,在保持性能的同时实现高达75%的缓存压缩和1.92倍解码加速。