DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
DepthKV:面向长上下文LLM推理的分层KV缓存压缩
机构 * Ruhr University Bochum(博德姆鲁尔大学) ; UAR Research Center for Trustworthy Data Science and Security(UAR可信数据科学与安全研究中心)
专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI
AI总结 本文提出DepthKV,通过分层敏感度分配优化KV缓存预算,提升长上下文LLM推理效率,优于统一剪枝方法。