OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond
OScaR:LLMs及更广泛场景中的极压缩KV缓存量化之奥卡姆之刀
机构 * Tsinghua University(清华大学) ; Meituan LongCat Team(美团LongCat团队) ; The University of Hong Kong(香港大学) ; The University of Edinburgh(爱丁堡大学) ; UCAS(中国科学技术大学) ; The Hong Kong Polytechnic University(香港理工大学)
AI总结 本文针对LLMs中KV缓存极压缩时的量化保真问题,提出OScaR框架,通过Canalized Rotation和Omni-Token Scaling有效缓解Token Norm Imbalance,实现近无损的INT2量化性能,同时提升解码速度和吞吐量。
Comments Under review