QAM-W: Joint 2D Codebook Quantization for LLM Weights via Hadamard Rotation and Activation-Aware Scaling
QAM-W: 通过哈达玛旋转和激活感知缩放实现LLM权重的联合2D码本量化
机构 * Independent Research(独立研究) ; Institute of Computing Science(计算科学研究所) ; Poznan University of Technology(波兹南技术大学)
专题命中 效率与部署 :LLM(title,title_cn);post-training(abstract);分类 cs.CL、cs.LG
AI总结 提出QAM-W方法,通过L2归一化、块哈达玛旋转和2D坐标配对量化,结合激活感知缩放,在约5.5 bpw下使困惑度接近BF16,优于极坐标编码,并在5-6 bpw范围内保持质量。