NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
NanoQuant: 大型语言模型高效子1位量化
机构 * KAIST(韩国科学技术院)
AI总结 本文提出NanoQuant,一种新型后训练量化方法,能够将大型语言模型压缩到二进制和子1位水平,通过低秩二进制分解问题实现高效压缩,并在消费者硬件上实现大规模部署。
Comments Accepted to ICML 2026. Hyochan Chong and Dongkyu Kim contributed equally to this work