Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention
全栈FP4:通过量化投影、优化器和注意力进行稳定的语言模型预训练
机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; Sichuan University(四川大学) ; Beijing Zhongguancun Academy(北京中关村科学院) ; Zhejiang University(浙江大学) ; Beijing Key Laboratory of Brain-Inspired General Intelligence Large Model(北京脑科学与类脑研究中心通用人工智能大模型北京市重点实验室) ; Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与脑机智能技术重点实验室)
专题命中 预训练与数据 :pretraining(title,abstract);LLM(title);分类 cs.AI、cs.LG
AI总结 研究针对4位预训练中优化器状态等未充分探索的问题,提出全栈FP4框架。通过模块精度策略解决稳定性瓶颈,如用LoRA-SVD抑制量化噪声,设计优化器变换,采用混合精度方案,实现稳定高效的端到端预训练。
Comments Fix experiment bugs and some statement, update current developed hardware efficiency result on 5090