WINDQuant: Weight-Informed Neural Decision-Making for Global Mixed-Precision LLM Quantization
WINDQuant: 权重感知的全局混合精度大语言模型量化神经决策
Phong Nam Huu Nguyen, Khoi M. Le, Cong-Duy T Nguyen, Anh Tuan Luu, Thong Thanh Nguyen, Tho Quan
机构
*
CAIR, VinUniversity, Vietnam(越南 VinUniversity 的 CAIR)
;
Ho Chi Minh City University of Technology (HCMUT), VNU-HCM, Ho Chi Minh City, Vietnam(越南胡志明市技术大学 (HCMUT))
;
CCDS, Nanyang Technological University, Singapore(新加坡南洋理工大学的 CCDS)
;
School of Computing, National University of Singapore, Singapore(新加坡国立大学计算机学院)
专题命中
效率与部署
:LLM(title,summary_cn);large language model(abstract);language model(abstract);post-training(abstract)
Communication-Efficient Hybrid Language Model via Uncertainty-Aware Opportunistic and Compressed Transmission
基于不确定性感知的机会主义与压缩传输的通信高效混合语言模型
Seungeun Oh, Jinhyuk Kim, Jihong Park, Seung-Woo Ko, Jinho Choi, Tony Q. S. Quek, Seong-Lyun Kim
机构
*
School of Electrical and Electronic Engineering, Yonsei University, South Korea(延世大学电气电子工程学院,韩国)
;
Information Systems Technology and Design pillar, Singapore University of Technology and Design, Singapore 487372(新加坡科技设计大学信息系统技术与设计支柱,新加坡487372)
;
Department of Smart Mobility Engineering, Inha University, South Korea(Inha大学智能移动工程系,韩国)
;
School of Electrical and Mechanical Engineering, The University of Adelaide, Australia(阿德莱德大学电气与机械工程学院,澳大利亚)
;
Singapore University of Technology and Design, Singapore 487372(新加坡科技设计大学,新加坡487372)
专题命中
效率与部署
:language model(title,abstract);LLM(abstract,abstract_cn);SLM(abstract,abstract_cn);large language model(abstract)
Comments17 pages, 13 figures, 5 tables; This article has been accepted for publication in IEEE Transactions on Communications. This is the author's accepted version; the final published version will be available via IEEE Xplore
AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile Devices
AHASD:异步异构架构用于移动设备上的LLM自适应草案推测解码
Ma Zirui, Fan Zhihua, Li Wenxing, Wu Haibin, Zhang Fulin, Ye Xiaochun, Li Wenming
机构
*
State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所处理器国家重点实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
专题命中
效率与部署
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI
Comments26 pages, 12 figures. Empirical benchmark comparing fine-tuned encoders and LLM prompting for text classification under cost and latency constraints