UltraSketchLLM: Sub-1-Bit LLM Compression via Sketch and Hardware-Friendly Operators
UltraSketchLLM:基于草图与硬件友好算子的低于1比特LLM压缩
Sunan Zou, Xueting Sun, Ziyun Zhang, Guojie Luo
机构
*
National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(国家多媒体信息处理重点实验室,计算机科学学院,北京大学)
;
School of Electronic Engineering and Computer Science, Peking University(电子工程与计算机科学学院,北京大学)
;
Center for Energy-efficient Computing and Applications, Peking University(能效计算与应用中心,北京大学)
专题命中
效率与部署
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
FP4量化LLM训练中均值偏差的诅咒与祝福
Hengjie Cao, Zhendong Huang, Mengyi Chen, Yifeng Yang, Fang Dong, Anrui Chen, Ruijun Huang, Xin Zhang, Mingzhi Dong, Yujiang Wang, Jinlong Hou, Qin Lv, Robert P. Dick, Yuan Cheng, Tun Lu, Fan Yang, Yixuan Chen, Li Shang
机构
*
Fudan University(复旦大学)
;
University of Bath(巴斯大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
University of Oxford(牛津大学)
;
Oxford Suzhou Centre for Advanced Research(牛津苏浙研究中心)
;
University of Colorado Boulder(科罗拉多大学波德格分校)
;
University of Michigan(密歇根大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
专题命中
效率与部署
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
CommentsThis is the author's accepted version of the paper accepted to appear at IEEE AIIoT 2025. The final version will be available via IEEE Xplore. \c{opyright}2025 IEEE. Personal use of this material is permitted
Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression
Sentinel: 通过注意力探测解码上下文利用以实现高效LLM上下文压缩
Yong Zhang, Heng Li, Yanwen Huang, Ning Cheng, Yang Guo, Yun Zhu, Yanmeng Wang, Shaojun Wang, Jing Xiao
机构
*
Ping An Technology (Shenzhen) Co., Ltd., China(平安科技(深圳)有限公司,中国)
;
University of Science and Technology of China(中国科学技术大学)
;
University of Electronic Science and Technology of China(电子科技大学)
CommentsAccepted at ICML 2026 (43rd International Conference on Machine Learning, Seoul, South Korea). Code available at https://github.com/yyyyhx/MINIM
MirrorCheck: Efficient Adversarial Defense for Vision-Language Models
MirrorCheck: 视觉-语言模型的高效对抗防御
Samar Fares, Klea Ziu, Toluwani Aremu, Nikita Durasov, Martin Takáč, Pascal Fua, Ivan Laptev, Karthik Nandakumar
机构
*
Mohamed Bin Zayed University of Artificial Intelligence(莫扎伊德大学人工智能大学)
;
NVIDIA
;
École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)
;
Michigan State University(密歇根州立大学)
I'm Sorry Driver, I'm Afraid I Can't Do That: Appraising the Safety of LLMs within Automotive Contexts
抱歉,司机,恐怕我不能这么做:评估LLMs在汽车环境中的安全性
Shaun Feakins, Ibrahim Habli, Kim Littler, Robert Palin
机构
*
UKRI AI Centre for Doctoral Training in Safe Artificial Intelligence Systems (SAINTS)(英国研究理事会安全人工智能系统博士培训中心(SAINTS))
;
University of York(约克大学)
;
Jaguar Land Rover(捷克·陆罗恩)
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
Carnegie Mellon University Africa Kigali Rwanda(卡内基梅隆大学非洲分校,基亚利,卢旺达)
;
Rochester Institute of Technology(罗切斯特理工学院)
;
Rochester New York USA(罗切斯特,纽约州,美国)
;
Carnegie Mellon University Pittsburgh Pennsylvania USA(卡内基梅隆大学匹兹堡,宾夕法尼亚州,美国)
;
University of California, Riverside(加州大学河滨分校)
专题命中
效率与部署
:large language model(abstract);language model(abstract);分类 cs.LG