Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding
学习承诺:通过学习异步解码扩展语言模型解码并行性
Tian Jin, Ellie Y. Cheng, Zack Ankner, Nikunj Saunshi, Blake M. Elias, Amir Yazdanbakhsh, Jonathan Ragan-Kelley, Suvinay Subramanian, Michael Carbin
机构
*
DeepMind, London, UK(深度思维公司,伦敦,英国)
;
Google Research, New York, NY, USA(谷歌研究院,纽约,纽约州,美国)
;
Stanford University, Stanford, CA, USA(斯坦福大学,斯坦福,加利福尼亚州,美国)
;
University of Toronto, Toronto, Ontario, Canada(多伦多大学,多伦多,安大略省,加拿大)
;
University of Washington, Seattle, WA, USA(华盛顿大学,西雅图,华盛顿州,美国)
机构
*
TMLR Group, Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学系 TMLR 组)
;
Stanford University(斯坦福大学)
;
Sydney AI Centre, The University of Sydney(悉尼大学人工智能中心)
机构
*
Stanford University(斯坦福大学)
;
Wuhan University of Science and Technology(武汉科技大学)
;
Yale University(耶鲁大学)
;
School of Medicine, Yale University(耶鲁大学医学院)
;
Hunan University(湖南大学)
;
Yuelushan Laboratory(岳麓实验室)
;
Kumo.AI
;
Toyota Technological Institute at Chicago(芝加哥技术研究所)
机构
*
Boston University School of Engineering(波士顿大学工程学院)
;
Stanford University(斯坦福大学)
;
University of Pittsburgh Medical Center(匹兹堡大学医学中心)
;
Boston University School of Medicine(波士顿大学医学院)