Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production
将文档AI operationalize:一种用于OCR和LLM流水线的微服务架构
Yao Fehlis, Benjamin Bengfort, Zhangzhang Si, Vahid Eyorokon, Prema Roman, Patrick Deziel, Devon Slonaker, Steve Veldman, Ben Johnson, Joyce Rigelo, Michael Wharton, Steve Kramer
专题命中
效率与部署
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
The Alpha Illusion: Reported Alpha from LLM Trading Agents Should Not Be Treated as Deployment Evidence
阿尔法幻觉:LLM交易代理报告的阿尔法不应被视为部署证据
Yuxuan Ye, Jun Han, Ao Hu, Juncheng Bu, Yiyi Chen, Liangjian Wen, Danilo Mandic, Danny Dongning Sun, Xu Yinghui, Zenglin Xu
机构
*
Fudan University(复旦大学)
;
Shanghai University of Finance and Economics(上海财经大学)
;
Southwest University of Finance and Economics(西南财经大学)
;
Northeastern University(东北大学)
;
Imperial College London(伦敦帝国理工学院)
;
Peng Cheng Laboratory(鹏城实验室)
MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings
MTRouter: 带成本意识的多轮LLM路由与历史-模型联合嵌入
Yiqun Zhang, Hao Li, Zihan Wang, Shi Feng, Xiaocui Yang, Daling Wang, Bo Zhang, Lei Bai, Shuyue Hu
机构
*
School of Computer Science and Engineering, Northeastern University Shenyang 110819, China(东北大学计算机科学与工程学院,中国沈阳110819)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中
效率与部署
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
AI总结
本文提出MTRouter,通过联合历史-模型嵌入和学习轨迹预测器,提升多轮任务的性能-成本平衡,实验显示在ScienceWorld和Humanity's Last Exam上均取得显著成本降低与性能提升。
Clinical Note Bloat Reduction for Efficient LLM Use
临床笔记去冗余以提高大语言模型使用效率
Jordan L. Cahoon, Chloe Stanwyck, Asad Aali, Rachel Madding, Emma Sun, Yixing Jiang, Renumathy Dhanasekaran, Emily Alsentzer
机构
*
Department of Biomedical Data Science, Stanford University, Stanford, CA(斯坦福大学生物医学数据科学系)
;
Department of Pathology, Stanford University, Stanford, CA(斯坦福大学病理学系)
;
Department of Anesthesiology, Perioperative and Pain Medicine, Stanford University, Stanford, CA(斯坦福大学麻醉学、围术期医学与疼痛医学系)
;
Department of Radiology, Stanford University, Stanford, CA(斯坦福大学放射学系)
;
Department of Obstetrics and Gynecology, Stanford University, Stanford, CA(斯坦福大学妇产科学系)
;
Division of Gastroenterology and Hepatology, Stanford University, Stanford, CA(斯坦福大学消化内科与肝病学系)
;
Department of Computer Science, Stanford University, Stanford, CA(斯坦福大学计算机科学系)
;
Weill Cancer Hub West(韦尔癌症中心西区)
专题命中
效率与部署
:LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Comments10 pages, 3 figures, 1 table. Empirical measurement study reporting new repeated-run experiments quantifying baseline nondeterministic drift in large language models. This manuscript presents original empirical results (not a review or position paper) and establishes a baseline reference for future drift-mitigation work
Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
从何处开始:通过子网络选择和蒸馏实现高效的预训练
Arjun Krishnakumar, Rhea Sanjay Sukthanker, Hannan Javed Mahadik, Gabriela Kadlecová, Vladyslav Moroshan, Timur Carstensen, Frank Hutter, Aaron Klein
机构
*
University of Freiburg, Germany(弗赖堡大学)
;
ELLIS Institute Tübingen, Germany(图宾根ELLIS研究所)
;
Charles University, Faculty of Mathematics and Physics(查尔斯大学数学与物理系)
;
PriorLabs
;
The Czech Academy of Sciences, Institute of Computer Science(捷克科学院计算机科学研究所)
专题命中
效率与部署
:pretraining(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)
Consolidating TinyML Lifecycle with Large Language Models: Reality, Illusion, or Opportunity?
Guanghan Wu, Sasu Tarkoma, Roberto Morabito
专题命中
效率与部署
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI、cs.LG
CommentsThis paper has been accepted for publication in the IEEE Internet of Things Magazine (Special Issue on Applications of Large Language Models in IoT). The copyright will be transferred to IEEE upon publication. A preliminary version of this work was presented at the Edge AI Foundation event Beyond LLMs and Chatbots: The Journey to Generative AI at the Edge (https://youtu.be/aFWfisdjQIs)
Prefill/Decode-Aware Evaluation of LLM Inference on Emerging AI Accelerators
新兴AI加速器上LLM推理的Prefill/Decode感知评估
Shun Usami, Venkatram Vishwanath, E. Wes Bethel
机构
*
Department of Computer Science(计算机科学系)
;
San Francisco State University(旧金山州立大学)
;
Argonne National Laboratory(阿贡国家实验室)
;
Lawrence Berkeley National Laboratory(伯克利国家实验室)
专题命中
效率与部署
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI