arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 22188 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 22188 篇

2605.06669 2026-05-22 cs.CR cs.AI cs.LG 90%

Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs

评估教育LLM导师的提示注入防御:安全-可用性-延迟的权衡

Alexandre Cristovão Maiorano

机构 * Lumytics

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI、cs.LG

AI总结 本文提出了一种评估提示注入防御方法的框架,探讨了在教育LLM导师中安全、可用性和延迟之间的权衡,并通过实验比较了不同防御机制的性能。

Comments 19 pages, 4 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12610 2026-05-22 cs.DB cs.AI cs.LG 90%

ScaleDoc: Scaling LLM-based Predicates over Large Document Collections

ScaleDoc: 通过大规模文档集合进行基于大语言模型的谓词扩展

Hengrui Zhang, Yulong Hui, Yihao Liu, Huanchen Zhang

机构 * Tsinghua University(清华大学)

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出ScaleDoc系统,通过将谓词执行分为离线表示阶段和优化的在线过滤阶段,解决了大规模文档分析中大语言模型高推理成本的问题,实现了端到端速度提升和LLM调用成本降低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18818 2026-05-20 cs.AI cs.LG cs.SE 90%

Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production

将文档AI operationalize:一种用于OCR和LLM流水线的微服务架构

Yao Fehlis, Benjamin Bengfort, Zhangzhang Si, Vahid Eyorokon, Prema Roman, Patrick Deziel, Devon Slonaker, Steve Veldman, Ben Johnson, Joyce Rigelo, Michael Wharton, Steve Kramer

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种微服务架构,用于在生产环境中实现文档理解,通过整合多个模型的流水线,包括分类、OCR和LLM结构字段提取,并展示了在每小时处理数千页文档的经验。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16895 2026-05-19 cs.CE cs.AI cs.CL 90%

The Alpha Illusion: Reported Alpha from LLM Trading Agents Should Not Be Treated as Deployment Evidence

阿尔法幻觉:LLM交易代理报告的阿尔法不应被视为部署证据

Yuxuan Ye, Jun Han, Ao Hu, Juncheng Bu, Yiyi Chen, Liangjian Wen, Danilo Mandic, Danny Dongning Sun, Xu Yinghui, Zenglin Xu

机构 * Fudan University(复旦大学) Shanghai University of Finance and Economics(上海财经大学) Southwest University of Finance and Economics(西南财经大学) Northeastern University(东北大学) Imperial College London(伦敦帝国理工学院) Peng Cheng Laboratory(鹏城实验室)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.CL、cs.AI

AI总结 本文指出,LLM交易代理报告的阿尔法不应被当作部署的证据,因为这些阿尔法需要通过结构有效性测试来验证其时间完整性、现实摩擦、反事实稳健性、预测校准、数值执行和多代理分解等关键指标,当前的公开证据无法区分稳健的预测能力与时间污染、未建模的摩擦、短窗口夏普不确定性、叙事拟合和参数先验等因素。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07799 2026-05-19 cs.CL cs.AI 90%

Dynamic Generation of Multi-LLM Agents Communication Topologies with Graph Diffusion Models

基于图扩散模型的多LLM代理通信拓扑动态生成

Eric Hanchen Jiang, Mengting Li, Guancheng Wan, Sophia Yin, Yuchen Wu, Xiao Liang, Xinfeng Li, Yizhou Sun, Wei Wang, Kai-Wei Chang, Ying Nian Wu

机构 * University of California Los Angeles(加州大学洛杉矶分校) University of Washington(华盛顿大学) Nanyang Technological University(南洋理工大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出Guided Topology Diffusion框架,通过迭代构建过程生成适应任务需求的高效通信拓扑,优于现有方法。

Comments ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14844 2026-05-15 cs.LG cs.AI 90%

XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference

XFP:面向LLM推理的质量导向自适应码本量化与稀疏异常分离

Thomas Witt

机构 * Gemini Stiftung(吉姆米基金会)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI、cs.LG

AI总结 XFP通过自适应码本量化和稀疏异常分离技术,实现高质量LLM推理,无需Hessian或校准数据,支持动态调整码本大小和异常预算,提升推理速度和精度。

Comments 17 pages, 3 figures, 17 tables, 1 algorithm. Code: https://github.com/flash7777/vllm/tree/multiquant

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12019 2026-05-13 cs.LG cs.AI 90%

Efficient and Adaptive Human Activity Recognition via LLM Backbones

通过LLM骨干实现高效且自适应的人体活动识别

Aleksandr Bredikhin, Philippe Lalanda, German Vega

机构 * Univ. Grenoble Alpes, France(格勒诺布尔阿尔卑斯大学,法国)

专题命中 效率与部署 :LLM(title,title_cn);language model(abstract);foundation model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出利用预训练大语言模型作为通用时间骨干,用于基于传感器的人体活动识别,通过结构化卷积投影和参数高效LoRA适应,提升跨数据集迁移能力和低数据场景下的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11999 2026-05-13 cs.DC cs.AI cs.LG cs.PF 90%

The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures

LLM解码中的功率限制幻觉:跨注意力架构的相位感知能耗分析

Bole Ma, Ayesha Afzal, Jan Eitzinger, Gerhard Wellein

机构 * Erlangen National High Performance Computing Center(埃朗根国家高性能计算中心) Friedrich-Alexander-Universität Erlangen-Nürnberg(埃朗根-纽伦堡弗里德里希-亚历山大大学)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI、cs.LG

AI总结 研究揭示LLM解码中功率限制的表面现象,通过分析四种注意力架构发现解码实际耗能较低,内存瓶颈导致带宽饱和而非计算瓶颈,钟速锁定比功率限制更有效,可提升解码能耗32%并最小化吞吐量损失。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11334 2026-05-13 cs.LG cs.CL cs.IR 90%

VERDI: Single-Call Confidence Estimation for Verification-Based LLM Judges via Decomposed Inference

VERDI:基于验证的LLM裁判的单次调用置信度估计方法

Jasmine Qi, Danylo Dantsev, Muyang Sun

机构 * Indeed Inc(Indeed公司)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.CL、cs.LG

AI总结 VERDI通过分解推理过程提取置信度信号,无需额外调用,提升验证型LLM裁判的可信度评估。

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11202 2026-05-13 cs.CR cs.AI cs.LG cs.SE 90%

Continuous Discovery of Vulnerabilities in LLM Serving Systems with Fuzzing

在LLM服务系统中通过模糊测试连续发现漏洞

Yunze Zhao, Yibo Zhao, Yuchen Zhang, Zaoxing Liu, Michelle L. Mazurek

机构 * University of Maryland(马里兰大学) New York University(纽约大学)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI、cs.LG

AI总结 本文提出GRIEF模糊器,通过时间多请求追踪发现LLM服务层漏洞,包括15个漏洞,10个已确认,涉及缓存隔离失败、跨请求性能干扰和崩溃问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11019 2026-05-13 cs.LG cs.AI 90%

Efficient LLM Reasoning via Variational Posterior Guidance with Efficiency Awareness

通过变分后验指导实现高效的LLM推理:具有效率意识

Zizhao Chen, Yuying Li, Siting Lin, Lianxi Wang

机构 * Guangdong University of Foreign Studies(广东外语外贸大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出VPG-EA框架,通过变分推断和效率感知证据下界提升LLM推理效率,实验显示在不同模型规模下效率指标提升显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04284 2026-05-12 cs.AI cs.LG 90%

Agent-Omit: Adaptive Context Omission for Efficient LLM Agents

Agent-Omit:适应性上下文省略以提高LLM代理效率

Yansong Ning, Jun Fang, Naiqiang Tan, Hao Liu

机构 * AI Thrust, The Hong Kong University of Science(香港科学与技术大学人工智能前沿) Didichuxing Co. Ltd(滴滴出行有限公司)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI、cs.LG

AI总结 本文提出Agent-Omit框架,通过省略冗余思考和观察提升LLM代理效率,实验表明其在多个基准测试中表现优异。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06320 2026-05-08 cs.MA cs.AI cs.CL 90%

Improving the Efficiency of Language Agent Teams with Adaptive Task Graphs

通过自适应任务图提高语言代理团队的效率

Elizabeth Mieczkowski, Alexander Ku, Tiwalayo Eisape, Dilip Arumugam, John Matters, Katherine M. Collins, Ilia Sucholutsky, Thomas L. Griffiths

机构 * Princeton University(普林斯顿大学) University of Cambridge(剑桥大学) MIT(麻省理工学院) New York University(纽约大学)

专题命中 效率与部署 :language agent(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文提出LATTE框架,通过自适应任务图提升语言代理团队效率,减少资源消耗并提高协作准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05696 2026-05-08 cs.DC cs.AI cs.LG 90%

Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving

Irminsul:基于MLA的原生位置无关缓存用于代理LLM服务

Bole Ma, Jan Eitzinger, Harald Köstler

机构 * Erlangen National High Performance Computing Center(埃朗根国家高性能计算中心)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI、cs.LG

AI总结 本文提出Irminsul,一种基于MLA的原生位置无关缓存,通过内容哈希键和δ旋转规则提升代理LLM服务性能,实现高达83%的提示词恢复率和63%的预填能量节省。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05485 2026-05-08 cs.CL cs.AI 90%

ReaComp: Compiling LLM Reasoning into Symbolic Solvers for Efficient Program Synthesis

ReaComp:将LLM推理编译为符号求解器以实现高效的程序合成

Atharva Naik, Yash Mathur, Prakam, Carolyn Rose, David Mortensen

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.CL、cs.AI

AI总结 通过编译推理轨迹生成符号求解器,提升程序合成效率与准确性,同时减少对LLM的依赖,适用于多个基准测试任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.03379 2026-05-08 cs.LG cs.CL 90%

Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference

两次调用、两次矩与重复LLM推理的投票准确性曲线

Yi Liu

机构 * York University(约克大学)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.CL、cs.LG

AI总结 研究重复LLM推理的二元正确性层,在条件独立同分布调用下,通过两次调用确定第二矩和相同示例正确性相关性,从而为固定多数投票预算提供分布无关的两调用区间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15954 2026-04-29 cs.LG cs.AI 90%

MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment

MobileLLM-Flash: 用于工业级部署的延迟引导设备端大语言模型设计

Hanxian Huang, Igor Fedorov, Andrey Gromov, Bernard Beckerman, Naveen Suda, David Eriksson, Maximilian Balandat, Rylan Conway, Patrick Huber, Chinnadhurai Sankar, Ayushi Dalmia, Zechun Liu, Lemeng Wu, Tarek Elgamal, Adithya Sagar, Vikas Chandra, Raghuraman Krishnamoorthi

机构 * Meta AI

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);foundation model(abstract)

AI总结 本文提出MobileLLM-Flash,通过硬件循环架构搜索在移动端延迟约束下设计高效的大语言模型,支持8k上下文长度,实现更快的预填和解码速度。

Comments Accepted to ACL Industry Track 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23530 2026-04-28 cs.CL cs.AI 90%

MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings

MTRouter: 带成本意识的多轮LLM路由与历史-模型联合嵌入

Yiqun Zhang, Hao Li, Zihan Wang, Shi Feng, Xiaocui Yang, Daling Wang, Bo Zhang, Lei Bai, Shuyue Hu

机构 * School of Computer Science and Engineering, Northeastern University Shenyang 110819, China(东北大学计算机科学与工程学院,中国沈阳110819) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MTRouter,通过联合历史-模型嵌入和学习轨迹预测器,提升多轮任务的性能-成本平衡,实验显示在ScienceWorld和Humanity's Last Exam上均取得显著成本降低与性能提升。

Comments This work has accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18697 2026-04-22 cs.CR cs.CL cs.LG 90%

Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs

超越不可区分性:衡量LLM API中的提取风险

Ruixuan Liu, David Evans, Li Xiong

机构 * Emory University(埃默里大学) University of Virginia(弗吉尼亚大学)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.CL、cs.LG

AI总结 本文提出$(l, b)$-不可提取性作为衡量LLM API提取风险的新标准,通过定义提取风险上界并实验证明其在不同模型中的有效性,为模型训练、API访问和解码配置提供缓解指南。

Comments Accepted by S&P 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09775 2026-04-22 cs.AR cs.AI cs.DC cs.LG 90%

MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference

MIST:一种用于异构、多阶段LLM推理的联合设计框架

Abhimanyu Rajeshkumar Bambhaniya, Hanjiang Wu, Suvinay Subramanian, Sudarshan Srinivasan, Souvik Kundu, Amir Yazdanbakhsh, Midhilesh Elavazhagan, Madhu Kumar, Minlan Yu, Arijit Raychowdhury, Tushar Krishna

机构 * Georgia Institute of Technology(佐治亚理工学院) Google(谷歌) Intel(英特尔) Intel Labs(英特尔实验室) Google DeepMind(谷歌DeepMind) Harvard University(哈佛大学) Infravana

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI、cs.LG

AI总结 MIST是一种用于异构、多阶段LLM推理的联合设计框架,通过模拟不同请求阶段和复杂硬件层次,优化硬件-软件协同设计,解决LLM推理中的配置空间导航和跨厂商PD配置问题。

Comments Inference System Design for Multi-Stage AI Inference Pipelines. 11 Pages, 10 Figues, 5 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07954 2026-04-21 cs.CL cs.AI 90%

Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation

Bielik Guard:高效的波兰语安全分类器用于LLM内容审核

Krzysztof Wróbel, Jan Maria Kowalski, Jerzy Surma, Igor Ciuciura, Maciej Szymański

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 Bielik Guard通过高效波兰语安全分类器提升LLM内容审核效果,0.5B模型在F1分数上表现最佳,0.1B模型在效率与低误报率上优于现有方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02718 2026-04-21 cs.LG cs.AI 90%

End-to-End Optimization of LLM-Driven Multi-Agent Search Systems via Heterogeneous-Group-Based Reinforcement Learning

通过异质组强化学习实现LLM驱动的多智能体搜索系统的端到端优化

Guanzhong Chen, Shaoxiong Yang, Chao Li, Wei Liu, Jian Luan, Zenglin Xu

机构 * MiLM Plus, Xiaomi Inc.(小米公司MiLM Plus实验室) Fudan University(复旦大学) Shanghai Academy of AI for Science(上海人工智能科学研究院)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出MHGPO方法,通过估计异质组间的相对优势,优化多智能体系统的整体成功而非单个智能体性能,提升了任务表现和计算效率。

Comments Accepted to ACL 2026 Main Conference. 20 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16364 2026-04-21 cs.CY cs.AI cs.CL 90%

Clinical Note Bloat Reduction for Efficient LLM Use

临床笔记去冗余以提高大语言模型使用效率

Jordan L. Cahoon, Chloe Stanwyck, Asad Aali, Rachel Madding, Emma Sun, Yixing Jiang, Renumathy Dhanasekaran, Emily Alsentzer

机构 * Department of Biomedical Data Science, Stanford University, Stanford, CA(斯坦福大学生物医学数据科学系) Department of Pathology, Stanford University, Stanford, CA(斯坦福大学病理学系) Department of Anesthesiology, Perioperative and Pain Medicine, Stanford University, Stanford, CA(斯坦福大学麻醉学、围术期医学与疼痛医学系) Department of Radiology, Stanford University, Stanford, CA(斯坦福大学放射学系) Department of Obstetrics and Gynecology, Stanford University, Stanford, CA(斯坦福大学妇产科学系) Division of Gastroenterology and Hepatology, Stanford University, Stanford, CA(斯坦福大学消化内科与肝病学系) Department of Computer Science, Stanford University, Stanford, CA(斯坦福大学计算机科学系) Weill Cancer Hub West(韦尔癌症中心西区)

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出TRACE方法,通过EHR元数据去除临床笔记中的冗余文本,降低LLM计算成本,实验证明在保持信息提取和预测性能的同时,减少47.3%的文本量,节省每年约950万美元的推理成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00136 2026-04-15 cs.LG cs.CL 90%

ParetoBandit: Budget-Paced Adaptive Routing for Non-Stationary LLM Serving

ParetoBandit: 面向非平稳LLM服务的预算驱动自适应路由

Annette Taberner-Miller

机构 * Independent Researcher(独立研究者)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.CL、cs.LG

AI总结 ParetoBandit提出一种基于成本感知上下文老虎机的自适应路由方法,解决非平稳环境下LLM服务的预算控制与在线适应问题,实现高质量与低成本的平衡。

Comments 27 pages, 15 figures, 13 tables. Code available at https://github.com/ParetoBandit/ParetoBandit

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19934 2026-01-29 cs.CL cs.AI 90%

Quantifying non deterministic drift in large language models

量化大型语言模型中的非确定性漂移

Claire Nicholson

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过实验量化大型语言模型在无操作员条件下的非确定性漂移,揭示了不同模型大小和部署类型下的输出变化模式,并探讨了语义方法在评估漂移缓解技术中的应用。

Comments 10 pages, 3 figures, 1 table. Empirical measurement study reporting new repeated-run experiments quantifying baseline nondeterministic drift in large language models. This manuscript presents original empirical results (not a review or position paper) and establishes a baseline reference for future drift-mitigation work

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07227 2026-01-15 cs.CL cs.AI 90%

Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation

从何处开始:通过子网络选择和蒸馏实现高效的预训练

Arjun Krishnakumar, Rhea Sanjay Sukthanker, Hannan Javed Mahadik, Gabriela Kadlecová, Vladyslav Moroshan, Timur Carstensen, Frank Hutter, Aaron Klein

机构 * University of Freiburg, Germany(弗赖堡大学) ELLIS Institute Tübingen, Germany(图宾根ELLIS研究所) Charles University, Faculty of Mathematics and Physics(查尔斯大学数学与物理系) PriorLabs The Czech Academy of Sciences, Institute of Computer Science(捷克科学院计算机科学研究所)

专题命中 效率与部署 :pretraining(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文提出一种高效预训练框架,通过子网络选择和蒸馏技术,使小型语言模型在资源消耗上显著低于大型模型,同时保持相近的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17092 2025-08-01 cs.LG cs.AI cs.DM 90%

Accumulator-Aware Post-Training Quantization for Large Language Models

Ian Colbert, Giuseppe Franco, Fabian Grob, Jinjie Zhang, Rayan Saab

机构 * AMD(AMD公司) TUM(慕尼黑工业大学) GSK(葛兰素史克) University of California San Diego(加州大学圣地亚哥分校)

专题命中 效率与部署 :post-training(title,abstract);large language model(title);language model(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12420 2025-04-08 cs.SE cs.AI cs.LG 90%

Consolidating TinyML Lifecycle with Large Language Models: Reality, Illusion, or Opportunity?

Guanghan Wu, Sasu Tarkoma, Roberto Morabito

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI、cs.LG

Comments This paper has been accepted for publication in the IEEE Internet of Things Magazine (Special Issue on Applications of Large Language Models in IoT). The copyright will be transferred to IEEE upon publication. A preliminary version of this work was presented at the Edge AI Foundation event Beyond LLMs and Chatbots: The Journey to Generative AI at the Edge (https://youtu.be/aFWfisdjQIs)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07419 2026-08-10 cs.LG 新提交 90%

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration

超越事后温度缩放:用于大语言模型校准的双层优化方法

Ruochen Jin, Zhanliang Wang, Zongyu Dai, Jiancong Xiao, Bojian Hou

机构 * Dartmouth College(达特茅斯学院) University of Pennsylvania(宾夕法尼亚大学) National University of Singapore(新加坡国立大学)

专题命中 效率与部署 :LLM(title,summary_cn);language model(abstract,comments);large language model(abstract);分类 cs.LG

AI总结 针对LLM偏好对齐导致的过度自信与校准问题,提出基于双层优化的校准方法,通过最大化预测分布熵实现,在多项选择与开放式问答任务中提升了校准效果与域外泛化能力。

Comments Third Conference on Language Modeling (COLM 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17104 2026-06-17 cs.AR cs.AI cs.DC 新提交 90%

Prefill/Decode-Aware Evaluation of LLM Inference on Emerging AI Accelerators

新兴AI加速器上LLM推理的Prefill/Decode感知评估

Shun Usami, Venkatram Vishwanath, E. Wes Bethel

机构 * Department of Computer Science(计算机科学系) San Francisco State University(旧金山州立大学) Argonne National Laboratory(阿贡国家实验室) Lawrence Berkeley National Laboratory(伯克利国家实验室)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文通过分离测量Prefill和Decode阶段,评估GPU与新兴AI加速器在Llama2-7B模型上的推理性能,发现GPU在计算密集的Prefill阶段占优,而GroqRack在Decode延迟上更优,但GPU随批处理增大在吞吐上反超。

Comments 8 pages, 5 figures. Accepted to the Workshop on HPC for AI Foundation Models & LLMs for Science (HPAI4S'26), co-located with IEEE IPDPS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏