arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 22188 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 22188 篇

2409.00084 2026-04-10 cs.CL cs.AI 90%

Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models

视觉-语言与大语言模型在胃肠病学中的表现:GPT、Claude、Llama、Phi、Mistral、Gemma及量化模型

Seyed Amir Ahmad Safavi-Naini, Shuhaib Ali, Omer Shahab, Zahra Shahhoseini, Thomas Savage, Sara Rafiee, Jamil S Samaan, Reem Al Shabeeb, Farah Ladak, Jamie O Yang, Juan Echavarria, Sumbal Babar, Aasma Shaukat, Samuel Margolis, Nicholas P Tatonetti, Girish Nadkarni, Bara El Kurdi, Ali Soroush

机构 * Icahn School of Medicine at Mount Sinai(西奈山伊坎医学院) University of Texas Health(德克萨斯大学健康科学中心) Virginia Hospital Center(弗吉尼亚医院中心) Shahid Beheshti University of Medical Sciences(沙希德·贝赫什提医科大学) Stanford University(斯坦福大学) Cedars-Sinai Medical Center(西达赛奈医疗中心) Inova Fairfax Medical Campus(伊诺瓦费尔法克斯医疗中心) University of California–Los Angeles(加州大学洛杉矶分校) NYU Grossman School of Medicine(纽约大学格罗斯曼医学院) Columbia University(哥伦比亚大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,comments);分类 cs.CL、cs.AI

AI总结 本研究评估了大语言模型和视觉-语言模型在胃肠病学中的医学推理性能,比较了不同模型配置、参数及提示工程策略对性能的影响,发现专有模型在准确性上优于开源模型,且图像描述对视觉-语言模型性能有显著影响。

Comments Manuscript Pages: 34, Figures: 7, Tables: 2, Supplementary File Pages: 35, Data Transparency Statement: Code is available at: https://github.com/Sdamirsa/LLM-VLM-in-Gastroenterology . Study data from American College of Gastroenterology (ACG) are restricted and available upon request with ACG permission. Correction: updated abstract considering Llama3.1 results

Journal ref npj Digital Medicine 8, 797 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14017 2025-10-01 cs.CL cs.LG 90%

Efficient Temporal Tokenization for Mobility Prediction with Large Language Models

Haoyu He, Haozheng Luo, Yan Chen, Qi R. Wang

机构 * Northeastern University, Boston, MA(东北大学) Northwestern University, Evanston, IL(西北大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL、cs.LG

Journal ref Proceedings of the 3rd Workshop on Efficient Systems for Foundation Models (ES-FoMo III) at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18002 2025-03-26 cs.NE cs.AI cs.AR cs.LG 90%

Neuromorphic Principles for Efficient Large Language Models on Intel Loihi 2

Steven Abreu, Sumit Bam Shrestha, Rui-Jie Zhu, Jason Eshraghian

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI、cs.LG

Comments Accepted to International Conference on Learning Representations (ICLR) Workshop on Scalable Optimization for Efficient and Adaptive Foundation Models (SCOPE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03034 2025-02-06 cs.CL cs.LG 90%

Knowledge Distillation from Large Language Models for Household Energy Modeling

Mohannad Takrouri, Nicolás M. Cuadrado, Martin Takáč

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,comments);分类 cs.CL、cs.LG

Comments Source code is available at https://github.com/Singularity-AI-Lab/LLM-Energy-Knowledge-Distillation

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14677 2023-11-28 cs.CY cs.CL cs.LG 90%

Filter bubbles and affective polarization in user-personalized large language model outputs

Tomo Lazovich

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL、cs.LG

Comments Accepted to NeurIPS 2023 Workshop "I Can't Believe It's Not Better: Failure Modes in the Age of Foundation Models"

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15127 2026-08-18 cs.OS cs.AI cs.DC cs.MA 新提交 90%

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

从大语言模型(LLM)推理到智能体工作负载:特征分析与服务系统启示

Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI

AI总结 本文提出AgentSysBench基准套件,分析智能体工作负载与传统LLM服务的6项差异,据此开展的4项设计探索可实现延迟降低、效率提升、内存减少及冗余调用节省等效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14065 2026-08-18 cs.SE cs.AI 版本更新 90%

Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-efficiency

重新思考自动程序修复:缺陷复杂度、缺陷定位与大语言模型成本效率的影响

Junchi Liu, Ali Bigdeli, Roya Daneshi, Atu Ambala, Sudipto Ghosh, Fabio Santos

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究通过实证分析,发现中等复杂度缺陷可超50%由低成本LLM修复,不精确缺陷定位会扩大APR技术差距,高成本LLM与强推理设置并非总能提升成本效率,DeepSeek-V3.2成本效率最优。

Comments 20 pages, 6 figures, 10 tables. Accepted at ESEM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12921 2026-08-17 cs.MA cs.AI 版本更新 90%

Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference

通过因果推理发现基于大语言模型的多智能体系统的高效且可解释的通信拓扑

Junzhi Li, Peng He, Qirui Ji, Wei Wang, Lixiang Liu, Chuxiong Sun

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究针对基于LLM的多智能体系统通信拓扑可解释性不足的问题,提出模型无关框架E2-Explainer,通过因果推理识别关键通信子图,在保持任务性能的同时降低了通信成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12915 2026-08-14 cs.DC cs.AI 新提交 90%

InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers

InFactPlanner:可持续地理分布式大语言模型(LLM)数据中心规划

Nicoletta Tsiopani, Moysis Symeonides, George Pallis, Marios D. Dikaiakos

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI

AI总结 InFactPlanner是用于LLM推理的可持续地理分布式数据中心部署假设分析的决策支持框架,可通过多维度建模分析得出可持续性与延迟最优选择存在差异等结论。

Comments Author copy of paper published at 34th International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication System (MASCOTS2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03645 2026-08-12 cs.CL cs.CY 90%

LLM-MC-Affect: LLM-Based Monte Carlo Modeling of Affective Trajectories and Latent Ambiguity for Interpersonal Dynamic Insight

LLM-MC-Affect: 基于大语言模型的蒙特卡洛建模:情感轨迹与潜在模糊性的人际动态洞察

Yu-Zheng Lin, Bono Po-Jen Shih, John Paul Martin Encinas, Elizabeth Victoria Abraham Achom, Karan Himanshu Patel, Jesus Horacio Pacheco, Sicong Shao, Jyotikrishna Dass, Soheil Salehi, Pratik Satam

机构 * University of Arizona(亚利桑那大学) Pennsylvania State University(宾夕法尼亚州立大学) Universidad de Sonora(索尔纳大学) University of North Dakota(北达科他大学)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.CL

AI总结 本文提出LLM-MC-Affect框架,通过概率建模方法,将情感视为连续的潜在概率分布,从而捕捉人际互动中的情感轨迹和潜在模糊性,为动态分析提供新的视角和方法。

Comments Accepted to the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03611 2026-08-12 cs.DC cs.AI 版本更新 90%

Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling

Astrolabe:基于随机预测引导调度的大语言模型服务负载均衡

Wei Da, Evangelia Kalyvianaki

机构 * University of Cambridge(剑桥大学)

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出Astrolabe,一种用于LLM服务的随机预测引导调度器,结合响应长度估计等策略实现负载均衡,在多组基准测试中提升SLO容量、降低延迟并减少资源开销。

Comments 16 pages. Accepted at SYSTOR 2026. Camera-ready version with expanded evaluation and revisions. Previously circulated as "Block"; renamed "Astrolabe" to match the SYSTOR publication title

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07583 2026-08-11 stat.ML cs.LG 新提交 90%

RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough

RouteGuard:当互补性不足时,对LLM多智能体系统中的路由增益进行认证

Anchen Sun, Kaiqi Yang

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.LG

AI总结 针对LLM多智能体路由的部署问题,提出RouteGuard框架,通过分解增益、结合Le Cam下界等实现认证,在两个基准中验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09577 2026-08-11 cs.AI 新提交 90%

ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization

ElasticBack:通过耦合触发-规则优化实现LLM智能体技能中的隐蔽条件后门

Hao Sui, Simeng Qin, Jie Liao, Xiaojun Jia, Bing Chen, Yang Liu

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI

AI总结 本研究提出ElasticBack,通过耦合触发-规则优化在LLM智能体技能中植入条件性单技能后门,实验显示其攻击性能优异且隐蔽性强,可推动技能供应链的防御研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09555 2026-08-11 cs.AI 新提交 90%

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents

面向基于技能的大语言模型智能体强化学习的双向上下文自蒸馏

Tianjun Pan, Yuan Li, Hongda Wang, Linbo Jin, Mengfei Song, Lei Gao, Qiming Shi, Shaokang Fu, Jiarong Zhao, Chengyu Wang, Chengfu Huo

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究针对基于技能的LLM智能体技能利用不足问题,提出BCSD框架,通过双向上下文自蒸馏结合强化学习,在ALFWorld和WebShop上取得最优性能,验证了互补视角的有效性。

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09225 2026-08-11 cs.CR cs.AI 新提交 90%

Governing the KV Cache: Preventing Timing Side-Channel Leakage in Multi-Tenant LLM Inference

管控KV缓存:防止多租户大语言模型推理中的时序侧信道泄露

Tejasvi C. Addagada

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文针对多租户LLM推理中KV缓存引发的时序侧信道泄露问题,提出KVGov治理层,结合盐值机制与ORIGAMI调度器,可抵御三类攻击,同时保留93%的缓存效率,在真实硬件与模拟环境中验证了防御效果。

Comments 12 pages, 5 figures, 9 tables. Measurements on NVIDIA A100 with vLLM 0.26.0, independently replicated on llama.cpp/Apple Metal. Experimental scripts and raw results available from the author on request

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08382 2026-08-11 cs.AI cs.DC 新提交 90%

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving

LLMVisor:面向多租户大语言模型(LLM)服务的实时延迟归因模型

Shuowei Jin, Xueshen Liu, Jiaxin Shan, Le Xu, Tieying Zhang, Liguang Xie, Z. Morley Mao

机构 * University of Michigan(密歇根大学) ByteDance(字节跳动)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI

AI总结 针对多租户LLM服务中联合批处理导致的延迟归因难题,提出基于屋顶线模型的LLMVisor实时延迟归因模型,在微秒级高效运行,相比基线方法大幅降低预填充和解码阶段的相对误差

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07974 2026-08-11 cs.LG cs.DC 新提交 90%

ZeroLock: Concurrent Memory-Efficient LLM Training via Modular Update Decoupling

ZeroLock:通过模块化更新解耦实现并发内存高效的大语言模型训练

Wentao Dai, Xuanran Li, Yuxiang Zhang, Ming Tang, Chao Huang

机构 * Southern University of Science and Technology(南方科技大学) Montclair State University(蒙特克莱尔州立大学)

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 针对边缘LLM微调的BP训练存在更新锁定瓶颈,本研究提出ZeroLock无BP算法,通过模块化更新解耦突破局限,经实验验证可降内存26.5%、提吞吐量4.9%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07223 2026-08-11 cs.AI 版本更新 90%

Reflex First, Reflect Later: Latency-Aware Embodied LLM Agents for Dynamic Response

先反射,后反思:面向动态响应的延迟感知具身大语言模型智能体

Yangqing Zheng, Shunqi Mao, Dingxin Zhang, Weidong Cai

机构 * School of Computer Science, The University of Sydney(计算机科学学院,悉尼大学)

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究针对动态环境中具身LLM智能体的推理延迟问题,提出RRARA智能体及相关评估指标,通过时间转换机制与预规划器实现决策质量与响应能力的平衡。

Comments Accepted by the CVPR 2025 Embodied AI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07126 2026-08-10 cs.HC cs.AI 新提交 90%

PHOENIX: Fine-Tuned SLM-Powered Autonomous Satellite Lifetime Extension via Predictive Self-Healing and Multi-Agent AI Recovery

PHOENIX:通过预测性自修复与多智能体AI恢复实现的经微调小型语言模型驱动的自主卫星寿命延长

Sumaiya Islam, Harsha Kumara Moraliyage

专题命中 效率与部署 :SLM(title,summary_cn);language model(abstract);small language model(abstract);分类 cs.AI

AI总结 该研究针对CubeSat在轨故障无法及时修复导致寿命不足的问题,提出PHOENIX系统,利用微调SLM与多智能体AI实现自主故障修复,基于ESA基准验证了初步效果。

Comments 6 pages, 2 figures. Accepted at IEEE IRAI 2026 (International Conference on Responsible Artificial Intelligence)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07091 2026-08-10 cs.HC cs.AI 新提交 90%

Human-Centered Explainable AI for TinyML Edge Devices: A Pareto-Based Selection Framework with LLM-Guided Design

面向TinyML边缘设备的以人为中心的可解释人工智能:基于帕累托的选择框架与大语言模型引导设计

Zeinab Dehghani, Dhavalkumar Thakker, Koorosh Aslansefat, Kuniko Paxton, Bhupesh Kumar Mishra, Baseer Ahmad, Rameez Raja Kureshi

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究提出一种整合LLM引导设计与帕累托优化的以人为中心XAI选择框架,用于TinyML边缘设备部署,经皮肤病变分类任务验证可识别帕累托有效权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04065 2026-08-06 cs.CR cs.AI 新提交 90%

An Inline Control Architecture for Language Models in Intelligent Transportation Systems

智能交通系统中语言模型的内联控制架构

Narendra Kumar Dewangan, Mounira Msahli

专题命中 效率与部署 :LLM(summary_cn,abstract);language model(title,abstract);large language model(abstract);分类 cs.AI

AI总结 针对V2X系统中LLM的提示级攻击问题,提出Guarded-V2X内联语义护栏架构,经四阶段实验验证其可降低入侵成功率且不超延迟预算。

Comments 18 pages, under minor revision (IEEE transactions)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23035 2026-08-06 cs.LG cs.AR 90%

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration

OASIS:基于查找表的离群点感知双端量化LLM推理加速通用矩阵乘法

Xueying Wu, Baijun Zhou, Zhihui Gao, Yuzhe Fu, Qilin Zheng, Yintao He, Hai Li

机构 * National University of Singapore(新加坡国立大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 提出OASIS架构,利用预计算笛卡尔积查找表实现非均匀量化权重与激活的高效通用矩阵乘法,通过离群点感知量化方案和实时离群点检测引擎Orizuru,在保持精度的同时显著提升推理速度和能效。

Journal ref 2026 ACM/IEEE 53rd Annual International Symposium on Computer Architecture (ISCA), pp. 307-321

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03961 2026-08-05 cs.AI 新提交 90%

Interpretable Adaptive Sampling for LLM Test-Time Scaling

面向大语言模型测试时缩放的可解释自适应采样

Mobina Kashaniyan, Ali Jannesari

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究针对LLM测试时推理的固定计算预算不灵活、不可解释的问题,提出带轻量级模糊控制器的自适应采样方法,在问答和数学任务上提升性能并减少平均采样数,为高效测试时推理提供实用方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03463 2026-08-05 cs.AI 新提交 90%

LeanMem: Simple and Efficient Long-Term Memory for LLM Agents

LeanMem:面向LLM智能体的简单高效长期记忆系统

Yuxin Liao, Le Wu, Min Hou, Hao Liu, Han Wu, Zishu Wang

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI

AI总结 LeanMem是面向LLM智能体的轻量级长期记忆框架,通过差异化处理历史内容、动态分配资源,在两个数据集上提升了记忆基线的准确率,同时降低了成本与延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01536 2026-08-04 cs.AR cs.LG 新提交 90%

Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference

Celty:用于高效双稀疏大语言模型推理的SpMspV GPU内核与SIMT协同设计

Ruokai Yin, Priyadarshini Panda

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 针对双稀疏LLM推理中spMspV workload处理低效问题,提出协同设计的Celty稀疏格式、GPU内核与SIMT微架构,实现最高5.3倍加速。

Comments ICCAD 2026. Will update with the camera-ready version once ready. The code is available on Github at https://github.com/RuokaiYin/Celty

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00303 2026-08-04 cs.AI 新提交 90%

CrystalMem: Elastic Memory for Self-Evolving LLM Agents via Knowledge Crystallization

CrystalMem:基于知识结晶的自进化大语言模型智能体弹性内存

Beining Wu, Jun Huang

机构 * South Dakota State University(南达科他州立大学)

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究针对LLM智能体的内存滞后问题,提出弹性内存组件CrystalMem,通过知识结晶策略恢复能力,在多环境实验中性能优于基准方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00026 2026-08-04 cs.AI cs.DC 新提交 90%

Request-Level Energy Attribution for Batched LLM Serving

批量大语言模型(LLM)服务的请求级能耗归因

Qi Luo, Kunlin Li, Ziwen Wang, Dongsheng Wang, Yun Chen

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI

AI总结 本研究提出JouleShare框架,通过计算Shapley能耗建立请求级基准,用JCalib校准模型提升批量LLM服务的请求级能耗归因精度,验证了token比例归因的偏差及JCalib的有效性。

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21633 2026-08-04 cs.LG 版本更新 90%

HERALD: High-Throughput Block Diffusion LLM Serving via CPU-GPU Cooperative KV Cache Retrieval

HERALD: 通过CPU-GPU协同KV缓存检索实现高吞吐量块扩散LLM服务

Omin Kwon, Doyeon Kim, Jongseok Park, Seung Yul Lee, Ion Stoica, Jae W. Lee

机构 * Seoul National University(首尔国立大学)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.LG

AI总结 针对块扩散LLM中KV缓存随上下文增长导致吞吐量下降的问题,提出HERALD系统,利用块内去噪步骤间KV条目一致性,通过稀疏检索和计算重叠实现高精度低开销的KV卸载,在5-10%缓存预算下达到近无损精度,延迟降低1.59倍,吞吐量提升2.47倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07665 2026-08-04 cs.PL cs.AI 版本更新 90%

AgentCompile: An LLM-Guided Compiler for Direct CUDA Inference

AgentCompile:一种用于直接CUDA推理的LLM引导编译器

Xuanzhe Li, Ziyan Weng, Zhiyu Zhu, Junhui Hou

机构 * City University of Hong Kong (Dongguan)(香港城市大学(东莞)) City University of Hong Kong(香港城市大学)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI

AI总结 提出AgentCompile,利用LLM提供语义建议,通过模板生成CUDA候选实现并验证,在多个Transformer模型上实现4-5.7倍加速。

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25018 2026-08-03 cs.LG 版本更新 90%

Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference

共形级联:多层大语言模型推理的无分布精度保证

Yifan Dou, Shikan Lian, Shibo Li

机构 * Department of Computer Science, Florida State University(佛罗里达州立大学计算机科学系)

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 研究针对LLM级联推理成本高、置信分数校准不当等问题,提出共形级联框架,以共形预测集大小为推迟规则,提供无分布精度保证,在多基准测试中表现优于启发式级联,且无需模型训练,仅需黑盒API访问。

详情

展开后加载摘要…

URL PDF HTML 收藏