arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-04-22 至 2026-04-22 共收录 74 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 74 篇

2604.10401 2026-04-22 cs.CL 92%

NameBERT: Scaling Name-Based Nationality Classification with LLM-Augmented Open Academic Data

NameBERT: 通过LLM增强的开放学术数据扩展基于名称的国籍分类

Cong Ming, Ruixin Shi, Yifan Hu

机构 * Northeastern University, United States(美国东北大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出NameBERT,利用LLM增强开放学术数据构建大规模名称-国籍数据集,通过生成低资源国家名称提升分类性能,实现高效且准确的国籍分类。

Comments 12 pages, 3 figures, 8 tables; accepted at the 39th Canadian Conference on Artificial Intelligence (Canadian AI 2026)

Journal ref Proceedings of Machine Learning Research 318 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04445 2026-04-22 cs.NI cs.CL cs.PF 92%

Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey

动态模型路由与级联用于高效LLM推理:综述

Yasmin Moslem, John D. Kelleher

机构 * ADAPT Centre(ADAPT中心) School of Computer Science & Statistics(计算机科学与统计学学院) Trinity College Dublin(都柏林三一学院)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文综述了多LLM路由与级联方法,分析了不同路由范式及关键权衡,提出框架揭示路由系统在决策时机、信息使用和计算方式上的特性,强调实际系统常整合多种范式以优化性能。

Comments Work funded by ADAPT Centre, Trinity College Dublin, and Huawei Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19241 2026-04-22 cs.DC 91%

UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training

UniEP:统一的专家并行MoE MegaKernel用于LLM训练

Size Zheng, Xuegui Zheng, Li-wen Chang, Jidong Zhai

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 本文提出UniEP,通过统一专家并行优化策略,将MoE通信与计算融合为MegaKernel,实现自动化适应性调整,提升LLM训练效率并保持精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18764 2026-04-22 cs.AR 91%

CHICO-Agent: An LLM Agent for the Cross-layer Optimization of 2.5D and 3D Chiplet-based Systems

CHICO-Agent:一种用于2.5D和3D芯片片上系统的跨层优化LLM代理

Qihang Wu, Aman Arora, Vidya A. Chhabria

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract)

AI总结 针对2.5D/3D芯片片上系统设计复杂性问题,提出CHICO-Agent框架,通过LLM驱动优化,降低成本并提供可解释的审计跟踪。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18963 2026-04-22 cs.LG cs.AI 91%

Distillation Traps and Guards: A Calibration Knob for LLM Distillability

知识蒸馏陷阱与守护:一种用于LLM可蒸馏性的校准调节器

Weixiao Zhan, Yongcheng Jing, Leszek Rutkowski, Dacheng Tao

机构 * Generative AI Lab, College of Computing and Data Science(生成人工智能实验室,计算与数据科学学院) Nanyang Technological University(南洋理工大学) Systems Research Institute of the Polish Academy of Sciences(波兰科学院系统研究所) AGH University of Krakow(克拉科夫AGH大学) SAN University(SAN大学)

专题命中 效率与部署 :LLM(title,title_cn);SFT(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文分析了知识蒸馏中的陷阱,提出了一种后处理校准方法,通过强化学习微调控制教师模型的可蒸馏性,提升蒸馏效果和模型安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19167 2026-04-22 cs.LG cs.AI 91%

LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation

LBLLM:通过三阶段蒸馏实现大语言模型的轻量二值化

Siqing Song, Chuang Wang, Yong Lang, Yi Yang, Xu-Yao Zhang

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Central Media Technology Institute, Huawei(华为中央媒体技术研究所)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 LBLLM通过三阶段蒸馏策略实现大语言模型的轻量二值化,采用W(1+1)A4量化方法,在单GPU上仅用0.016B tokens训练,超越现有二值化方法,在语言模型、常识问答和语言理解任务中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18697 2026-04-22 cs.CR cs.CL cs.LG 90%

Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs

超越不可区分性:衡量LLM API中的提取风险

Ruixuan Liu, David Evans, Li Xiong

机构 * Emory University(埃默里大学) University of Virginia(弗吉尼亚大学)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.CL、cs.LG

AI总结 本文提出$(l, b)$-不可提取性作为衡量LLM API提取风险的新标准,通过定义提取风险上界并实验证明其在不同模型中的有效性,为模型训练、API访问和解码配置提供缓解指南。

Comments Accepted by S&P 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09775 2026-04-22 cs.AR cs.AI cs.DC cs.LG 90%

MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference

MIST:一种用于异构、多阶段LLM推理的联合设计框架

Abhimanyu Rajeshkumar Bambhaniya, Hanjiang Wu, Suvinay Subramanian, Sudarshan Srinivasan, Souvik Kundu, Amir Yazdanbakhsh, Midhilesh Elavazhagan, Madhu Kumar, Minlan Yu, Arijit Raychowdhury, Tushar Krishna

机构 * Georgia Institute of Technology(佐治亚理工学院) Google(谷歌) Intel(英特尔) Intel Labs(英特尔实验室) Google DeepMind(谷歌DeepMind) Harvard University(哈佛大学) Infravana

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI、cs.LG

AI总结 MIST是一种用于异构、多阶段LLM推理的联合设计框架,通过模拟不同请求阶段和复杂硬件层次,优化硬件-软件协同设计,解决LLM推理中的配置空间导航和跨厂商PD配置问题。

Comments Inference System Design for Multi-Stage AI Inference Pipelines. 11 Pages, 10 Figues, 5 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19049 2026-04-22 cs.CR cs.AI cs.SE 90%

Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery

反驳或促进:一种对抗性阶段门多智能体评审方法,用于高精度LLM辅助缺陷发现

Abhinav Agarwal

机构 * OpenSSL

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.AI

AI总结 本文提出Refute-or-Promote方法,通过对抗性阶段门多智能体评审,提高LLM辅助缺陷发现的精度,通过多种技术过滤虚假正例,最终发现多个漏洞并推动标准制定。

Comments 10 pages, 3 tables. Artifacts: https://github.com/abhinavagarwal07/refute-or-promote (Zenodo DOI: 10.5281/zenodo.19668799)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19299 2026-04-22 cs.CL cs.AI 90%

Rethinking Scale: Deployment Trade-offs of Small Language Models under Agent Paradigms

重新思考规模:在代理范式下小型语言模型的部署权衡

Xinlin Wang, Mats Brorsson

机构 * Proximus Luxembourg S.A.(普罗米修斯卢森堡股份有限公司) University of Luxembourg(卢森堡大学)

专题命中 效率与部署 :language model(title,abstract);small language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

AI总结 本文探讨了在代理范式下小型语言模型的部署权衡,通过对比基础模型、单代理工具系统和多代理协作系统,发现单代理系统在性能与成本间达到最佳平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18146 2026-04-22 cs.IR cs.AI cs.CL 90%

Modular Representation Compression: Adapting LLMs for Efficient and Effective Recommendations

模块化表示压缩:适应LLM以实现高效且有效的推荐

Yunjia Xi, Menghui Zhu, Jianghao Lin, Bo Chen, Ruiming Tang, Yong Yu, Weinan Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei Noah's Ark Lab(华为诺亚实验室) Antai College of Economics and Management, Shanghai Jiao Tong University(上海交通大学安泰经济管理学院)

专题命中 效率与部署 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MARC方法,通过模块化调整和任务解耦,解决LLM表示压缩中的MRA问题,提升推荐系统效率与效果。

Comments SIGIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19398 2026-04-22 cs.AI 89%

GRASPrune: Global Gating for Budgeted Structured Pruning of Large Language Models

GRASPrune:基于预算的大型语言模型结构剪枝的全局门控

Ziyang Wang, Jiangfeng Xiao, Chuan Xiao, Ruoxiang Li, Rui Mao, Jianbin Qin

机构 * Beijing Institute of Technology(北京理工大学) Shenzhen University(深圳大学) Osaka University(大阪大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);pretraining(abstract);分类 cs.AI

AI总结 GRASPrune通过全局预算约束联合剪枝FFN通道和KV头部组,采用轻量门分数学习和投影直通估计器,在保持模型权重冻结的情况下实现高效剪枝,并通过校准缩放因子提升推理性能。

Comments Accepted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19342 2026-04-22 cs.CL 89%

Are Large Language Models Economically Viable for Industry Deployment?

大语言模型在工业部署中经济上可行吗?

Abdullah Mohammad, Sushant Kumar Ray, Pushkar Arora, Rafiq Ali, Ebad Shabbir, Gautam Siddharth Kashyap, Jiechao Gao, Usman Naseem

机构 * DSEU-Okhla University of Delhi(德里大学) Macquarie University(麦考瑞大学) Center for SDGC(SDGC研究中心) Stanford University(斯坦福大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract_cn);分类 cs.CL

AI总结 本文提出EDGE-EVAL框架,评估大语言模型在工业场景中的全生命周期性能,引入五个部署指标,揭示模型在经济性和能效上的效率前沿及异常现象。

Comments Accepted at ACL 2026 (Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18592 2026-04-22 cs.CL cs.AI 89%

Two-dimensional early exit optimisation of LLM inference

二维早退优化大语言模型推理

Jan Hůla, David Adamczyk, Tomáš Filip, Martin Pavlíček, Petr Sosík

机构 * Institute for Research and Applications of Fuzzy Modelling(模糊建模研究与应用研究所) University of Ostrava(奥斯特拉瓦大学) Institute of Computer Science(计算机科学研究所) Faculty of Philosophy and Science(哲学与科学学院) Silesian University in Opava(奥帕瓦西里西亚大学)

专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出一种二维早退策略,通过协调层间和句子间早退实现大语言模型分类任务的计算优化,实验表明在不同规模的LLM上能显著提升推理速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29078 2026-04-22 cs.CL cs.LG 89%

PolarQuant: Optimal Gaussian Weight Quantization via Hadamard Rotation for LLM Compression

PolarQuant:通过Hadamard旋转实现最优高斯权重量化用于LLM压缩

Caio Vicentino

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 PolarQuant通过Hadamard旋转实现高斯分布的权重量化,显著降低压缩后的语言模型 perplexity,达到近无损压缩效果,并提升后续INT4量化性能。

Comments Found some errors, I need to fix

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09642 2026-04-22 cs.CL cs.AI 88%

MATA: Multi-Agent Framework for Reliable and Flexible Table Question Answering

MATA:用于可靠和灵活的表格问答的多智能体框架

Sieun Hyeon, Jusang Oh, Sunghwan Steve Cho, Jaeyoung Do

机构 * Department of Electrical and Computer Engineering, Seoul National University(首尔国立大学电气与计算机工程系) Interdisciplinary Program in Artificial Intelligence, Seoul National University(首尔国立大学人工智能跨学科项目)

专题命中 效率与部署 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);small language model(abstract)

AI总结 MATA通过多智能体框架和小语言模型工具,提升表格问答的可靠性与灵活性,实验显示其在不同LLM上均取得最佳准确率和高效推理。

Comments Accepted to ACL 2026 (findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07761 2026-04-22 cs.AI cs.LG 88%

TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards

TROJail: 多轮对话中针对大语言模型的多轮攻击优化

Xiqiao Xiong, Ouxiang Li, Zhuo Liu, Moxin Li, Wentao Shi, Fengbin Zhu, Qifan Wang, Fuli Feng

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Meta AI

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出TROJail,通过多轮强化学习优化攻击策略,引入过程奖励提升攻击成功率,实验显示在多个模型和基准上效果显著。

Comments Accepted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18610 2026-04-22 cs.NE cs.AI 88%

SpikeMLLM: Spike-based Multimodal Large Language Models via Modality-Specific Temporal Scales and Temporal Compression

基于脉冲的多模态大语言模型:通过模态特定的时间尺度和时间压缩

Han Xu, Zhiyong Qin, Di Shang, Jiahong Zhang, Xuerui Qiu, Bo Lei, Tiejun Huang, Bo Xu, Guoqi Li

机构 * 1 Institute of Automation, Chinese Academy of Sciences 2 University of Chinese Academy of Sciences 3 Beijing Academy of Artificial Intelligence 4 Zhongguancun Academy 5 Peking University 6 Key Laboratory of Brain Cognition Brain-inspired Intelligence Technology 7 Spiking Intelligence Lab, Tianqiao \& Chrissy Chen Institute [0.5em] Equal contribution Corresponding authors

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本文提出SpikeMLLM,首个基于脉冲的多模态大语言模型框架,通过模态特定时间尺度和时间压缩技术,在减少时间步数的同时保持高性能,实验显示其在多个基准上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19157 2026-04-22 cs.LG 87%

SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving

SAW-INT4: 为现实世界LLM服务的系统感知4位KV缓存量化

Jinda Jia, Jisen Li, Zhongzhu Zhou, Jung Hwan Heo, Jue Wang, Tri Dao, Shuaiwen Leon Song, Ben Athiwaratkun, Chenfeng Xu, Tianyi Zhang, Xiaoxia Wu

机构 * Together AI

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.LG

AI总结 本文提出SAW-INT4方法,通过系统感知的4位KV缓存量化,在满足实际服务约束下实现高精度与高效率的平衡,恢复近全部精度损失。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18788 2026-04-22 cs.LG 87%

Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs

高效混合专家LLM推理与Apple Silicon NPUs

Afsara Benazir, Felix Xiaozhu Lin

机构 * University of Virginia(弗吉尼亚大学)

专题命中 效率与部署 :LLM(title,title_cn);分类 cs.LG

AI总结 本文提出NPUMoE,通过将密集静态计算卸载到Apple Silicon NPUs,提升混合专家LLM推理效率,减少延迟、能耗和CPU使用率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09427 2026-04-22 cs.AR cs.AI 87%

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators

ODMA:面向LLM服务的LPDDR类加速器按需内存分配策略

Guoqiang Zou, Wanyu Wang, Hao Zheng, Longxiang Yin, Yinhe Han

机构 * University of Chinese Academy of Sciences(中国科学院大学) Beijing Information Science and Technology University(北京信息科技大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 ODMA针对LPDDR类加速器的随机访问带宽限制,提出按需内存分配策略,通过动态调整内存桶边界和安全池提升KV缓存利用率和吞吐量。

Comments 4 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06798 2026-04-22 cs.LG cs.AI 84%

MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization

MoBiE: 一种在后训练量化下高效混合二进制专家的推理方法

Zhixiong Zhao, Zukang Xu, Zhixuan Chen, Dawei Yang

专题命中 效率与部署 :post-training(title);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出MoBiE,一种专为基于混合专家(MoE)的大型语言模型(LLMs)设计的二进制化框架,通过减少交叉专家冗余、增强权重重要性估计和缓解路由扭曲,提升效率与性能。

Comments Although previously revised, per strict university regulations regarding incorrect affiliation, I am unauthorized to retain this manuscript. Furthermore, fundamental derivation errors in the NGES section compromise the mathematical framework, alongside misleading overlapping wording. The paper is therefore withdrawn

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19664 2026-04-22 cs.IR 83%

ECLASS-Augmented Semantic Product Search for Electronic Components

基于ECLASS的语义产品搜索增强

Nico Baumgart, Markus Lange-Hegermann, Jan Henze

专题命中 效率与部署 :LLM(summary_cn,abstract);foundation model(abstract)

AI总结 本文提出利用LLM密集检索和ECLASS层次语义提升电子元件语义搜索效果,实验显示其在准确率和效率上均优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16368 2026-04-22 cs.CL 83%

Cross-Family Speculative Decoding for Polish Language Models on Apple~Silicon: An Empirical Evaluation of Bielik~11B with UAG-Extended MLX-LM

跨家族推测解码用于波兰语言模型在苹果Silicon:对Bielik~11B与扩展MLX-LM的实证评估

Krzysztof Fonal

机构 * Wrocław University of Science and Technology(沃拉夫大学科学与技术学院)

专题命中 效率与部署 :language model(title);LLM(abstract,abstract_cn);分类 cs.CL

AI总结 本文评估了在苹果Silicon上使用扩展MLX-LM框架实现跨 tokenizer 推测解码的有效性,探讨了波兰语言模型的性能差异及统一内存架构下的实证结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04800 2026-04-22 cs.CL 83%

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights

语言模型的混合架构:系统分析与设计洞察

Sangmin Bae, Bilge Acun, Chien-Yu Lin, Haroun Habeeb, Seungyeon Kim, Liang Luo, Junjie Wang, Carole-Jean Wu

机构 * FAIR at Meta(Meta的FAIR) Meta

专题命中 效率与部署 :language model(title,abstract);large language model(abstract);分类 cs.CL

AI总结 本文系统分析了混合架构的设计,探讨了不同融合策略对语言模型性能、长上下文能力及效率的影响,提出优化设计方法。

Comments 41 pages, 8 figures, 22 tables;

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17789 2026-04-22 cs.CV cs.AI cs.CL 82%

DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization

DuQuant++: 细粒度旋转增强微缩放FP4量化

Haokun Lin, Xinle Jia, Haobo Xu, Bingchen Yao, Xianglong Guo, Yichen Wu, Zhichao Lu, Ying Wei, Qingfu Zhang, Zhenan Sun

机构 * CASIA(中国科学院自动化研究所) NJU(南京大学) THU(清华大学) ZJU(浙江大学) Harvard(哈佛大学) CityU(城市大学)

专题命中 效率与部署 :LLM(summary_cn,abstract);分类 cs.CL、cs.AI

AI总结 DuQuant++通过细粒度旋转优化MXFP4微缩放格式,解决激活异常带来的量化误差问题,提升LLM推理效率。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08899 2026-04-22 cs.CL cs.LG 82%

ConFu: Contemplate the Future for Better Speculative Sampling

ConFu:为更好的推测采样展望未来

Zongyue Qin, Raghavv Goel, Mukul Gagrani, Risheek Garrepalli, Mingu Lee, Yizhou Sun

机构 * University of California Los Angeles, United States(美国加州大学洛杉矶分校) Qualcomm AI Research, United States(高通人工智能研究)

专题命中 效率与部署 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG

AI总结 ConFu提出一种新的推测解码框架,通过引入展望令牌和软提示,使草案模型能利用目标模型的未来信号,提升生成速度和接受率。

Comments v3: Added stress test with long drafts (DL=12, top-k=1) and tail-acceptance (survival) analysis. Earlier versions added Qwen3-4B results and ablations

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18909 2026-04-22 cs.AR 82%

ChipLight: Cross-Layer Optimization of Chiplet Design with Optical Interconnects for LLM Training

ChipLight:基于光学互连的芯片片级设计跨层优化用于大语言模型训练

Kangbo Bai, Zhantong Zhu, Yifan Ding, Tianyu Jia

专题命中 效率与部署 :LLM(title,abstract)

AI总结 ChipLight通过跨层多目标设计优化方法,结合芯片片和光学互连技术,提升训练集群的通信效率与性能。

Comments Accepted by DATE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14170 2026-04-22 cs.CL cs.CY 81%

Comparing energy consumption and accuracy in text classification inference

在文本分类推理中比较能耗与准确性

Johannes Zschache, Tilman Hartwig

机构 * Application Lab for AI and Big Data(人工智能与大数据应用实验室)

专题命中 效率与部署 :LLM(summary_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究比较了文本分类推理中模型准确性和能耗的权衡,发现高准确度模型可能也具备能效,LLM在零样本分类中与传统模型有相似或更低的准确率,能耗受模型类型、规模和硬件影响显著。

Comments Key results in Figure 2, accepted in Nature Sci Rep, 32 pages

Journal ref Sci Rep 16, 12717 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18913 2026-04-22 cs.CL 81%

LogosKG: Hardware-Optimized Scalable and Interpretable Knowledge Graph Retrieval

LogosKG:硬件优化的可扩展且可解释的知识图谱检索

He Cheng, Yifu Wu, Saksham Khatwani, Maya Kruse, Dmitriy Dligach, Timothy A. Miller, Majid Afshar, Yanjun Gao

机构 * LARK Lab, University of Colorado Anschutz(洛克拉克实验室,科罗拉多大学安施图茨分校) University of Colorado Boulder(科罗拉多大学波德分校) Loyola University Chicago(芝加哥洛克拉克大学) Harvard Medical School(哈佛医学院) Boston Children’s Hospital(波士顿儿童医院) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 效率与部署 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 LogosKG通过符号知识图谱和硬件高效操作实现大规模知识图谱的多跳检索,提升效率与可解释性,展示出在生物医学知识与大语言模型推理对齐分析中的应用价值。

Comments Accepted to the ACL 2026 Main Conference. 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏