arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 22280 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 22280 篇

2604.13556 2026-04-16 cs.CL 89%

YOCO++: Enhancing YOCO with KV Residual Connections for Efficient LLM Inference

YOCO++:通过KV残差连接增强YOCO以实现高效的LLM推理

You Wu, Ziheng Chen, Yizhen Zhang, Haoyi Wu, Chengting Yu, Yuchi Xu, Wenbo Su, Bo Zheng, Kewei Tu

机构 * Alibaba Group(阿里巴巴集团) ShanghaiTech University(上海科技大学) Zhejiang University(浙江大学)

专题命中 效率与部署 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出YOCO++,通过在底层层与底层之间加入加权残差连接,提升YOCO的性能,实现50%的KV缓存压缩率下的最佳表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08003 2026-04-10 eess.AS cs.CL cs.SD 89%

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs

重新思考基于大语言模型的自动语音识别中的熵分配:理解语音编码器与大语言模型之间的动态关系

Yuan Xie, Jiaqi Song, Guang Qiu, Xianliang Wang, Ming Lei, Jie Gao, Jie Wu

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);pretraining(abstract)

AI总结 本文从熵分配角度重新审视基于大语言模型的ASR,提出三个指标刻画训练范式如何在语音编码器和LLM之间分配熵减少,通过多阶段训练策略优化参数效率和抗幻觉能力,实验显示在2.3B参数下达到与最新模型相当的性能并有效缓解幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12933 2026-03-16 cs.AI 89%

Efficient and Interpretable Multi-Agent LLM Routing via Ant Colony Optimization

通过蚁群优化实现高效且可解释的多智能体大语言模型路由

Xudong Wang, Chaoning Zhang, Jiaquan Zhang, Chenghao Li, Qigan Sun, Sung-Ho Bae, Peng Wang, Ning Xie, Jie Zou, Yang Yang, Hengtao Shen

机构 * School of Computing, Kyung Hee University(Kyung Hee 大学计算机学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(中国电子科技大学信息与软件工程学院) School of Computer Science and Engineering, University of Electronic Science and Technology of China(中国电子科技大学计算机科学与工程学院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);small language model(abstract)

AI总结 本文提出AMRO-S框架,通过意图推理、任务特定信息素专家和质量门控异步更新机制,提升多智能体系统路由效率与可解释性,实验显示其在质量-成本权衡上优于现有方法。

Comments 11 pages, 3 figures, submitted to IEEE Transactions on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22132 2026-01-30 cs.LG 89%

Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference

为提示付费,而非答案:用于高效推断的LLM牧师

Ziming Dong, Hardik Sharma, Evan O'Toole, Jaya Prakash Champati, Kui Wu

机构 * Department of Computer Science, University of Victoria, Victoria, Canada(维多利亚大学计算机科学系) Department Of Information Technology, Manipal University, Manipal, India(马那尔大学信息科技系)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);small language model(abstract)

AI总结 LLM牧师通过请求LLM的短提示来提高SLM的准确性,显著降低推断成本,同时保持准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22788 2025-12-01 cs.CR cs.CL 89%

PRISM: Privacy-Aware Routing for Adaptive Cloud-Edge LLM Inference via Semantic Sketch Collaboration

PRISM:通过语义草图协作实现隐私感知的自适应云-边LLM推理

Junfei Zhan, Haoxun Shen, Zheng Lin, Tengjiao He

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);small language model(abstract)

AI总结 PRISM通过语义草图协作实现隐私感知的自适应云-边LLM推理,动态平衡隐私与推理质量,降低能耗和延迟,提升输出质量。

Comments Accepted to AAAI 2026. This is the arXiv preprint version

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17257 2025-10-29 cs.LG q-bio.GN 89%

JanusDNA: A Powerful Bi-directional Hybrid DNA Foundation Model

Qihao Duan, Bingding Huang, Zhenqiao Song, Irina Lehmann, Lei Gu, Roland Eils, Benjamin Wild

机构 * Berlin Institute of Health, Charité – Universitätsmedizin Berlin(柏林健康研究所,柏林夏里特大学医学中心) College of Big Data and Internet, Shenzhen Technology University(大数据与互联网学院,深圳技术大学) Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学) Epigenetics Laboratory, Max Planck Institute for Heart and Lung Research(表观遗传学实验室,马克斯·普朗克心脏病和肺部研究所以) Department of Mathematics and Computer Science, Freie Universität Berlin(数学与计算机科学系,柏林自由大学) Health Data Science Unit, Heidelberg University Hospital and BioQuant(健康数据科学单元,海德堡大学医院和BioQuant) Intelligent Medicine Institute, Fudan University(智能医学研究所,复旦大学)

专题命中 效率与部署 :foundation model(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02133 2025-09-03 cs.CL 89%

AMBEDKAR-A Multi-level Bias Elimination through a Decoding Approach with Knowledge Augmentation for Robust Constitutional Alignment of Language Models

Snehasis Mukhopadhyay, Aryan Kasat, Shivam Dubey, Rahul Karthikeyan, Dhruv Sood, Vinija Jain, Aman Chadha, Amitava Das

机构 * Indian Institute of Information Technology, Kalyani(印度信息技术学院,卡里尼) BITS Pilani Goa(比尔·斯图尔特学院,果阿) IIT Madras(马德拉斯理工学院) DTU(达丁理工大学) Artificial Intelligence Institute, University of South Carolina(南卡罗来纳大学人工智能研究所) Meta AI Amazon GenAI(亚马逊生成人工智能)

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);small language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06402 2025-08-08 cs.RO cs.AI cs.HC 89%

Camera Control at the Edge with Language Models for Scene Understanding

Alexiy Buynitsky, Sina Ehsani, Bhanu Pallakonda, Pragyana Mishra

机构 * Purdue University(普渡大学)

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);SFT(abstract)

Comments 7 pages, 6 figures. This work was presented and published at the 11th IEEE International Conference on Control, Automation and Robotics (ICCAR) in 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16056 2025-04-23 cs.CL 89%

Honey, I Shrunk the Language Model: Impact of Knowledge Distillation Methods on Performance and Explainability

Daniel Hendriks, Philipp Spitzer, Niklas Kühl, Gerhard Satzger

机构 * Karlsruhe Institute of Technology (KIT)(卡尔斯鲁厄理工学院) Institute for Information Systems (WIN)(信息系统研究所) University of Bayreuth(拜罗伊特大学)

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);small language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12167 2025-03-20 cs.CL 89%

PLM: Efficient Peripheral Language Models Hardware-Co-Designed for Ubiquitous Computing

Cheng Deng, Luoyang Sun, Jiwen Jiang, Yongcheng Zeng, Xinjian Wu, Wenxin Zhao, Qingfa Xiao, Jiachuan Wang, Haoyang Li, Lei Chen, Lionel M. Ni, Haifeng Zhang, Jun Wang

专题命中 效率与部署 :language model(title,abstract);large language model(abstract);small language model(abstract);SFT(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11499 2024-12-17 cs.AI cs.RO 89%

Embodied CoT Distillation From LLM To Off-the-shelf Agents

Wonje Choi, Woo Kyung Kim, Minjong Yoo, Honguk Woo

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);small language model(abstract)

Comments Accepted at ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09758 2024-10-01 cs.CL 89%

OrchestraLLM: Efficient Orchestration of Language Models for Dialogue State Tracking

Chia-Hsuan Lee, Hao Cheng, Mari Ostendorf

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);small language model(abstract)

Comments updated version (NAACL camera ready)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06668 2024-03-19 cs.LG cs.CV 89%

Large Language Models and Foundation Models in Smart Agriculture: Basics, Opportunities, and Challenges

Jiajia Li, Mingle Xu, Lirong Xiang, Dong Chen, Weichao Zhuang, Xunyuan Yin, Zhaojian Li

专题命中 效率与部署 :large language model(title);language model(title);foundation model(title);分类 cs.LG

Comments 18 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06801 2024-02-20 cs.AI 89%

Graph-of-Thought: Utilizing Large Language Models to Solve Complex and Dynamic Business Problems

Ye Li

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

Comments Keywords: Graph-of-Thought (GoT), Workflow Automation, Large Language Models (LLMs), Task Execution, Data-Driven Decision Making, Complexity Management

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00819 2023-10-03 cs.CL 89%

Parameter-Efficient Tuning Helps Language Model Alignment

Tianci Xue, Ziqi Wang, Heng Ji

专题命中 效率与部署 :language model(title,abstract);large language model(abstract);RLHF(abstract);preference optimization(abstract)

Comments 21 pages, 11 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07093 2024-07-10 cs.CL cs.AI cs.LG 89%

FBI-LLM: Scaling Up Fully Binarized LLMs from Scratch via Autoregressive Distillation

Liqun Ma, Mingjie Sun, Zhiqiang Shen

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);pretraining(abstract)

Comments Github at https://github.com/LiqunMa/FBI-LLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07725 2026-08-19 cs.CL cs.AI 版本更新 88%

SOD: Step-wise On-policy Distillation for Small Language Model Agents

SOD:分步式在线蒸馏用于小型语言模型代理

Qiyong Zhong, Mao Zheng, Mingyang Song, Xin Lin, Jie Sun, Houcheng Jiang, Xiang Wang, Junfeng Fang

机构 * Zhejiang University(浙江大学) Large Language Model Department, Tencent(腾讯大语言模型部门) University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学)

专题命中 效率与部署 :language model(title,abstract);small language model(title,abstract);分类 cs.CL、cs.AI

AI总结 针对小型语言模型中工具集成推理的稳定性问题,提出SOD分步式在线蒸馏框架,通过动态调整蒸馏强度缓解教师信号误导,提升推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14644 2026-08-18 cs.LG cs.CL 新提交 88%

DUET: Dual-Teacher On-Policy Distillation via Same-Weight Disagreement for Prohibition Compliance

DUET:基于同权重分歧的双教师在线策略蒸馏用于禁止合规性

Zihan Li, Feifei Li, Wenhui Que

专题命中 效率与部署 :LLM(summary_cn,abstract);SFT(abstract,abstract_cn);post-training(abstract);分类 cs.CL、cs.LG

AI总结 本研究针对LLM部署中的动态禁止规则合规问题,提出DUET双教师在线策略蒸馏方法,构建工业基准,在Qwen模型上实现高合规性与效用保留,性能优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14604 2026-08-18 cs.CL cs.AI 新提交 88%

Wiola 13M, a Gated Spiral Attention Architecture for Parameter Efficient Small Language Models

Wiola 13M:一种用于参数高效小型语言模型的门控螺旋注意力架构

Aryuemaan Kumar Chowdhury, Praveen Oosa, Vineesha Reddy

专题命中 效率与部署 :language model(title,abstract);small language model(title,abstract);分类 cs.CL、cs.AI

AI总结 针对10M-100M参数小型语言模型未适配小规模Transformer的问题,提出含三个创新组件的Wiola 13M模型,验证其等价性并发布开源实现,提升长程区分与梯度流性能。

Comments 6

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03089 2026-08-14 cs.LG cs.AI 版本更新 88%

Constitutional On-Policy Safe Distillation

宪法性在策略安全蒸馏

Ming Wen, Yuxuan Liu, Kun Yang, Yunhao Feng, Zhuoer Xu, Yuhao Sun, Shiwen Cui, Xiang Zheng, Yi Liu, Xingjun Ma, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI(可信具身人工智能研究院) Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Ant Group(蚂蚁集团) Zhejiang University(浙江大学) City University of Hong Kong(香港城市大学)

专题命中 效率与部署 :SFT(summary_cn,abstract);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 针对在策略自蒸馏在安全对齐中因宪法条件导致教师分布收缩、表达能力下降的问题,提出宪法性在策略安全蒸馏(COPSD),通过交叉SFT冷启动校准教师分布,再进行宪法条件在策略蒸馏,在12个基准上实现了更优的安全-有用性权衡并降低安全税。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02068 2026-08-12 cs.CR cs.CL cs.LG 版本更新 88%

Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign

基于ML/密码学协同设计的大语言模型鲁棒安全代码水印

Ruisi Zhang, Neusha Javidnia, Nojan Sheybani, Farinaz Koushanfar

专题命中 效率与部署 :large language model(title);language model(title);LLM(abstract,abstract_cn);分类 cs.CL、cs.LG

AI总结 该研究提出首个ML/密码学协同设计的RoSeMary框架,通过端到端训练结合CodeT5实现高质量代码水印,用零知识证明安全验证,兼具高检测准确率、抗攻击能力与高效性。

Comments accept to ACM TAISAP

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15208 2026-08-11 cs.LG cs.AI 88%

Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels

量化消除了对齐:跨模型和精度水平的压缩LLM中的偏见涌现

Plawan Kumar Rath, Rahul Maliakkal

机构 * Meta

专题命中 效率与部署 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 研究通过压缩LLM的量化过程,发现3位量化导致6-21%的原本无偏项目出现新偏见,且标准质量指标无法检测到公平性关键的退化。

Comments 7 pages, 4 figures, 4 tables. Accepted at IEEE Cloud Summit 2026. This is the author's accepted version; the version of record will appear in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03873 2026-08-07 cs.LG cs.CL 版本更新 88%

SODA: Semi On-Policy Black-Box Distillation for Large Language Models

SODA:大语言模型的半在线盒式知识蒸馏

Xiwen Chen, Jingjing Wang, Wenhui Zhu, Peijie Qiu, Xuanzhao Dong, Yueyue Deng, Hejian Sang, Zhipeng Wang, Alborz Geramifard, Feng Luo

机构 * Clemson University(克莱姆森大学) LinkedIn(领英) Washington University in St. Louis(圣路易斯华盛顿大学) Arizona State University(亚利桑那州立大学) Columbia University(哥伦比亚大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.LG

AI总结 本文提出SODA,一种高效的半在线盒式知识蒸馏方法,通过对比教师最优响应与学生静态输出,实现高质量分布对齐,无需动态回放和对抗平衡,提升蒸馏效率和稳定性。

Comments Efficient Reasoning@COLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28639 2026-08-05 cs.CL cs.AI cs.CY 版本更新 88%

The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

知识蒸馏对小型语言模型偏见的非对称影响

Plawan Kumar Rath

机构 * Meta

专题命中 效率与部署 :language model(title,abstract);small language model(title);SFT(abstract_cn);分类 cs.CL、cs.AI

AI总结 该研究发现知识蒸馏对小型语言模型偏见存在非对称影响,明确任务下提升上下文遵循能力,模糊任务下破坏弃权校准,提出PCCD协议可捕捉聚合指标遗漏的危害。

Comments 18 pages, 5 figures. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01043 2026-07-30 cs.LG cs.AI 版本更新 88%

Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities

大语言模型的低精度训练:方法、挑战与机遇

Zhiwei Hao, Jianyuan Guo, Li Shen, Yong Luo, Han Hu, Guoxia Wang, Dianhai Yu, Yonggang Wen, Dacheng Tao

机构 * School of Information and Electronics, Beijing Institute of Technology(信息与电子学院,北京理工大学) Department of Computer Science, City University of Hong Kong(计算机科学系,香港城市大学) School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(网络科学与技术学院,中山大学深圳校区) School of Computer Science, Wuhan University(计算机科学学院,武汉大学) Baidu Inc.(百度公司) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.AI、cs.LG

AI总结 本综述全面梳理大语言模型低精度训练方法,按数值格式分类并探讨相关挑战、鲁棒性及研究方向,助力统一该领域研究视角。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22757 2026-07-28 cs.LG cs.AI 新提交 88%

Hierarchical Grading in Large Language Models

大语言模型中的分层分级

T. Shaska

机构 * Oakland University(奥克兰大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.AI、cs.LG

AI总结 研究提出分级大语言模型(GLLMs)框架,通过代数方法为变换器表示空间分级并传播加权标量作用,扩展相关理论到自回归语言模型。证明分层目标下分级与无分级的极小极大分离,最优分级可离线估计,训练后编译成标准变换器,兼具多种优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21702 2026-07-28 cs.CL cs.AI 版本更新 88%

CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference

CSV-Decode: 可证的子词汇解码以实现高效的大型语言模型推理

Dong Liu, Shu Wang, Yanxuan Yu, Haisheng Wang, Ben Lengerich

机构 * Department of Computer Science, Yale University(耶鲁大学计算机科学系) College of Engineering, Columbia University(哥伦比亚大学工程学院) Department of Statistics, University of Wisconsin-Madison(威斯康星大学麦迪逊分校统计学系)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 CSV-Decode通过构建子词汇表实现高效的大语言模型推理,提供精确的top-k认证和ε-认证的softmax近似,显著提升速度并保持分布保证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18046 2026-07-21 cs.LG cs.AI cs.CR 版本更新 88%

NanoZK: Privacy-Preserving Verifiable Inference for Large Language Models via Layerwise Zero-Knowledge Proofs

NANOZK:用于可验证大语言模型推理的分层零知识证明

Zhaohui Wang

机构 * USC Viterbi School of Engineering(USC维特里维学院工程学院)

专题命中 效率与部署 :large language model(title);language model(title);LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 NANOZK通过分层零知识证明验证大语言模型推理,采用分层证明框架实现常数大小证明,提升可扩展性与并行证明效率,同时保持形式正确性。

Comments Extended version. The first 12 pages correspond to the ICICS 2026 (Springer LNCS) camera-ready paper. This version supersedes the earlier ICLR 2026 VeriFAI Workshop preprint and adds full security proofs, Halo2 circuit details, lookup-table derivations, extended experiments, verifier-cost analysis, GPU scaling, reproducibility instructions, and appendices omitted from the proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08991 2026-07-13 cs.LG cs.CL 新提交 88%

Sensitivity-Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models

大语言模型中用于激活稀疏化的灵敏度感知阈值处理和令牌路由

Bishmoy Paul, Youngmin Yi, Hoeseok Yang

机构 * Santa Clara University(圣塔克拉拉大学) Sogang University(西江大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.LG

AI总结 研究大语言模型中如何在保留模型质量时减少计算,提出灵敏度感知阈值处理方法SATS及轻量级令牌路由框架,经实验评估,SATS在匹配稀疏度时优于基线,令牌路由有更好的质量-吞吐量权衡,改进方法可提升大语言模型权衡效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21082 2026-07-13 cs.CL cs.LG stat.ML 版本更新 88%

Accelerating Large Language Model Inference with Self-Supervised Early Exits

通过自监督早期退出加速大语言模型推理

Florian Valade

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.LG

AI总结 研究通过在中间层添加自监督早期退出头加速大语言模型推理,评估多种置信度指标,实验表明可降成本保精度,还应用于推测性解码提出DSSD,高令牌接受率且少超参数调整。

详情

展开后加载摘要…

URL PDF HTML 收藏