arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 22280 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 22280 篇

2509.15518 2025-09-22 cs.CL cs.AI cs.LG 87%

How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages

Siyang Wu, Zhewei Sun

机构 * Data Science Institute University of Chicago(芝加哥大学数据科学研究所) Toyota Technological Institute at Chicago(芝加哥丰田技术研究所)

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07824 2025-09-22 cs.CL cs.AI cs.LG 87%

Efficient Real-time Refinement of Language Model Text Generation

Joonho Ko, Jinheon Baek, Sung Ju Hwang

机构 * KAIST(韩国科学技术院)

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13541 2025-09-18 cs.CL cs.AI cs.LG 87%

Annotation-Efficient Language Model Alignment via Diverse and Representative Response Texts

Yuu Jinnai, Ukyo Honda

机构 * CyberAgent / Tokyo, Japan(CyberAgent)

专题命中 效率与部署 :language model(title,abstract);large language model(abstract);preference optimization(abstract);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP Findings, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02694 2025-09-15 cs.LG cs.AI cs.CL cs.CR cs.DC 87%

MeanCache: User-Centric Semantic Caching for LLM Web Services

Waris Gill, Mohamed Elidrisi, Pallavi Kalapatapu, Ammar Ahmed, Ali Anwar, Muhammad Ali Gulzar

机构 * Virginia Tech(弗吉尼亚理工大学) Cisco(思科) University of Minnesota(明尼苏达大学)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at 2025 IEEE 39th International Parallel and Distributed Processing Symposium (IPDPS)

Journal ref 2025 IEEE International Parallel and Distributed Processing Symposium (IPDPS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15486 2025-09-04 cs.CL cs.AI cs.LG 87%

SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention

Qianchao Zhu, Jiangfei Duan, Chang Chen, Siran Liu, Guanyu Feng, Xin Lv, Xiao Chuanfu, Dahua Lin, Chao Yang

专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);pretraining(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04676 2025-08-07 cs.CL cs.AI cs.LG 87%

GeRe: Towards Efficient Anti-Forgetting in Continual Learning of LLM via General Samples Replay

Yunan Zhang, Shuoran Jiang, Mengchen Zhao, Yuefeng Li, Yang Fan, Xiangping Wu, Qingcai Chen

机构 * Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学)

专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);pretraining(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09820 2025-07-15 cs.SE cs.CY 87%

Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications

Jia Yi Goh, Shaun Khoo, Nyx Iskandar, Gabriel Chua, Leanne Tan, Jessica Foo

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14432 2025-06-13 cs.CL cs.AI cs.LG 87%

PLAY2PROMPT: Zero-shot Tool Instruction Optimization for LLM Agents via Tool Play

Wei Fang, Yang Zhang, Kaizhi Qian, James Glass, Yada Zhu

机构 * Massachusetts Institute of Technology(麻省理工学院) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2025 Long Paper (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12872 2025-06-06 cs.CL cs.AI cs.CY cs.LG 87%

Not All Options Are Created Equal: Textual Option Weighting for Token-Efficient LLM-Based Knowledge Tracing

JongWoo Kim, SeongYeub Chu, Bryan Wong, Mun Yi

机构 * KAIST(韩国科学技术院)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04358 2025-05-30 cs.CL cs.AI cs.CC cs.LG cs.NE 87%

Position: Scaling LLM Agents Requires Asymptotic Analysis with LLM Primitives

Elliot Meyerson, Xin Qiu

机构 * Elliot Meyerson Xin Qiu

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments In Proceedings of the 42nd International Conference on Machine Learning (ICML 2025); 13 pages including references

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22192 2025-05-29 cs.MA 87%

Efficient Leave-one-out Approximation in LLM Multi-agent Debate Based on Introspection

Yue Cui, Liuyi Yao, Zitao Li, Yaliang Li, Bolin Ding, Xiaofang Zhou

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15576 2025-05-28 cs.RO cs.CV 87%

QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning

Xinyang Tong, Pengxiang Ding, Yiguo Fan, Donglin Wang, Wenjie Zhang, Can Cui, Mingyang Sun, Han Zhao, Hongyin Zhang, Yonghao Dang, Siteng Huang, Shangke Lyu

机构 * MiLAB, Westlake University, Hangzhou, 310030, China(西交利物浦大学微实验室,杭州,310030,中国) Zhejiang University, Hangzhou, 310027, China(浙江大学,杭州,310027,中国) Beijing University of Posts and Telecommunications, Beijing, 100876, China(北京邮电大学,北京,100876,中国)

专题命中 效率与部署 :language model(title,abstract);large language model(abstract);foundation model(abstract);instruction tuning(abstract)

Comments Accepted to ICRA 2025; Github page: https://quart-online.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16638 2025-05-27 cs.AI cs.CL cs.LG 87%

LLMScan: Causal Scan for LLM Misbehavior Detection

Mengdi Zhang, Kai Kiat Goh, Peixin Zhang, Jun Sun, Rose Lin Xin, Hongyu Zhang

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17548 2025-05-26 cs.DC 87%

H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips

Ding Tang, Jiecheng Zhou, Jiakai Hu, Shengwei Li, Huihuang Zheng, Zhilin Pei, Hui Wang, Xingcheng Zhang

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17074 2025-05-26 cs.CL cs.AI cs.LG 87%

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency

Ruixiao Li, Fahao Chen, Peng Li

机构 * School of Cyber Science and Engineering, Xi’an Jiaotong University(西安交通大学计算机科学与工程学院) School of Computer Science and Engineering, The University of Aizu(台湾花莲大学计算机科学与工程学院)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12162 2025-05-20 cs.CL cs.AI cs.DC cs.LG 87%

AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding

Zikun Li, Zhuofu Chen, Remi Delacourt, Gabriele Oliaro, Zeyu Wang, Qinghan Chen, Shuhuai Lin, April Yang, Zhihao Zhang, Zhuoming Chen, Sean Lai, Xinhao Cheng, Xupeng Miao, Zhihao Jia

机构 * Carnegie Mellon University(卡内基梅隆大学) Tongji University(同济大学) EPFL(苏黎世联邦理工学院) Amazon Web Services(亚马逊网络服务)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11271 2025-05-19 cs.CL cs.AI cs.IR cs.LG 87%

Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models

Camille Couturier, Spyros Mastorakis, Haiying Shen, Saravan Rajmohan, Victor Rühle

机构 * Microsoft 365 Research(微软365研究) University of Virginia(弗吉尼亚大学)

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Preprint. Paper accepted at ICCCN 2025, the final version will appear in the proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01259 2025-05-09 cs.HC 87%

Facilitating Instructors-LLM Collaboration for Problem Design in Introductory Programming Classrooms

Muntasir Hoq, Jessica Vandenberg, Shuyin Jiao, Seung Lee, Bradford Mott, Narges Norouzi, James Lester, Bita Akram

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

Comments Accepted at CHI 2025 Workshop on Augmented Educators and AI: Shaping the Future of Human and AI Cooperation in Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01459 2025-05-06 cs.CL cs.AI cs.LG 87%

MoxE: Mixture of xLSTM Experts with Entropy-Aware Routing for Efficient Language Modeling

Abdoul Majid O. Thiombiano, Brahim Hnich, Ali Ben Mrad, Mohamed Wiem Mkaouer

机构 * FSM, University of Monastir(FSM,蒙斯塔尔大学) CES Lab, ENIS, University of Sfax(CES实验室,ENIS,萨菲大学) Department of Computer Science, College of Computer, Qassim University(计算机科学系,计算机学院,卡西姆大学) University of Michigan-Flint(密歇根大学弗林特分校)

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04532 2025-05-02 cs.CL cs.AI cs.LG cs.PF 87%

QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, Song Han

机构 * MIT(麻省理工学院)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments The first three authors contribute equally to this project and are listed in the alphabetical order. Yujun Lin leads the quantization algorithm, Haotian Tang and Shang Yang lead the GPU kernels and the serving system. Code is available at https://github.com/mit-han-lab/omniserve

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19720 2025-04-29 cs.CL cs.AI cs.DC cs.LG 87%

Taming the Titans: A Survey of Efficient LLM Inference Serving

Ranran Zhen, Juntao Li, Yixin Ji, Zhenlin Yang, Tong Liu, Qingrong Xia, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Min Zhang

机构 * Soochow University(苏州大学) Huawei Cloud(华为云)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments work in progress;11 pages of main paper with 7 main figures, overall 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14866 2025-04-22 cs.CL cs.AI cs.DC cs.LG cs.PF 87%

LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention

Shang Yang, Junxian Guo, Haotian Tang, Qinghao Hu, Guangxuan Xiao, Jiaming Tang, Yujun Lin, Zhijian Liu, Yao Lu, Song Han

机构 * MIT(麻省理工学院)

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by MLSys 2025. Code available at: https://github.com/mit-han-lab/omniserve

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04667 2025-04-03 cs.CL cs.AI cs.LG cs.SE 87%

Non-Determinism of "Deterministic" LLM Settings

Berk Atil, Sarp Aykent, Alexa Chittams, Lisheng Fu, Rebecca J. Passonneau, Evan Radcliffe, Guru Rajan Rajagopal, Adam Sloan, Tomasz Tudrej, Ferhan Ture, Zhe Wu, Lixinyu Xu, Breck Baldwin

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12820 2025-04-01 cs.CL cs.AI cs.LG 87%

PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Hailin Zhang, Xiaodong Ji, Yilin Chen, Fangcheng Fu, Xupeng Miao, Xiaonan Nie, Weipeng Chen, Bin Cui

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09949 2025-03-26 cs.LG cs.AI cs.AR cs.CL 87%

Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models

Wenqi Jiang, Marco Zeller, Roger Waleffe, Torsten Hoefler, Gustavo Alonso

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by VLDB 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13427 2025-03-18 cs.LG cs.AI cs.CL 87%

xLSTM 7B: A Recurrent LLM for Fast and Efficient Inference

Maximilian Beck, Korbinian Pöppel, Phillip Lippe, Richard Kurle, Patrick M. Blies, Günter Klambauer, Sebastian Böck, Sepp Hochreiter

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Code available at: https://github.com/NX-AI/xlstm and https://github.com/NX-AI/xlstm-jax

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11794 2025-03-18 cs.CV cs.AI cs.CL cs.LG 87%

Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection

Bangzheng Li, Fei Wang, Wenxuan Zhou, Nan Xu, Ben Zhou, Sheng Zhang, Hoifung Poon, Muhao Chen

专题命中 效率与部署 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04412 2025-03-05 cs.LG cs.AI cs.CL 87%

Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment

Dongyoung Kim, Kimin Lee, Jinwoo Shin, Jaehyung Kim

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICLR 2025 Oral Presentation, 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06842 2025-03-03 cs.LG cs.AI cs.CL 87%

SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training

Tianjin Huang, Ziquan Zhu, Gaojie Jin, Lu Liu, Zhangyang Wang, Shiwei Liu

专题命中 效率与部署 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08155 2025-02-26 cs.LG cs.AI cs.CL 87%

QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts

Pingzhi Li, Xiaolong Jin, Zhen Tan, Yu Cheng, Tianlong Chen

专题命中 效率与部署 :post-training(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Our code for reproducing all our experiments is provided at https://github.com/UNITES-Lab/moe-quantization

详情

展开后加载摘要…

URL PDF HTML 收藏