arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 22116 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 22116 篇

2604.00626 2026-06-19 cs.LG cs.CL 版本更新 91%

A Survey of On-Policy Distillation for Large Language Models

大型语言模型的在线策略蒸馏综述

Mingyang Song, Mao Zheng

机构 * Tencent, China(腾讯,中国)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);RLHF(abstract,abstract_cn);分类 cs.CL、cs.LG

AI总结 本文综述了大型语言模型的在线策略蒸馏方法,探讨了蒸馏过程中如何通过反馈减少累积误差,提出了基于f-散度最小化的蒸馏框架,并分析了蒸馏与强化学习之间的联系。

Comments Ongoing Work

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11605 2026-06-11 cs.LG cs.AI 新提交 91%

Physics-Distilled Neural Network enabled by Large Language Models for Manufacturing Process-Property Predictive Modeling

基于大语言模型的物理蒸馏神经网络用于制造过程-性能预测建模

Ge Song, Kiarash Naghavi Khanghah, Anandkumar Patel, Rajiv Malhotra, Hongyi Xu

机构 * School of Mechanical, Aerospace and Manufacturing Engineering, University of Connecticut(康奈尔大学机械、航空航天与制造工程学院) Department of Mechanical & Aerospace Engineering, Rutgers, the State University of New Jersey(新泽西州立大学鲁特大学机械与航空航天工程学院)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 提出一种知识蒸馏框架,利用大语言模型从文献中提取物理先验,通过图掩码注意力层捕获变量依赖,蒸馏至轻量学生模型,在数据稀缺下实现高精度预测与实时部署。

Comments Under review, Journal of Computing and Information Science in Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00369 2026-06-08 cs.LG cs.AI 版本更新 91%

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees

InvEvolve:通过具有性能保证的大语言模型进化白盒库存策略

Chenyu Huang, Jianghao Lin, Zhengyang Tang, Bo Jiang, Ruoqing Jiang, Benyou Wang, Lai Wei

机构 * Shanghai University of Finance and Economics(上海财经大学) Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Tsinghua University(清华大学) Boston College(波士顿大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 提出InvEvolve框架,利用强化学习训练的大语言模型,结合置信区间认证,在线生成具有统计安全保证的白盒库存策略,在合成和真实零售数据上优于经典和深度学习方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13056 2026-06-05 cs.CL cs.AI 91%

Channel-Wise Mixed-Precision Quantization for Large Language Models

通道级混合精度量化用于大语言模型

Zihan Chen, Bike Xie, Jundong Li, Cong Shen

机构 * Department of Electrical and Computer Engineering, University of Virginia(电气与计算机工程系,弗吉尼亚大学) Kneron Inc.(芯驰科技)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本文提出通道级混合精度量化(CMPQ),通过根据激活分布分配不同精度级别来优化大语言模型的量化过程,从而在低比特范围内实现任意平均比特宽度,并在内存使用增加有限的情况下提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02964 2026-06-03 cs.AR cs.CL cs.LG 91%

Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving

多段注意力:实现高效KV缓存管理以加速大型语言模型服务

Chunan Shi, Yilei Chen, Yilin Chen, Xupeng Miao, Bin Cui

机构 * Peking University(北京大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.LG

AI总结 提出AsymCache,一种计算延迟感知的KV缓存管理系统,通过多段注意力、缓存驱逐策略和自适应分块调度器,在保持无损精度的同时显著降低TTFT和TPOT。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09711 2026-06-03 cs.CL cs.AI 91%

ReaLM: Residual Quantization Bridging Knowledge Graph Embeddings and Large Language Models

ReaLM:残差量化桥接知识图谱嵌入与大型语言模型

Wenbin Guo, Xin Wang, Jiaoyan Chen, Lingbing Guo, Zhao Li, Zirui Chen

机构 * Tianjin University(天津大学) The University of Manchester(曼彻斯特大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 提出ReaLM框架,通过残差向量量化将知识图谱嵌入离散化为可学习标记,融入大型语言模型词汇表,结合本体约束实现结构化知识与语言模型的语义对齐,在知识图谱补全任务上取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01099 2026-06-02 cs.CL cs.AI 91%

MiCU: End-to-End Smart Home Command Understanding with Large Language Model

MiCU: 基于大语言模型的端到端智能家居指令理解

Haowei Han, Kexin Hu, Weiwei Cai, Debiao Zhang, Bin Qin, Yuxiang Wang, Jiawei Jiang, Xiao Yan, Bo Du

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Xiaomi Corporation(小米公司) Institute for Math & AI, Wuhan University(武汉大学数学与人工智能研究院)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 提出MiCU,一种利用课程学习、强化学习和令牌压缩技术的领域特定大语言模型,用于解决智能家居中模糊指令理解问题,平均准确率提升20.01%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23485 2026-06-02 cs.CL cs.AI cs.CY 91%

Failure of contextual invariance in large language models

大型语言模型中语境不变性的失效

Sagar Kumar, Ariel Flint, Luca Maria Aiello, Andrea Baronchelli

机构 * Network Science Institute, Northeastern University(网络科学研究所,东北大学) Center for Health Informatics Program, Boston Children’s Hospital(健康信息学计划中心,波士顿儿童医院) Dept. of Mathematics, City St George’s, University of London(伦敦大学城市圣乔治学院数学系) IT University of Copenhagen(哥本哈根IT大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 通过代词选择任务发现,在语境等价但无信息量的干扰下,大语言模型输出发生系统性偏移,表明其违反语境不变性,影响偏见评估与高风险应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25837 2026-06-02 cs.LG cs.AI 91%

Distillation of Large Language Models via Concrete Score Matching

通过具体分数匹配进行大型语言模型的蒸馏

Yeongmin Kim, Donghyeok Shin, Mina Kang, Byeonghu Na, Il-Chul Moon

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 提出具体分数蒸馏(CSD)目标,通过离散分数匹配克服softmax平滑和logit平移不变性限制,实现学生与教师模型间所有词汇对相对logit差异的灵活加权,在GPT-2、OpenLLaMA和GEMMA上优于现有蒸馏方法。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08204 2026-06-01 cs.CL cs.AI 91%

Human-Alignment and Calibration of Inference-Time Uncertainty in Large Language Models

大型语言模型中推理时间不确定性的人类对齐与校准

Kyle Moore, Jesse Roberts, Daryl Watson

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本文评估了多种推理时间不确定性度量,发现它们与人类群体不确定性高度对齐,尽管与人类答案偏好不一致,但在正确性相关性和分布分析上表现出中等到强校准证据。

Comments We have discovered a critical error in the normalized entropy calculation that may have substantially inflated nearly all results herein. We have since fixed this error in a new work, but we believe that the new work is sufficiently dissimilar in focus, methods, dataset, and results as to be misleading if presented as a simple replacement. As such, we propose removal and retraction instead

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15236 2026-05-29 cs.CR cs.AI cs.LG 91%

Jailbreaking and Mitigation of Vulnerabilities in Large Language Models

大语言模型的越狱与漏洞缓解

Benji Peng, Hanxuan Chen, Keyu Chen, Qian Niu, Ziqian Bi, Ming Liu, Pohsun Feng, Tianyang Wang, Lawrence K. Q. Yan, Yizhu Wen, Yichao Zhang, Caitlyn Heqi Yin, Xinyuan Song, Riyang Bao, Jiacheng Shi

机构 * Hunan University Changsha, PRC Georgia Institute of Technology Atlanta, USA Kyoto University Kyoto, Japan Purdue University West Lafayette, USA National Taiwan Normal University Taipei, ROC University of Liverpool Suzhou, PRC Hong Kong University of Science University of Hawaii Honolulu, USA The University of Texas at Dallas Dallas, USA University of Wisconsin-Madison Madison, USA Emory University Atlanta, USA College of William \& Mary Williamsburg, USA

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

AI总结 本文综述了大语言模型在提示注入和越狱攻击下的漏洞,分类攻击方法并评估防御策略,指出研究空白与未来方向。

Journal ref Eureka 1(1) (2026) 26-61

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11774 2026-05-13 cs.CL cs.LG 91%

From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction

从标记到标记对:用于临床预测的大型语言模型高效提示压缩

Mingcheng Zhu, Zhiyao Luo, Yu Liu, Tingting Zhu

机构 * Department of Engineering Science, University of Oxford, Oxford, United Kingdom(牛津大学工程科学系,英国牛津)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.LG

AI总结 本文提出MedTPE方法,通过将频繁共现的医疗标记对合并为复合标记,实现无损压缩,减少输入标记长度和推理延迟,同时保持或提升预测性能。

Comments 21 pages, 6 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11513 2026-05-13 cs.CL cs.AI 91%

A Study on Hidden Layer Distillation for Large Language Model Pre-Training

针对大语言模型预训练的隐藏层蒸馏研究

Maxime Guigon, Lucas Dixon, Michaël E. Sander

机构 * Google DeepMind(谷歌DeepMind)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本文研究了隐藏层蒸馏在大语言模型预训练中的应用,对比了蒸馏与基于输出logits的方法,发现蒸馏在某些配置下能提升困惑度,但未在下游任务中超越标准蒸馏。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21159 2026-05-13 cs.AI cs.LG cs.MA 91%

MAC: Masked Agent Collaboration Boosts Large Language Model Medical Decision-Making

MAC:掩码代理协作提升大语言模型医疗决策能力

Zhihao Peng, Liuxin Bao, Yixuan Yuan

机构 * School of Automation, Hangzhou Dianzi University(杭州电子大学自动化学院) Department of Electronic Engineering, Chinese University of Hong Kong(香港中文大学电子工程系)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 本文提出MAC框架,通过帕累托最优代理构建和交叉一致性最大化机制,提升医疗决策能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27527 2026-05-12 cs.LG cs.AI 91%

TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control

TetraJet-v2:用于大语言模型的准确NVFP4训练方法,具有振荡抑制和异常值控制

Yuxiang Chen, Yifan Liu, Xiaoming Xu, Pengle Zhang, Michael Beyer, Martin Rapp, Jun Zhu, Jianfei Chen

机构 * Dept. of Comp. Sci. and Tech., Institute for AI, BNRist Center, THBI Lab, Tsinghua-Bosch Joint ML Center, Tsinghua University(计算机科学与技术系,人工智能研究所,BNRist中心,THBI实验室,清华-博世联合机器学习中心,清华大学) Zhili College, Tsinghua University(紫荆学院,清华大学) Bosch AI Research, Renningen, Germany(博世人工智能研究,德国Renningen)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 TetraJet-v2通过NVFP4格式实现低精度训练,解决权重振荡和异常值问题,提升大语言模型性能,减少与BF16的性能差距达51.3%,并实现1.67倍速度提升。

Journal ref Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 (ICML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05693 2026-05-11 cs.AI cs.LG 91%

Saliency-Aware Regularized Quantization Calibration for Large Language Models

具有显著性感知的正则化量化校准用于大语言模型

Yanlong Zhao, Xiaoyuan Cheng, Huihang Liu, Baihua He, Xinyu Zhang, Harrison Bo Hua Zhu, Wenlong Chen, Li Zeng, Zhuo Sun

机构 * University of Science and Technology of China(中国科学技术大学) University College London(伦敦大学学院) Shanghai University of Finance and Economics(上海财经大学) Academy of Mathematics and Systems Science, Chinese Academy of Sciences(中国科学院数学与系统科学研究院) University of Copenhagen(哥本哈根大学) Imperial College London(伦敦帝国学院) Technical University of Denmark(丹麦技术大学) Peking University(北京大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract_cn);post-training(abstract)

AI总结 本文提出SARQC,通过引入显著性感知正则化改进量化校准,提升大语言模型的泛化能力,实验表明在密集和专家混合模型中提升了困惑度和零样本准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09316 2026-05-08 cs.LG cs.CL 91%

Large Language Model Prompt Datasets: An In-depth Analysis and Insights

大语言模型提示数据集:深入分析与洞察

Yuanming Zhang, Yan Lin, Arijit Khan, Huaiyu Wan

机构 * School of Computer Science and Technology(计算机科学与技术学院) Beijing Jiaotong University(北京交通大学) Department of Computer Science(计算机科学系) Aalborg University(奥尔堡大学)

专题命中 效率与部署 :LLM(summary_cn,abstract);large language model(title);language model(title);分类 cs.CL、cs.LG

AI总结 本文深入分析了129个异构LLM提示数据集,通过多层级语言分析揭示提示与一般文本的区别模式,并通过三个下游实验验证了提示过滤、领域分类和提示质量预测的实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20398 2026-04-23 cs.CL cs.LG cs.SE 91%

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning

WebGen-R1: 通过强化学习激励大语言模型生成功能性和美观的网站

Juyong Jiang, Chenglin Cai, Chansung Park, Jiasi Shen, Sunghun Kim, Jianguo Li, Yue Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Tongyi Lab, Alibaba Group(阿里云实验室) Electronics and Telecommunications Research Institute(电子电信研究院) Ant Group(蚂蚁集团)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.LG

AI总结 本文提出WebGen-R1框架,通过强化学习生成多页面功能性和美观的网站,解决了传统方法在多页面交互和审美评估上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19167 2026-04-22 cs.LG cs.AI 91%

LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation

LBLLM:通过三阶段蒸馏实现大语言模型的轻量二值化

Siqing Song, Chuang Wang, Yong Lang, Yi Yang, Xu-Yao Zhang

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Central Media Technology Institute, Huawei(华为中央媒体技术研究所)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 LBLLM通过三阶段蒸馏策略实现大语言模型的轻量二值化,采用W(1+1)A4量化方法,在单GPU上仅用0.016B tokens训练,超越现有二值化方法,在语言模型、常识问答和语言理解任务中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18141 2026-04-21 cs.CL cs.AI 91%

Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models

稀疏特征共激活揭示大语言模型中的因果语义模块

Ruixuan Deng, Xiaoyang Hu, Miles Gilberti, Shane Storks, Aman Taxali, Mike Angstadt, Chandra Sripada, Joyce Chai

机构 * Georgia Institute of Technology(佐治亚理工学院) Brown University(布朗大学) University of Michigan(密歇根大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 通过稀疏自编码器特征的共激活,研究发现大语言模型中存在语义连贯且上下文一致的网络组件,揭示了模型的模块化结构及高效操控方法。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13552 2026-04-16 cs.CL cs.AI 91%

Training-Free Test-Time Contrastive Learning for Large Language Models

无需训练的测试时对比学习用于大语言模型

Kaiwen Zheng, Kai Zhou, Jinwu Hu, Te Gu, Mingkai Peng, Fei Liu

机构 * South China University of Technology(南方科技大学) Pazhou Laboratory(琶洲实验室)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本文提出无需训练的测试时对比学习(TF-TTCL),通过动态'探索-反思-引导'循环提升冻结大语言模型的在线推理能力,优于零样本基线和代表性测试时适应方法。

Comments Accepted by Findings ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20112 2026-03-17 cs.CL cs.AI 91%

ERC-SVD: Error-Controlled SVD for Large Language Model Compression

ERC-SVD: 基于误差控制的大型语言模型压缩SVD

Haolei Bai, Siyong Jian, Tuo Liang, Yu Yin, Huan Wang

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);post-training(abstract)

AI总结 本文提出ERC-SVD方法,通过利用截断过程生成的残差矩阵减少截断损失,并选择性压缩模型最后一层以降低误差传播,从而提升压缩模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13765 2026-03-17 cs.CL cs.AI 91%

Knowledge Distillation for Large Language Models

为大型语言模型进行知识蒸馏

Alejandro Paredes La Torre, Barbara Flores, Diego Rodriguez

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);post-training(abstract);prompting(abstract)

AI总结 本文提出一种高效的压缩框架,结合引导的思维链强化学习,通过知识蒸馏和链式思维提示提升模型效率,实现更小规模的模型部署。

Comments Code and data are available at: https://github.com/AlejandroParedesLT/knowledge_distillLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07596 2026-02-10 cs.LG cs.AI 91%

Astro: Activation-guided Structured Regularization for Outlier-Robust LLM Post-Training Quantization

Astro: 基于激活引导的结构正则化用于鲁棒性LLM后训练量化

Xi Chen, Ming Li, Junxi Li, Changsheng Li, Peisong Wang, Lizhong Ding, Ye Yuan, Guoren Wang

机构 * Beijing Institute of Technology(北京理工大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Hebei Province Key Laboratory of Big Data Science and Intelligent Technology(河北省大数据科学与智能技术重点实验室)

专题命中 效率与部署 :LLM(title,abstract);post-training(title,abstract);large language model(abstract);language model(abstract)

AI总结 Astro通过激活引导的结构正则化方法,在高效且无延迟的情况下提升LLM后训练量化的鲁棒性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10157 2026-01-27 cs.CL cs.AI 91%

BILLY: Steering Large Language Models via Merging Persona Vectors for Creative Generation

BILLY:通过合并人设向量来引导大型语言模型进行创造性生成

Tsung-Min Pai, Jui-I Wang, Li-Chun Lu, Shao-Hua Sun, Hung-Yi Lee, Kai-Wei Chang

机构 * Department of Electrical Engineering, National Taiwan University(国立台湾大学电子工程系) Department of Computer Science & Information Engineering, National Taiwan University(国立台湾大学计算机科学与信息工程系) Graduate Institute of Communication Engineering, National Taiwan University(国立台湾大学通信工程研究所) CSAIL, Massachusetts Institute of Technology(麻省理工学院计算机科学与人工智能实验室)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

AI总结 BILLY通过合并人设向量在单个模型中实现多视角生成,提升创造力并降低计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05271 2026-01-12 cs.CL cs.LG 91%

Enhancing Foundation Models in Transaction Understanding with LLM-based Sentence Embeddings

利用基于大语言模型的句子嵌入增强交易理解的基础模型

Xiran Fan, Zhimeng Jiang, Chin-Chia Michael Yeh, Yuzhong Chen, Yingtong Dou, Menghai Pan, Yan Zheng

机构 * Visa Research(Visa研究)

专题命中 效率与部署 :LLM(title,abstract);foundation model(title,abstract);large language model(abstract);language model(abstract)

AI总结 本文提出一种结合大语言模型生成的句子嵌入与轻量级交易模型的混合框架,以提升交易理解任务的性能和效率。

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track (EMNLP 2025), pages 903-911

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07705 2025-12-09 cs.LG cs.AI 91%

In-Context and Few-Shots Learning for Forecasting Time Series Data based on Large Language Models

基于大语言模型的时序数据预测中的上下文学习与少样本学习

Saroj Gopali, Bipin Chhetri, Deepika Giri, Sima Siami-Namini, Akbar Siami Namin

机构 * Department of Computer Science(计算机科学系) Texas Tech University(德克萨斯理工大学) Advanced Academic Programs Science(高级学术计划科学) Johns Hopkins University(约翰霍普金斯大学) Cumberland University(坎布里奇大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);foundation model(abstract)

AI总结 本文研究了基于大语言模型的时序数据预测,探讨了上下文学习、零样本和少样本学习方法,发现TimesFM在预测性能上表现最佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02056 2025-12-05 cs.CL cs.AI 91%

Reversing Large Language Models for Efficient Training and Fine-Tuning

反向大型语言模型以实现高效的训练和微调

Eshed Gal, Moshe Eliasof, Javier Turek, Uri Ascher, Eran Treister, Eldad Haber

机构 * Department of Computer Science, University of British Columbia(计算机科学系,不列颠哥伦比亚大学) Department of Computer Science, Ben Gurion University(计算机科学系,本· Gurion大学) EarthDynamics AI Department of Earth, Ocean and Atmospheric Sciences, University of British Columbia(地球、海洋和大气科学系,不列颠哥伦比亚大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);foundation model(abstract)

AI总结 本研究提出可逆架构以高效训练和微调LLM,通过时间可逆动力学减少内存消耗,提升吞吐量,并通过微调将现有模型转换为可逆架构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24722 2025-11-07 cs.LG cs.AI 91%

HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts

Neil He, Rishabh Anand, Hiren Madhu, Ali Maatouk, Smita Krishnaswamy, Leandros Tassiulas, Menglin Yang, Rex Ying

机构 * Yale University, USA(耶鲁大学)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);pretraining(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19334 2025-10-23 stat.ML cs.AI cs.IR cs.LG 91%

Metadata Extraction Leveraging Large Language Models

Cuize Han, Sesh Jalagam

机构 * Box AI Platform(Box人工智能平台)

专题命中 效率与部署 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏