arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 11547 信号源:cs.CL, cs.AI, cs.LG

1. 指令微调 11547 篇

2305.18752 2023-05-31 cs.CV cs.CL 89%

GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction

Rui Yang, Lin Song, Yanwei Li, Sijie Zhao, Yixiao Ge, Xiu Li, Ying Shan

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18153 2023-05-31 cs.CL 89%

Do Large Language Models Know What They Don't Know?

Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, Xuanjing Huang

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);instruction tuning(abstract);分类 cs.CL

Comments 10 pages, 9 figures, accepted by Findings of ACL2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11159 2023-05-19 cs.CL 89%

Aligning Instruction Tasks Unlocks Large Language Models as Zero-Shot Relation Extractors

Kai Zhang, Bernal Jiménez Gutiérrez, Yu Su

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL

Comments ACL 2023 Findings; The code is available at https://github.com/OSU-NLP-Group/QA4RE

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07001 2023-05-12 cs.IR cs.CL 89%

Recommendation as Instruction Following: A Large Language Model Empowered Recommendation Approach

Junjie Zhang, Ruobing Xie, Yupeng Hou, Wayne Xin Zhao, Leyu Lin, Ji-Rong Wen

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.15042 2023-05-02 cs.CR cs.CL 89%

Privately Fine-Tuning Large Language Models with Differential Privacy

Rouzbeh Behnia, Mohamamdreza Ebrahimi, Jason Pacheco, Balaji Padmanabhan

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL

Comments Publised at IEEE ICDM Workshop on Machine Learning for Cybersecurity (MLC) 2022

Journal ref 2022 IEEE International Conference on Data Mining Workshops (ICDMW), pp. 560-566

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08109 2023-04-19 cs.CL 89%

A Comparative Study between Full-Parameter and LoRA-based Fine-Tuning on Chinese Instruction Data for Instruction Following Large Language Model

Xianghui Sun, Yunjie Ji, Baochang Ma, Xiangang Li

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);instruction tuning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.11140 2022-12-22 cs.PL cs.LG cs.SE 89%

Benchmarking Large Language Models for Automated Verilog RTL Code Generation

Shailja Thakur, Baleegh Ahmad, Zhenxing Fan, Hammond Pearce, Benjamin Tan, Ramesh Karri, Brendan Dolan-Gavitt, Siddharth Garg

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.LG

Comments Accepted in DATE 2023. 7 pages, 4 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03767 2022-05-12 cs.CL 89%

Context-Aware Abbreviation Expansion Using Large Language Models

Shanqing Cai, Subhashini Venugopalan, Katrin Tomanek, Ajit Narayanan, Meredith Ringel Morris, Michael P. Brenner

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL

Comments 15 pages, 7 figures, 8 tables. Accepted as a long paper at NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.03910 2022-04-01 cs.CL 89%

A Recipe For Arbitrary Text Style Transfer with Large Language Models

Emily Reif, Daphne Ippolito, Ann Yuan, Andy Coenen, Chris Callison-Burch, Jason Wei

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11919 2025-03-07 cs.CV 89%

LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Models for Referring Expression Comprehension

Amaia Cardiel, Eloi Zablocki, Elias Ramzi, Oriane Siméoni, Matthieu Cord

专题命中 指令微调 :LLM(title,abstract);language model(title,abstract);large language model(abstract)

Comments LLM-wrapper (v3) is published as a conference paper at ICLR 2025. (v1 was presented at EVAL-FoMo workshop, ECCV 2024.)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22067 2026-08-18 cs.CL cs.AI 版本更新 89%

Multimodal Language Models Benchmarked Against the NRC Reactor Operator Licensing Examination: Fine-Tuning and Retrieval Strategies

基于美国核管理委员会反应堆操作员执照考试的多模态语言模型微调与检索策略基准测试

Isak Hwang, Yoon Pyo Lee, Syed Bahauddin Alam

机构 * organization= Department of Nuclear Engineering, Hanyang University , addressline= 222 Wangsimni-ro , postcode= 04763 , state= Seongdong-gu , city= Seoul , country= South Korea organization= The Grainger College of Engineering, Nuclear, Plasma \& Radiological Engineering, University of Illinois Urbana-Champaign , city= Urbana , state= IL , country= USA

专题命中 指令微调 :SFT(summary_cn,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 该研究针对美国核管理委员会反应堆操作员执照考试,评估310亿参数多模态模型应用核知识的能力,通过对比基础模型与多种微调及检索配置,发现固定大小分块RAG的SFT配置表现最佳,并揭示了分块策略规律及RAFT与SFT的性能差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11715 2026-08-13 cs.CL cs.AI 新提交 89%

When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use

当API“说错”语言:重新审视多语言工具使用的后训练

Siddharth Chauhan, Thomas Butler, Abhishek Singhania, Pankaj Porwal, Honey Gupta

机构 * Amazon(亚马逊)

专题命中 指令微调 :post-training(title,abstract);SFT(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 针对多语言API调用中存在的参数语言不匹配问题,研究发现监督微调可实现接近或优于复杂强化学习方法的性能,强化学习仅能提供渐进式改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18960 2026-07-22 cs.LG cs.AI cs.CR 新提交 89%

SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement

SFGA:一种用于可信SFT数据采购的具有裁决升级的统计优先门控架构

Arther Tian, Alex Ding, Simon Wu, Aaron Chan

机构 * DGrid AI(DGrid人工智能)

专题命中 指令微调 :SFT(title,title_cn);LLM(abstract);分类 cs.AI、cs.LG

AI总结 研究如何采购监督微调数据,提出统计优先门控架构SFGA,将采购视为成本感知路由问题,经实验在准确率、F1值和成本上取得良好平衡,还报告辩论路径负面诊断,为测量和校准构建合成基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30077 2026-06-30 cs.LG cs.AI 89%

Online Data Selection for Instruction Tuning via Gaussian Processes

基于高斯过程的指令调优在线数据选择

Jun Wang, Quoc Phong Nguyen, Julien Monteil, Vu Nguyen

机构 * Amazon(亚马逊)

专题命中 指令微调 :instruction tuning(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 提出GAIA框架,利用高斯过程回归建模语义空间中的连续效用流形,通过自适应策略融合动态选择高质量样本,在三个数据集上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07528 2026-06-23 cs.CL cs.AI 版本更新 89%

From RAG to Agentic RAG for Faithful Islamic Question Answering

从RAG到智能体RAG:面向可靠的伊斯兰问答

Gagan Bhatia, Hamdy Mubarak, Mustafa Jarrar, George Mikros, Fadi Zaraket, Mahmoud Alhirthani, Mutaz Al-Khatib, Logan Cochrane, Kareem Darwish, Rashid Yahiaoui, Firoj Alam

机构 * Qatar Computing Research Institute, HBKU, Qatar(卡塔尔计算研究中心,HBKU,卡塔尔) College of Humanities and Social Sciences, HBKU, Qatar(人文与社会科学学院,HBKU,卡塔尔) Arab Center for Research and Policy Studies, Qatar(阿拉伯研究中心与政策研究所,卡塔尔) College of Islamic Studies, HBKU, Qatar(伊斯兰研究学院,HBKU,卡塔尔) College of Public Policy, HBKU, Qatar(公共政策学院,HBKU,卡塔尔)

专题命中 指令微调 :LLM(summary_cn,abstract_cn);SFT(abstract,abstract_cn);large language model(abstract,comments);language model(abstract,comments)

AI总结 针对LLM在伊斯兰问答中的幻觉与弃权问题,构建双语基准IslamicFaithQA,并提出基于结构化工具调用的智能体RAG框架,显著提升正确性与鲁棒性。

Comments Islamic Question Answering; Faithful Question Answering; Retrieval-Augmented Generation; Agentic RAG; Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16074 2026-06-16 cs.CL cs.AI 新提交 89%

PVminerLLM2: Improving Structured Extraction of Patient Voice via Preference Optimization

PVminerLLM2:通过偏好优化改进患者声音的结构化提取

Samah Fodeh, Linhai Ma, Ganesh Puthiaraju, Srivani Talakokkul, Afshan Khan, Elyas Irankhah, Sreeraj Ramachandran, Ashley Hagaman, Sarah Lowe, Aimee Roundtree

机构 * Yale School of Medicine(耶鲁大学医学院) Yale School of Public Health(耶鲁大学公共卫生学院) Texas State University(德克萨斯州立大学)

专题命中 指令微调 :preference optimization(title,abstract);LLM(abstract,abstract_cn);SFT(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 提出PVminerLLM2,通过偏好优化和令牌级门控稳定项、混淆感知偏好对构建等技术,解决监督微调难以处理的细粒度错误,在患者声音结构化提取任务上优于基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12426 2026-06-12 cs.CY cs.CL cs.LG 新提交 89%

Two Wrongs, No Right: Auditing Social-Desirability Bias in LLM Annotators for Computational Social Science

两个错误,没有正确:审计计算社会科学中LLM标注者的社会期望偏差

Varun Kotte

机构 * Varun Kotte

专题命中 指令微调 :LLM(title,title_cn);prompting(abstract);分类 cs.CL、cs.LG

AI总结 研究审计了三个开源指令微调模型在TweetEval任务中的社会期望偏差,发现模型存在宽大、过度纠正和中性偏差,且提示干预无法纠正,聚合指标可能掩盖实质结论错误。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03685 2026-06-03 cs.LG cs.AI 89%

A Close Look At World Model Recovery In Supervised Fine-Tuned LLM Planners

监督微调的大语言模型规划器中世界模型恢复的深入探究

Patrick Emami, Nan Qiang, Peter Graf

机构 * National Laboratory of the Rockies(落基山国家实验室)

专题命中 指令微调 :LLM(title,abstract_cn);SFT(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 通过可解释性实验,研究监督微调如何影响大语言模型在经典规划任务中恢复世界模型的能力,发现微调使模型线性编码动作有效性和状态谓词,且更广泛的状态空间覆盖有助于更准确的世界模型恢复。

Comments 17 pages. Under review at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10765 2026-05-12 cs.CV cs.AI cs.LG 89%

Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning

动态跨模态提示生成用于多模态持续指令微调

Tao Hu, Da-Wei Zhou

机构 * School of Artificial Intelligence, Nanjing University(南京大学人工智能学院) State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室)

专题命中 指令微调 :instruction tuning(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文提出DRAPE框架,通过生成连续实例特定的软提示提升多模态持续指令微调性能,采用跨模态注意力和投影梯度投影减少遗忘,实验显示优于现有基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23750 2026-05-12 cs.LG cs.AI 89%

The Override Gap: A Magnitude Account of Knowledge Conflict Failure in Hypernetwork-Based Instant LLM Adaptation

覆盖差距:基于超网络的即时LLM适应中知识冲突失败的幅度解释

Shuaizhi Cheng, Xiang Shi, Zhiwei Zhang, Mingwei Li

机构 * Harbin Institute of Technology(哈尔滨工程大学) Imperial College London(伦敦帝国理工学院) KigLand Machine Learning Lab(KigLand机器学习实验室)

专题命中 指令微调 :LLM(title,title_cn);pretraining(abstract);分类 cs.AI、cs.LG

AI总结 研究揭示超网络方法在知识冲突中的失败源于幅度问题,通过选择性层提升和意识冲突内部化技术,显著提升深度冲突准确性,同时保持新知识召回。

Comments 35 pages, 15 figures v2: minor layout fixes and author list update

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08368 2026-05-12 cs.AI cond-mat.stat-mech cs.LG 89%

On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective

在训练后区分能力激发与能力创造:从自由能视角出发

Yuhao Li, Shengchao Liu

机构 * Department of Computer Science and Engineering(计算机科学与工程系)

专题命中 指令微调 :post-training(title,abstract);SFT(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文从自由能视角出发,区分训练后的能力激发与能力创造,通过引入可及支持的概念,探讨训练过程对行为概率和可达行为空间的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06654 2026-05-08 cs.LG cs.AI math.OC 89%

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less

优化器-模型一致性:使用与预训练相同的优化器进行全微调可减少遗忘

Yuxing Liu, Jianyu Wang, Tong Zhang

机构 * UIUC(伊利诺伊大学香槟分校) Apple(苹果公司)

专题命中 指令微调 :pretraining(title,abstract);SFT(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文发现使用与预训练相同的优化器进行全微调,在监督微调阶段能更少遗忘并保持性能,提出优化器-模型一致性概念,通过实验和理论分析揭示优化器对模型的影响及微调策略的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00650 2026-05-04 cs.LG cs.AI 89%

AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments

AdaMeZO:一种用于LLM微调的Adam风格零阶优化器,无需维护动量

Zhijie Cai, Haolong Chen, Guangxu Zhu

机构 * Shenzhen Research Institute of Big Data(深圳大数据研究院) The Chinese University of Hong Kong-Shenzhen(香港中文大学(深圳)) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 指令微调 :LLM(title,title_cn);分类 cs.AI、cs.LG

AI总结 本文提出AdaMeZO,一种基于Adam风格的零阶优化器,通过估计一阶和二阶动量而不存储在内存中,有效减少GPU内存需求,同时在微调LLM时表现出色,比MeZO快70%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08275 2026-02-16 cs.CL cs.AI 89%

MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis

MLLM-CTBench: 一个用于持续指令微调的基准,包含推理过程诊断

Haiyun Guo, Zhiyan Hou, Yandu Sun, Jinghan He, Yu Chen, Yuzhe Zhou, Yuheng Jia, Jinqiao Wang, Tat-Seng Chua

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Southeast University(东南大学) School of Computing, National University of Singapore(新加坡国立大学计算机学院)

专题命中 指令微调 :instruction tuning(title,abstract);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 MLLM-CTBench提出一个用于持续指令微调的基准,通过多维评估框架和强化微调方法,分析跨任务知识保留和灾难性遗忘问题。

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08239 2026-02-10 cs.LG cs.AI 89%

Linearization Explains Fine-Tuning in Large Language Models

线性化解释大语言模型中的微调

Zahra Rahimi Afzal, Tara Esmaeilbeig, Mojtaba Soltanalian, Mesrob I. Ohannessian

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文通过线性化视角揭示了大语言模型微调的机制,分析了正则化对NTK特征值谱的影响,并通过LoRA实验证明了微调优化的改进。

Journal ref Afzal, Z.R., Esmaeilbeig, T., Soltanalian, M. and Ohannessian, M.I., 2025. Linearization Explains Fine-Tuning in Large Language Models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05032 2026-02-03 cs.CL cs.AI 89%

Enhancing Human-Like Responses in Large Language Models

增强大型语言模型的人类样响应

Ethem Yağız Çalık, Talha Rüzgar Akkuş

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);foundation model(comments,journal_ref);分类 cs.CL、cs.AI

AI总结 本文提出通过提升自然语言理解、对话连贯性和情感智能来增强大型语言模型的人类化响应,展示了改进用户交互和跨领域应用的潜力。

Comments Presented at the AAAI-26 Workshop on Personalization in the Era of Large Foundation Models (PerFM), Singapore, January 2026

Journal ref Presented at the AAAI-26 Workshop on Personalization in the Era of Large Foundation Models (PerFM), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11344 2026-01-19 cs.CL cs.AI 89%

How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting

临床医生会修改这个草稿多少?评估LLM对患者信息回复草稿的对齐情况

Parker Seegmiller, Joseph Gatto, Sarah E. Greer, Ganza Belise Isingizwe, Rohan Ray, Timothy E. Burdick, Sarah Masud Preum

机构 * Department of Computer Science, Dartmouth College(计算机科学系,达特茅斯学院) Department of Community and Family Medicine, Dartmouth Health(社区与家庭医学系,达特茅斯健康) The Dartmouth Institute, Dartmouth College(达特茅斯研究所,达特茅斯学院)

专题命中 指令微调 :LLM(title,abstract);large language model(abstract);language model(abstract);preference optimization(abstract)

AI总结 本研究评估了LLM在患者信息回复起草任务中的对齐情况,提出了一种新的评估框架,并发现需要根据临床医生的偏好进行适应以提高可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09438 2025-11-13 cs.LG cs.AI 89%

LLM-Guided Dynamic-UMAP for Personalized Federated Graph Learning

Sai Puppala, Ismail Hossain, Md Jahangir Alam, Tanzim Ahad, Sajedul Talukder

机构 * University of Texas at El Paso(德克萨斯理工大学) Southern Illinois University Carbondale(南方伊利诺伊大学卡本代尔分校)

专题命中 指令微调 :LLM(title,abstract);large language model(abstract);language model(abstract);foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10717 2025-08-27 cs.CL cs.AI 89%

A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment

Jean-Philippe Corbeil, Amin Dada, Jean-Michel Attendu, Asma Ben Abacha, Alessandro Sordoni, Lucas Caccia, François Beaulieu, Thomas Lin, Jens Kleesiek, Paul Vozila

专题命中 指令微调 :instruction tuning(title,abstract);large language model(abstract);language model(abstract);small language model(abstract)

Journal ref ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05945 2025-08-26 cs.CL cs.AI 89%

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models

Paul Darm, Annalisa Riccardi

机构 * University of Strathclyde(斯特拉思克莱德大学)

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

Comments Published at Transaction of Machine Learning Research 08/2025, Large Language Models (LLMs), Interference-time activation shifting, Steerability, Explainability, AI alignment, Interpretability

详情

展开后加载摘要…

URL PDF HTML 收藏