arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 11511 信号源:cs.CL, cs.AI, cs.LG

1. 指令微调 11511 篇

2406.13542 2024-07-19 cs.CL cs.AI cs.LG 91%

Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Guanting Dong, Keming Lu, Chengpeng Li, Tingyu Xia, Bowen Yu, Chang Zhou, Jingren Zhou

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);SFT(abstract);RLHF(abstract)

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17440 2024-05-29 cs.LG cs.AI cs.CL 91%

CataLM: Empowering Catalyst Design Through Large Language Models

Ludi Wang, Xueqing Chen, Yi Du, Yuanchun Zhou, Yang Gao, Wenjuan Cui

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02823 2024-04-04 cs.CL cs.AI cs.LG 91%

Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models

Haoran Sun, Lixin Liu, Junjie Li, Fengyu Wang, Baohua Dong, Ran Lin, Ruohui Huang

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);LLM(abstract);instruction tuning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12255 2024-04-04 cs.CR cs.AI cs.CL cs.LG 91%

Instructional Fingerprinting of Large Language Models

Jiashu Xu, Fei Wang, Mingyu Derek Ma, Pang Wei Koh, Chaowei Xiao, Muhao Chen

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);LLM(abstract);instruction tuning(abstract)

Comments Accepted at NAACL 2024; 30 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13995 2024-02-27 cs.CL cs.AI cs.IR cs.LG 91%

On Bilingual Lexicon Induction with Large Language Models

Yaoyiran Li, Anna Korhonen, Ivan Vulić

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)

Comments EMNLP 2023 Main Conference

Journal ref Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 9577-9599

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16405 2024-02-05 cs.CL cs.AI cs.LG 91%

Scaling Sparse Fine-Tuning to Large Language Models

Alan Ansell, Ivan Vulić, Hannah Sterz, Anna Korhonen, Edoardo M. Ponti

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);LLM(abstract);SFT(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08577 2024-01-17 cs.CV cs.AI cs.CL cs.LG cs.RO 91%

MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World

Yining Hong, Zishuo Zheng, Peihao Chen, Yian Wang, Junyan Li, Chuang Gan

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);LLM(abstract);instruction tuning(abstract)

Comments Project page: https://vis-www.cs.umass.edu/multiply

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14936 2023-07-28 cs.CL cs.AI cs.LG cs.PL cs.SE 91%

PanGu-Coder2: Boosting Large Language Models for Code with Ranking Feedback

Bo Shen, Jiaxin Zhang, Taihong Chen, Daoguang Zan, Bing Geng, An Fu, Muhan Zeng, Ailun Yu, Jichuan Ji, Jingyang Zhao, Yuenan Guo, Qianxiang Wang

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);LLM(abstract);instruction tuning(abstract)

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17545 2024-07-26 cs.SE cs.AI cs.CL 91%

Large Language Models for Anomaly Detection in Computational Workflows: from Supervised Fine-Tuning to In-Context Learning

Hongwei Jin, George Papadimitriou, Krishnan Raghavan, Pawel Zuk, Prasanna Balaprakash, Cong Wang, Anirban Mandal, Ewa Deelman

专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);SFT(abstract);prompting(abstract)

Comments 12 pages, 14 figures, paper is accepted by SC'24, source code, see: https://github.com/PoSeiDon-Workflows/LLM_AD

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14150 2026-08-17 cs.CL 新提交 91%

Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge

面向第二届MLC-SLM挑战赛的前置静音增强与多阶段合成监督

Kexin Shi, Renhe Sun, Yuge Huang, Ximeng Wang, Jiayi Zhou, Jian Liu, Malu Zhang

机构 * Ant Group(蚂蚁集团) UESTC(电子科技大学)

专题命中 指令微调 :SLM(title,title_cn);language model(abstract);分类 cs.CL

AI总结 针对第二届MLC-SLM挑战赛的两项任务,分别采用前置静音增强等策略微调VibeVoice-ASR-7B、用合成问答对微调Qwen3-Omni-30B-A3B-Instruct,提升了任务1和任务2的评估性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05156 2026-08-07 cs.CL 新提交 91%

Scaffold-Mediated Post-Training: Co-Evolving Model Parameters and Procedural Scaffold Graphs

支架介导的后训练:协同演化模型参数与程序性支架图

Fei Ding, Yongkang Zhang, Runhao Liu, Yuhao Liao, Zijian Zeng, Huiming Yang

机构 * Alibaba Group(阿里巴巴集团) Tsinghua University(清华大学)

专题命中 指令微调 :SFT(summary_cn,abstract);post-training(title,abstract);large language model(abstract);language model(abstract)

AI总结 该研究针对大型语言模型后训练中参数与推理支架脱节的问题,提出支架介导的后训练范式,通过协同演化实现技能提升,在FeatureBench上较标准SFT取得显著性能优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04697 2026-08-06 cs.AI 新提交 91%

Traceable LLM-Generated Hazard Scenarios for Operational Safety Analysis of Aviation Systems Using ASRS Reports

基于ASRS报告的可追踪LLM生成航空系统运行安全分析危险场景

Cristian Mascia, Roberto Pietrantuono, Daniel Rodriguez, Stefano Russo

专题命中 指令微调 :LLM(title,title_cn);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 该研究针对航空系统运行安全分析,提出AI辅助方法结合ASRS报告,用LLM生成可追踪危险场景,通过进化溯因优化混合变体,经评估验证了模型与提示对生成场景有效性的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27443 2026-07-31 cs.AI 新提交 91%

Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems

利用轨迹图实现智能体大语言模型系统的预执行错误诊断

Xu Zheng, Zhuomin Chen, Chaohao Lin, Hua Wei, Haifeng Chen, Wei Cheng, Dongsheng Luo

机构 * Florida International University(佛罗里达国际大学) Arizona State University(亚利桑那州立大学) NEC Laboratories America(美国 NEC 实验室) Singapore Management University(新加坡管理大学)

专题命中 指令微调 :LLM(title,summary_cn);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 该研究提出Trajectory Graph Copilot框架,以Graph Debugger为核心,通过预执行错误诊断提升LLM智能体完成长周期任务的能力,在多基准测试中平均提升14.69%通过率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22651 2026-07-28 cs.AI 新提交 91%

ARdena: Scenario-driven control of real-time LLM agents

ARdena:实时大语言模型智能体的场景驱动控制

Luka Borozan, Domagoj Matijević

机构 * Faculty of Applied Mathematics and Informatics, University of Osijek Croatia(奥西耶克大学应用数学与信息学院,克罗地亚) Meet Intelligent Innovations LLC(Meet智能创新有限责任公司)

专题命中 指令微调 :LLM(title,summary_cn);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 研究如何在实时交互环境中控制大语言模型智能体行为,提出分层场景驱动的LLM控制框架,通过结构化提示实现运行时行为控制,在ARDena中实现该框架并评估,结果显示场景驱动提示能有效控制LLM智能体。

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.19523 2026-07-23 cs.CL 新提交 91%

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play

当推理缩小行动范围:大语言模型游戏中的多样性崩溃

Junyi Sha, Renfei Tan, David Simchi-Levi

机构 * Institute for Data, Systems, and Society(数据、系统与社会研究所)

专题命中 指令微调 :SFT(summary_cn,abstract);LLM(title,abstract);large language model(abstract);language model(abstract)

AI总结 研究大语言模型游戏中监督微调对行为多样性的影响,发现推理模式生成常抑制行动多样性,标准SFT会致过早多样性崩溃,行动增强可部分缓解,指出窄支持模仿是策略崩溃源,SFT中保持行动支持对维持探索行为很重要。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24434 2026-07-02 cs.CL 版本更新 91%

LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data

LuxIT:从单语种子数据生成的卢森堡语指令微调数据集

Julian Valline, Cedric Lothritz, Siwen Guo, Jordi Cabot

机构 * Luxembourg Institute of Science and Technology(卢森堡科学技术研究院)

专题命中 指令微调 :LLM(summary_cn,abstract);instruction tuning(title,abstract);large language model(abstract);language model(abstract)

AI总结 本文提出LuxIT数据集,通过合成高质量卢森堡语指令-回答对,提升低资源语言LLM性能,实验显示在语言考试和NLP任务中均取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16517 2026-07-01 cs.LG q-bio.QM 新提交 91%

How Post-Training Shapes Biological Reasoning Models

后训练如何塑造生物学推理模型

Lukas Fesser, Hanlin Zhang, Michelle M. Li, Eric Wang, Bryan Perozzi, Shekoofeh Azizi, Sham M. Kakade, Marinka Zitnik

机构 * Harvard University(哈佛大学) Google DeepMind(谷歌DeepMind) Google Research(谷歌研究院)

专题命中 指令微调 :SFT(summary_cn,abstract);post-training(title,abstract);language model(abstract);foundation model(abstract)

AI总结 研究后训练各阶段(CPT、SFT、RL)对生物学推理模型领域内和领域外性能的影响,发现SFT提升领域内性能但损害泛化,RL可部分恢复泛化,最佳策略是短SFT加长RL。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22942 2026-06-23 cs.CL 新提交 91%

Understanding Knowledge Distillation in Post-Training: When It Helps and When It Fails

理解后训练中的知识蒸馏:何时有效与何时失败

Xin Liu, Simin Ma, Shujian Liu, Song Wang, Sathish Reddy Indurthi, Haoyun Deng, Lu Wang, Kaiqiang Song

机构 * University of Michigan(密歇根大学) Zoom Video Communications(Zoom视频通信公司)

专题命中 指令微调 :SFT(summary_cn,abstract);post-training(title,abstract);large language model(abstract);language model(abstract)

AI总结 本文系统研究后训练阶段的知识蒸馏,发现低数据场景下KD优于SFT,但数据充足时优势减弱;从更强教师蒸馏可恢复增益,并提出两阶段KD策略提升数据稀缺环境下的学生模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11786 2026-06-11 cs.CL 新提交 91%

Lius: Translation Model Based Instructional Lingustic Using Continual Instruction Tuning In Kupang Malay

Lius:基于持续指令调优的库邦马来语教学语言学翻译模型

Joanito Agili Lopo, Yunita Sari, Guntur Budi Herwanto

机构 * Universitas Gadjah Mada(加札马达大学)

专题命中 指令微调 :LLM(summary_cn,abstract);instruction tuning(title,abstract);large language model(abstract);language model(abstract)

AI总结 针对低资源语言库邦马来语,提出利用双语词典的词汇和语义特征设计指令,并采用持续指令调优(CIT)范式微调大语言模型,在多个指标上超越基线4-6分,优于NMT和多语言LLM 10-13分。

Comments This paper is the result of the Master Thesis in Master of Artificial Intelligence at Universitas Gadjah Mada

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29427 2026-05-29 cs.CL 91%

FinGuard: Detecting Financial Regulatory Non-Compliance in LLM Interactions

FinGuard:检测LLM交互中的金融监管违规

Huaixia Dou, Jie Zhu, Minghao Wu, Shuo Jiang, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang

机构 * Qwen DianJin Team, Alibaba Cloud Computing(阿里云计算Qwen金融团队) Tongyi Lab, Alibaba Group(阿里集团通义实验室) School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院)

专题命中 指令微调 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 针对金融领域LLM交互中的监管违规检测问题,提出基于监管文档的自动化管道,构建首个金融合规检测基准FinGuard-Bench,并训练FinGuard模型,在基准上显著优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28802 2026-05-28 cs.CL 91%

Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization

人类标注变异作为稳定信号:通过跨标注者偏好优化学习标注者特定的解释行为

Beiduo Chen, Pingjun Hong, Ziyun Zhang, Benjamin Roth, Anna Korhonen, Barbara Plank

机构 * MaiNLP Center for Information and Language Processing(信息与语言处理中心) LMU Munich(慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) University of Vienna(维也纳大学) LTL University of Cambridge(剑桥大学)

专题命中 指令微调 :preference optimization(title,abstract);SFT(abstract,abstract_cn);LLM(abstract_cn);large language model(abstract)

AI总结 研究大语言模型能否学习并复现标注者特定的标签-解释行为,提出跨标注者偏好优化(CAPO)方法,通过对比目标标注者与其他有效但非目标标注者的响应来提升模仿和归因能力。

Comments 43 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22166 2026-05-28 cs.AI 91%

Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents

适配接口而非模型:面向确定性LLM智能体的运行时框架适配

Tianshi Xu, Huifeng Wen, Meng Li

机构 * Peking University(北京大学)

专题命中 指令微调 :LLM(title,title_cn);language model(abstract);分类 cs.AI

AI总结 提出Life-Harness运行时框架,通过从训练轨迹中演化出可复用的环境侧干预,在不修改模型权重或评估环境的情况下,显著提升冻结LLM智能体在确定性任务中的性能。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02709 2026-05-22 cs.AI 91%

ATLAS: A Multi-LLM Training Framework for EvoDPO with Adaptive Reference Evolution

ATLAS:一种用于EvoDPO的多LLM训练框架,具有自适应参考进化

Ujin Jeon, Jiyong Kwon, Madison Ann Sullivan, Caleb Eunho Lee, Guang Lin

机构 * School of Electrical and Computer Engineering(电气与计算机工程学院) Purdue University West Lafayette(韦伯州立大学) School of Mechanical Engineering(机械工程学院) Department of Mathematics(数学系) Department of Computer Science(计算机科学系) Department of Mathematics and Mechanical Engineering(数学与机械工程系)

专题命中 指令微调 :LLM(title,title_cn);preference optimization(abstract);分类 cs.AI

AI总结 本文提出ATLAS框架,通过自适应参考进化解决多LLM代理系统中固定参考模型导致的更新保守或训练停滞问题,结合支持者驱动探索与EvoDPO驱动的稳定性,提升长期评估驱动的自我改进能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07239 2026-05-19 cs.CL 91%

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts

Red-Bandit:通过带引导的LoRA专家实现LLM红队测试的测试时适应

Christos Ziakas, Nicholas Loo, Nishita Jain, Alessandra Russo

机构 * Imperial College London(伦敦帝国学院)

专题命中 指令微调 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出Red-Bandit框架,通过带引导的LoRA专家在不同攻击风格下实现LLM的测试时适应,通过强化学习生成不安全提示,并利用多臂老虎机策略动态选择攻击风格专家,从而在AdvBench上取得最佳结果,同时生成更易读的提示。

Comments Accepted to the Main Conference at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15393 2026-05-18 cs.LG 91%

LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling

LPDS:通过逻辑保持难度扩展评估LLM鲁棒性

Philipp Mondorf, Samuel J. Bell, Jesse Dodge, Dieuwke Hupkes

机构 * FAIR at Meta(Meta 的 FAIR 部门) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 指令微调 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出LPDS框架,通过系统搜索逻辑保持变化来评估LLM鲁棒性,发现难度增加导致性能下降,且微调困难变体能获得更一致的鲁棒性提升。

Comments 41 pages, 31 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11907 2026-05-15 cs.LG 91%

Procedural-skill SFT across capacity tiers: A W-Shaped pre-SFT Trajectory and Regime-Asymmetric Mechanism on 0.8B-4B Qwen3.5 Models

跨容量层级的程序技能SFT:一种W形的预SFT轨迹和 regime-不对称机制在0.8B-4B Qwen3.5模型上

Igor Strozzi

机构 * Applied Mathematics Department(应用数学系)

专题命中 指令微调 :SFT(title,title_cn);LLM(abstract);分类 cs.LG

AI总结 本文研究了不同规模Qwen3.5模型在程序技能SFT上的贡献,发现预SFT轨迹呈W形,且SFT效果在不同模型容量上呈现不对称性,通过实验验证了这一机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13411 2026-05-14 cs.CR cs.CL 91%

Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution

基于外部化攻击-防御共演的模型无关持续LLM安全

Xiaozhe Zhang, Chaozhuo Li, Hui Liu, Shaocheng Yan, Bingyu Yan, Qiwei Ye, Haoliang Li

机构 * City University of Hong Kong(香港城市大学) Beijing University of Posts and Telecommunications(北京邮电大学) Wuhan University(武汉大学) Beihang University(北京航空航天大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 指令微调 :LLM(title,title_cn);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 本文提出EvoSafety框架,通过外部结构实现持续安全改进,采用对抗技能库和轻量防御模型,在单一训练中实现攻击探测与防御提升,实验显示其在Guard模式下防御成功率高达99.61%。

Comments 48 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03412 2026-04-30 cs.CL 91%

Verified Critical Step Optimization for LLM Agents

用于大语言模型代理的验证关键步骤优化

Mukai Li, Qingcheng Zeng, Tianqing Fang, Zhenwen Liang, Linfeng Song, Qi Liu, Haitao Mi, Dong Yu

机构 * Tencent AI Lab(腾讯AI实验室) The University of Hong Kong(香港大学) Northwestern University(西北大学)

专题命中 指令微调 :SFT(summary_cn,abstract);LLM(title);large language model(abstract);language model(abstract)

AI总结 本文提出CSO方法,通过验证关键步骤进行强化学习,以提升代理在复杂任务中的表现,实验显示其在GAIA-Text-103和XBench-DeepSearch上优于SFT基线。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20835 2026-04-23 cs.CL 91%

Parallel-SFT: Improving Zero-Shot Cross-Programming-Language Transfer for Code RL

并行SFT:改进代码RL的零样本跨编程语言转移

Zhaofeng Wu, Shiqi Wang, Boya Peng, Anuj Goyal, Melanie Kambadur, Sebastian Ruder, Yoon Kim, Chloe Bi

机构 * Meta Superintelligence Labs(Meta超智能实验室) MIT(麻省理工学院)

专题命中 指令微调 :SFT(title,title_cn);language model(abstract);分类 cs.CL

AI总结 本文提出零样本跨编程语言转移任务,通过并行SFT策略提升代码RL在不同语言间的迁移能力,发现通用化SFT初始化能提升模型泛化性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17210 2026-04-21 cs.LG 91%

Guardrails in Logit Space: Safety Token Regularization for LLM Alignment

对数空间中的Guardrails:用于LLM对齐的安全令牌正则化

Thong Bach, Truyen Tran

机构 * Applied Artificial Intelligence Initiative (A2I2)(应用人工智能倡议(A2I2)) Deakin University(德克萨斯大学)

专题命中 指令微调 :LLM(title,title_cn);large language model(abstract);language model(abstract);pretraining(abstract)

AI总结 本文提出安全令牌正则化(STR)方法,通过约束关键令牌的logits来保持细调模型的安全性,实验显示其在安全性和任务性能方面表现优异。

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏