When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff
当强化学习在监督微调后失效:恢复模型可塑性以实现稳健的SFT到RL交接
Runze Liu, Jiashun Liu, Xu Wan, Yuqian Fu, Ling Pan
机构
*
Hong Kong University of Science and Technology(香港科技大学)
;
Zhejiang University(浙江大学)
;
State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(多模态人工智能系统国家重点实验室,CASIA)
专题命中
指令微调
:SFT(title,title_cn);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)
机构
*
Mila - Quebec AI Institute(魁北克人工智能研究所)
;
Polytechnique Montréal(蒙特利尔理工学院)
;
Université de Montréal(蒙特利尔大学)
;
McGill University(麦吉尔大学)
;
CIFAR AI Chair(CIFAR人工智能主席)
;
Google DeepMind(谷歌DeepMind)
专题命中
指令微调
:SFT(title,title_cn);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)
机构
*
School of Electronic, Electrical and Physics, Fujian University of Technology(福建工程学院电子、电气与物理系)
;
School of Humanities, Fujian University of Technology(福建工程学院人文学院)
;
Department of Investigation, Fujian Police College(福建警察学院调查系)
;
Institute of Applied Physics and Materials Engineering, University of Macau(澳门大学应用物理与材料工程研究院)
专题命中
指令微调
:LLM(title,title_cn);SFT(abstract,abstract_cn);RLHF(abstract,abstract_cn);large language model(abstract)
机构
*
McCormick School of Engineering, Northwestern University(西北大学工程学院)
;
Wenzhou Buyi Pharmacy Chain Co., Ltd.(温州-buyi药链有限公司)
;
College of Computer Science and Artificial Intelligence, Wenzhou University(温州大学计算机科学与人工智能学院)
;
Department of Decision Analytics and Operations, City University of Hong Kong(香港城市大学决策分析与运营部门)
;
Institute of Operations Research and Analytics, National University of Singapore(新加坡国立大学运筹学与分析研究所)
专题命中
指令微调
:LLM(title,title_cn);SFT(abstract,abstract_cn);large language model(abstract);language model(abstract)
Property Enhanced Instruction Tuning for Multi-task Molecule Generation with Large Language Models
属性增强指令微调用于大型语言模型的多任务分子生成
Xuan Lin, Long Chen, Yile Wang, Yangyang Chen, Xiangxiang Zeng
机构
*
School of Computer Science, Xiangtan University(湘潭大学计算机科学学院)
;
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
;
Department of Computer Science, University of Tsukuba(东京大学理工学部)
;
College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院)
专题命中
指令微调
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);instruction tuning(title)
ReasonXL: Shifting LLM Reasoning Language Without Sacrificing Performance
ReasonXL: 不牺牲性能的LLM推理语言迁移
Daniil Gurgurov, Tom Röhr, Sebastian von Rohrscheidt, Josef van Genabith, Alexander Löser, Simon Ostermann
机构
*
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI))
;
Saarland University(萨尔兰州立大学)
;
Berliner Hochschule für Technik(柏林技术学院)
专题命中
指令微调
:LLM(title,title_cn);SFT(summary_cn,abstract);large language model(abstract);language model(abstract)
机构
*
National Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室)
;
Nanjing University(南京大学)
;
Institute of AI Industry Research (AIR)(人工智能产业研究院)
;
Tsinghua University(清华大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
ChemBIC(化学信息学中心)
专题命中
指令微调
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);SFT(abstract,abstract_cn)
BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens
BudgetThinker: 通过控制令牌赋能预算感知的LLM推理
Hao Wen, Xinrui Wu, Yi Sun, Feifei Zhang, Liye Chen, Jie Wang, Yunxin Liu, Yunhao Liu, Ya-Qin Zhang, Yuanchun Li
机构
*
Institute for AI Industry Research (AIR) Tsinghua University(人工智能产业研究院(AIR)清华大学)
;
Global Innovation Exchange & Department of Automation Tsinghua University(全球创新交流中心及自动化系 清华大学)
专题命中
指令微调
:LLM(title,title_cn);SFT(abstract,abstract_cn);large language model(abstract);language model(abstract)
Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation
通过经验知识集成与激活推动LLM工具调用极限
Yupu Hao, Zhuoran Jin, Huanxuan Liao, Kang Liu, Jun Zhao
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所复杂系统认知与决策智能重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
专题命中
指令微调
:LLM(title,title_cn);large language model(abstract);language model(abstract);post-training(abstract)
Case Study: Fine-tuning Small Language Models for Accurate and Private CWE Detection in Python Code
案例研究:微调小型语言模型以在Python代码中实现准确且隐私的CWE检测
Md. Azizul Hakim Bappy, Hossen A Mustafa, Prottoy Saha, Rajinus Salehat
机构
*
Institute of Information and Communication Technology, Bangladesh University of Engineering Technology(孟加拉工程科技大学信息与通信技术研究所)
;
Hajee Mohammad Danesh Science and Technology University(海杰莫哈默德丹什科学与技术大学)
专题命中
指令微调
:language model(title,abstract);small language model(title,abstract);LLM(abstract,abstract_cn);SLM(abstract,abstract_cn)
Large Language Models Align with the Human Brain during Creative Thinking
大型语言模型在创造性思维过程中与人类大脑对齐
Mete Ismayilzada, Simone A. Luchini, Abdulkadir Gokce, Badr AlKhamissi, Antoine Bosselut, Antonio Laverghetta, Lonneke van der Plas, Roger E. Beaty
机构
*
EPFL(瑞士联邦理工学院洛桑)
;
Università della Svizzera italiana (USI)(意大利语区大学(瑞士))
;
Wesleyan University(卫斯理大学)
;
Paris Brain Institute (ICM)(巴黎大脑研究所)
;
Pennsylvania State University(宾夕法尼亚州立大学)
专题命中
指令微调
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);post-training(abstract)