Guikun Chen, Yuqian Chen, Yijie Li, Yogesh Rathi, Nikos Makris, Fan Zhang, Wenguan Wang, Lauren J. O'Donnell
机构
*
The State Key Lab of Brain-Machine Intelligence, Zhejiang University, Hangzhou(脑机智能国家重点实验室,浙江大学,杭州)
;
Department of Radiology, Brigham and Women’s Hospital, Mass General Brigham, Boston(放射科,布里洛妇女医院,马萨诸塞总医院,波士顿)
;
Harvard Medical School, Boston(哈佛医学院,波士顿)
;
Academy of Medical Engineering and Translational Medicine, Tianjin University, Tianjin(医学工程与转化医学研究院,天津大学,天津)
;
School of Information and Communication Engineering, University of Electronic Science and Technology of China, Chengdu(信息与通信工程学院,电子科技大学,成都)
;
Psychiatry Neuroimaging Laboratory, Brigham and Women’s Hospital, Mass General Brigham, Boston(精神病神经影像实验室,布里洛妇女医院,马萨诸塞总医院,波士顿)
;
Department of Psychiatry, Center for Morphometric Analysis, Massachusetts General Hospital, Boston(精神病科,形态分析中心,马萨诸塞总医院,波士顿)
STORM: Stepwise Token Optimization with Reward-Guided Beam Search
STORM: 基于奖励引导束搜索的逐步令牌优化
Arthur Satouf, Giulio D'Erasmo, Yuxuan Zong, Habiboulaye Amadou Boubacar, Pablo Piantanida, Benjamin Piwowarski
机构
*
MILA – Quebec AI Institute & ILLS(魁北克人工智能研究所与ILLs)
;
Université Paris-Saclay & CentraleSupélec & CNRS(巴黎-萨克雷大学及CentraleSupélec与CNRS)
;
Air Liquide
;
Sorbonne Université & ISIR & CNRS(索邦大学及ISIR与CNRS)
;
Sapienza, University of Rome(罗马大学Sapienza)
Toan Nguyen, Yang Liu, Trung Le, Celso de Melo, Flora D. Salim
机构
*
School of Computer Science and Engineering, University of New South Wales(新南威尔士大学计算机科学与工程学院)
;
Department of Data Science & AI, Monash University(莫纳什大学数据科学与人工智能系)
;
DEVCOM Army Research Laboratory(DEVCOM陆军研究实验室)
机构
*
Shanghai Key Lab of Intelligent Information Processing, Fudan University(复旦大学上海智能信息处理重点实验室)
;
School of Computer Science, Fudan University(复旦大学计算机科学技术学院)
;
Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心)
;
Youtu Lab, Tencent(腾讯优图实验室)
;
Meta AI
;
Shanghai AI Laboratory(上海人工智能实验室)
When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff
当强化学习在监督微调后失效:恢复模型可塑性以实现稳健的SFT到RL交接
Runze Liu, Jiashun Liu, Xu Wan, Yuqian Fu, Ling Pan
机构
*
Hong Kong University of Science and Technology(香港科技大学)
;
Zhejiang University(浙江大学)
;
State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(多模态人工智能系统国家重点实验室,CASIA)
专题命中
指令微调
:SFT(title,title_cn);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)
Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation
通过经验知识集成与激活推动LLM工具调用极限
Yupu Hao, Zhuoran Jin, Huanxuan Liao, Kang Liu, Jun Zhao
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所复杂系统认知与决策智能重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
专题命中
指令微调
:LLM(title,title_cn);large language model(abstract);language model(abstract);post-training(abstract)
机构
*
Beijing Key Laboratory of Security and Privacy in Intelligent Transportation, Beijing Jiaotong University(北京交通大学智能交通信息安全与隐私保护北京市重点实验室)
;
College of Computer Science and Technology, Taiyuan University of Technology(太原理工大学计算机科学与技术学院)
;
Institute of Computing Technologies, China Academy of Railway Sciences Corporation Limited(中国铁道科学研究院集团有限公司计算技术研究所)
专题命中
指令微调
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI
Qi Xiong, Renzhi Chen, Bowei Wang, Yuqing Xiong, Libo Huang, Lei Wang
机构
*
College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机科学与技术学院)
;
Defense Innovation Institute, Academy of Military Science (AMS)(军事科学院创新院)
;
School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院)
;
Key Laboratory of Advanced Microprocessor Chips and Systems, Changsha, China, and College of Computer Science and Technology, National University of Defense Technology, Changsha, China(先进微处理器芯片与系统重点实验室,长沙,中国,和国防科技大学计算机科学与技术学院,长沙,中国)
;
Defense Innovation Institute, AMS, Beijing, China and Qiyuan Lab, Beijing, China(军事科学院创新院,北京,中国和启元实验室,北京,中国)
专题命中
指令微调
:LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.AI