Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
基于重新定义的逐步优势的自引导过程奖励优化用于过程强化学习
Wu Fei, Shuxian Liang, Yibo Yang, Yang Lin, Jing Tang, Lei Chen, Xiansheng Hua, Hao Kong
机构
*
Terminus Group(Terminus集团)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹大学)
Large Language Models Explore by Latent Distilling
通过潜在蒸馏探索大语言模型
Yuanhao Zeng, Ao Lu, Lufei Li, Zheng Zhang, Yexin Li, Kan Ren
机构
*
State Key Laboratory of General Artificial Intelligence, BIGAI, Beijing, China(人工智能通用基础理论国家重点实验室,BIGAI,北京,中国)
;
School of Information Science and Technology, ShanghaiTech University, Shanghai, China(信息科学与技术学院,上海交通大学,上海,中国)
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
最强的教师并不总是最好的教师:以学生为中心的答案选择
Zhengyu Hu, Zheyuan Xiao, Linxin Song, Fengqing Jiang, Yuetai Li, Zhihan Xiong, Yue Liu, Junhao Lin, Yao Su, Lijie Hu, Kaize Ding, Teng Xiao, Radha Poovendran
机构
*
University of Washington(华盛顿大学)
;
University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
University of Southern California(南加州大学)
;
Independent Researcher(独立研究者)
;
National University of Singapore(新加坡国立大学)
;
Microsoft(微软)
;
Google(谷歌)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Northwestern University(西北大学)
;
Allen Institute for AI (AI2)(人工智能研究院(AI2))
机构
*
Tencent HY LLM Frontier(腾讯HY大模型前沿团队)
;
University of Georgia(佐治亚大学)
;
University of Maryland, College Park(马里兰大学帕克分校)
;
University of Pennsylvania(宾夕法尼亚大学)
;
University of Minnesota, Twin Cities(明尼苏达大学双城分校)
;
Indiana University(印第安纳大学)
;
National University of Singapore(新加坡国立大学)
;
Hong Kong Polytechnic University(香港理工大学)
CommentsTitle changed from "POISE: Position-Aware Undetectable Skill Injection on LLM Agents" to "Poise: Position-Aware One-Instruction Skill Injection for Silent Execution on LLM Agents"; the manuscript, figures, appendices, and reproducibility details have been updated. 15 pages, 2 figures, 4 tables
MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph
MedKGent:用于构建随时间演变的医学知识图谱的大语言模型智能体框架
Duzhen Zhang, Zixiao Wang, Zhong-Zhi Li, Yahan Yu, Shuncheng Jia, Jiahua Dong, Haotian Xu, Xing Wu, Yingying Zhang, Tielin Zhang, Jie Yang, Xiuying Chen, Le Song
机构
*
Mohamed bin Zayed University of Artificial Intelligence(莫扎德大学人工智能学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Kyoto University(京都大学)
;
Tsinghua University(清华大学)
;
East China Normal University(华东师范大学)
;
Center for Excellence in Brain Science and Intelligence Technology(脑科学与智能技术卓越中心)
;
Brigham and Women’s Hospital, Harvard Medical School(哈佛医学院布里特妇女医院)
;
GenBio AI
机构
*
City University of Hong Kong(香港城市大学)
;
Tsinghua University(清华大学)
;
Shenzhen University of Advanced Technology(深圳理工大学)
;
Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
Comments15 pages, 1 figure, 8 tables. Major revision with locked MIMIC-IV transfer evaluation, blinded pairwise human assessment, matched multi-seed ablations, revised title, and revised author list
机构
*
City University of Hong Kong(香港城市大学)
;
The Institute of Statistical Mathematics(统计数学研究所)
;
University of Sydney(悉尼大学)
;
Nanyang Technological University(南洋理工大学)
;
The University of Tokyo(东京大学)