RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models
RewardFlow: 面向大语言模型智能体强化学习的拓扑感知状态图奖励传播
Xiao Feng, Bo Han, Zhanke Zhou, Jiaqi Fan, Jiangchao Yao, Ka Ho Li, Dahai Yu, Michael Kwok-Po Ng
机构
*
TMLR Group(TMLR小组)
;
Hong Kong Baptist University(香港 Baptist大学)
;
TCL Corporate Research (HK) Co Ltd(TCL企业研究(香港)有限公司)
;
Cooperative Medianet Innovation Center Shanghai Jiao Tong University(合作中位网创新中心上海交通大学)
;
Department of Mathematics Hong Kong Baptist University(香港 Baptist大学数学系)
机构
*
Nanyang Technological University(南洋理工大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Illinois Chicago(伊利诺伊大学香槟分校)
;
Tsinghua University(清华大学)
;
Sun Yat-sen University(中山大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
机构
*
Department of Electrical and Computer Engineering, University of California, Santa Barbara, Santa Barbara, CA 93106, USA(电子工程系,加州大学圣芭芭拉分校)
;
College of Engineering, Northeastern University, Boston, MA 02115, USA(工程学院,东北大学)
;
Department of Biomedical Engineering, New Jersey Institute of Technology, Newark, NJ 07102, USA(生物医学工程系,新泽西理工学院)
;
United Imaging Intelligence, Burlington, MA 01803, USA(联合影像智能公司)
;
Department of Radiation Oncology, Mayo Clinic, Scottsdale, AZ 85259, USA(放射肿瘤科,梅奥诊所)
;
Department of Radiology, Massachusetts General Hospital and Harvard Medical School, Boston, MA 02115, USA(放射科,麻省总医院和哈佛医学院)
;
School of Computing, University of Georgia, Athens, GA 30602, USA(计算学院,佐治亚大学)