Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data
平均场 PhiBE:基于离散时间数据的连续时间平均场强化学习
Erhan Bayraktar, Martin Hernandez, Qinxin Yan, Yuhua Zhu
机构
*
Department of Mathematics, University of Michigan, Ann Arbor, MI, USA(密歇根大学数学系,安阿伯,密歇根州,美国)
;
Department of Statistics and Data Science, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校统计与数据科学系,加利福尼亚州,美国)
;
Program in Applied and Computational Mathematics, Princeton University, Princeton, NJ, USA(普林斯顿大学应用与计算数学项目,普林斯顿,新泽西州,美国)
机构
*
Brain-inspired Cognitive AI Lab, Institute of Automation, Chinese Academy of Sciences, Beijing, China(脑启发认知人工智能实验室,自动化研究所,中国科学院,北京,中国)
;
Beijing Institute of AI Safety and Governance, China(北京人工智能安全与治理研究院,中国)
;
State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与脑启发智能技术国家重点实验室)
;
Beijing Key Laboratory of Safe AI and Superalignment, China(北京安全人工智能与超对齐重点实验室,中国)
;
University of Chinese Academy of Sciences (UCAS), Beijing, China(中国科学院大学(UCAS),北京,中国)
;
Long-term AI,Beijing,China(长期人工智能,北京,中国)