An Actor-Critic Framework for Continuous-Time Jump-Diffusion Controls with Normalizing Flows
一个用于连续时间跳跃扩散控制的Actor-Critic框架
Liya Guo, Ruimeng Hu, Xu Yang, Yi Zhu
机构
*
Yau Mathematical Sciences Center, Tsinghua University(清华大学丘成桐数学科学中心)
;
Department of Mathematics, Tsinghua University(清华大学数学系)
;
Department of Mathematics, University of California, Santa Barbara(加州大学圣塔芭芭拉分校数学系)
;
Department of Statistics and Applied Probability, University of California, Santa Barbara(加州大学圣塔芭芭拉分校统计与应用概率系)
;
Yanqi Lake Beijing Institute of Mathematical Sciences and Applications(北京雁栖湖应用数学研究院)
机构
*
School of Cyber Science and Engineering, Sichuan University(四川大学网络空间安全学院)
;
Institute for Network Sciences and Cyberspace, Tsinghua University(清华大学网络科学与网络空间研究院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
;
National University of Singapore(新加坡国立大学)
;
College of Electronic Engineering, National University of Defense Technology(国防科技大学电子工程学院)
;
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
;
School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)
Vintix II: Decision Pre-Trained Transformer is a Scalable In-Context Reinforcement Learner
Vintix II:决策预训练变压器是一种可扩展的上下文强化学习者
Andrei Polubarov, Lyubaykin Nikita, Alexander Derevyagin, Artyom Grishin, Igor Saprygin, Aleksandr Serkov, Mark Averchenko, Daniil Tikhonov, Maksim Zhdanov, Alexander Nikulin, Ilya Zisman, Albina Klepach, Alexey Zemtsov, Vladislav Kurenkov
机构
*
Yale University(耶鲁大学)
;
Broad Institute of MIT and Harvard(麻省理工学院-哈佛大学博德研究所)
;
Google DeepMind(谷歌DeepMind)
;
Stanford University(斯坦福大学)
;
Genentech(基因泰克)
;
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
Cornell University(康奈尔大学)
;
Harvard University(哈佛大学)
Joint Knowledge Base Completion and Question Answering by Combining Large Language Models and Small Language Models
通过结合大语言模型和小语言模型实现知识库补全与问答的联合处理
Yinan Liu, Dongying Lin, Sigang Luo, Xiaochun Yang, Bin Wang
机构
*
School of Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程学院,沈阳,中国)
;
National Frontiers Science Center for Industrial Intelligence and Systems optimization, Northeastern University, Shenyang, China(东北大学工业智能与系统优化国家级前沿科学中心,沈阳,中国)
How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators
人类如何帮助大语言模型:评估和激励人类偏好标注者
Shang Liu, Hanzhao Wang, Zhongyao Ma, Xiaocheng Li
机构
*
Imperial College Business School, Imperial College London(帝国理工学院商学院,帝国理工学院)
;
University of Sydney Business School, University of Sydney(悉尼大学商学院,悉尼大学)
;
Meta
Sparse Gain Radio Map Reconstruction With Geometry Priors and Uncertainty-Guided Measurement Selection
稀疏增益无线电地图重建:结合几何先验与不确定性引导的测量选择
Zhihan Zeng, Ning Wei, Muhammad Baqer Mollah, Kaihe Wang, Phee Lep Yeoh, Fei Xu, Yue Xiu, Zhongpei Zhang
机构
*
National Key Laboratory of Wireless Communications, University of Electronic Science and Technology of China (UESTC)(电子科技大学通信抗干扰全国重点实验室)
;
University of Electronic Science and Technology of China (UESTC)(电子科技大学)
;
Department of Information Science Technology, University of Houston(休斯顿大学信息科学与技术系)
;
School of Science, Technology and Engineering, University of the Sunshine Coast(阳光海岸大学科学、技术与工程学院)
;
ZGC Institute of Ubiquitous-X Innovation and Applications(中关村泛在融合创新研究院)