GDEPO: Group Dual-dynamic and Equal-right Advantage Policy Optimization with Enhanced Training Data Utilization for Sample-Constrained Reinforcement Learning
GDEPO:基于增强训练数据利用的群体双动态和等权优势策略优化
Zhengqing Yan, Xinyang Liu, Yi Zhang, Fan Guo, ChengXun Jia, Junchen Wan, Yao Liu, Qi Liu, Jihao Huang, Kang Song
机构
*
State Key Laboratory of Engines, Tianjin University(发动机国家重点实验室,天津大学)
;
Artificial Intelligence Center, Cylingo Group(Cylingo集团人工智能中心)
Agnieszka Mensfelt, David Tena Cucala, Santiago Franco, Angeliki Koutsoukou-Argyraki, Vince Trencsenyi, Kostas Stathis
机构
*
Department of Computer Science, Royal Holloway, University of London(伦敦大学皇家霍洛威学院计算机科学系)
;
Department of Computer Science and Technology, University of Cambridge(剑桥大学计算机科学与技术系)
CommentsPresented at NeLaMKRR@KR, 2025 (arXiv:2511.09575). A shorter version of this work will appear in the Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2026)
A Comprehensive Survey on Benchmarks and Solutions in Software Engineering of LLM-Empowered Agentic System
Jiale Guo, Suizhi Huang, Mei Li, Dong Huang, Xingsheng Chen, Regina Zhang, Zhijiang Guo, Han Yu, Siu-Ming Yiu, Pietro Lio, Kwok-Yan Lam
机构
*
Digital Trust Centre, Nanyang Technological University, Singapore(南洋理工大学数字信任中心)
;
College of Computing and Data Science, Nanyang Technological University, Singapore(南洋理工大学计算与数据科学学院)
;
The Hong Kong University of Science and Technology, Hong Kong(香港科学与技术大学)
;
School of Computer Science, Shanghai Jiao Tong University, Shanghai, China(上海交通大学计算机科学学院)
;
School of Computing and Data Science, The University of Hong Kong, Hong Kong(香港大学计算与数据科学学院)
;
School of Computer Science and Technology, The University of Cambridge, UK(剑桥大学计算机科学与技术学院)
AI's Euclid's Elements Moment: From Language Models to Computable Thought
Xinmin Fang, Lingfeng Tao, Zhengxiong Li
机构
*
Department of Computer Science and Engineering(计算机科学与工程系)
;
University of Colorado Denver (CU Denver)(科罗拉多大学丹佛分校)
;
Department of Robotics and Mechatronics(机器人与机电学系)
;
Kennesaw State University (KSU)(凯斯维尔州立大学)
机构
*
Peking University International Innovation Center, Lin-gang Special Area (PKU-IICSH)(北京大学国际创新中心,临港特殊区域(PKU-IICSH))
;
Onesyn (Shanghai) Technology Co., Ltd(上海奥森科技有限公司)
;
CATARC Automotive Technology(Shanghai) Co.,Ltd(CATARC汽车技术(上海)有限公司)
;
ZEEKR Intelligent Technology Holding Limited(ZEKR智能技术控股有限公司)
;
Department of Automation, Tsinghua University(清华大学自动化系)
;
State Key Laboratory for Management and Control of Complex Systems, Chinese Academy of Sciences(复杂系统管理与控制国家重点实验室,中国科学院)
;
Macao Institute of Systems Engineering, Macau University of Science and Technology(澳门系统工程研究院,澳门科技大学)
CommentsThis paper is a more detailed version of the following publication: Lavindra de Silva, "HTN Acting: A Formalism and an Algorithm", in Proceedings of AAMAS 2018
A Neurosymbolic Approach to Natural Language Formalization and Verification
一种用于自然语言形式化和验证的神经符号方法
Chenyang An, Sam Bayless, Stefano Buliani, Darion Cassel, Byron Cook, Duncan Clough, Rémi Delmas, Nafi Diallo, Ferhat Erata, Nick Feng, Dimitra Giannakopoulou, Aman Goel, Aditya Gokhale, Joe Hendrix, Victor Heorhiadi, Marc Hudak, Dejan Jovanović, Andrew M. Kent, Benjamin Kiesl-Reiter, Jeffrey J. Kuna, Nadia Labai, Joseph Lilien, Divya Raghunathan, Zvonimir Rakamarić, Niloofar Razavi, Michael Tautschnig, Ali Torkamani, Nathaniel Weir, Michael W. Whalen, Jianan Yao
机构
*
Amazon Web Services(亚马逊网络服务)
;
University College London(伦敦大学学院)
;
University of Toronto(多伦多大学)
;
Queen Mary University of London(伦敦大学女王学院)