SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
SCOPE-RL: 稳定和定量控制强化学习后训练中的策略熵
Chen Wang, Zhaochun Li, Jionghao Bai, Hexuan Deng, Ge Lan, Yue Wang
机构
*
College of Software, Nankai University(南开大学软件学院)
;
Zhongguancun Academy(中关村学院)
;
Beijing Institute of Technology(北京理工大学)
;
Zhejiang University(浙江大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
Quantum Sidecar Architectures for Hybrid AI Training and Inference: Stateful Protected Registers, Stateless Reset-and-Reprepare Circuits and Quantum Weight-State Outlook
机构
*
New York University Tandon School of Engineering(纽约大学Tandon工程学院)
;
University of California, San Diego(加州大学圣地亚哥分校)
;
Stanford University(斯坦福大学)
;
ServiceNow Research(ServiceNow研究)
;
Mila - Quebec AI Institute(魁北克AI研究院)
;
Université de Montréal(蒙特利尔大学)
KGPFN: Unlocking the Potential of Knowledge Graph Foundation Model via In-Context Learning
KGPFN:通过上下文学习解锁知识图谱基础模型的潜力
Yisen Gao, Jiaxin Bai, Haoyu Huang, Zhongwei Xie, Yufei Li, Hong Ting Tsang, Sirui Han, Yangqiu Song
机构
*
Department of Computer Science and Engineering, HKUST, Hong Kong, China(香港科技大学计算机科学与工程系)
;
Department of Computer Science and Engineering, HKBU, Hong Kong, China(香港城市大学计算机科学与工程系)
Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
计算作为教师:将推理计算转化为无参考监督
Dulhan Jayalath, Shashwat Goel, Thomas Foster, Parag Jain, Suchin Gururangan, Cheng Zhang, Anirudh Goyal, Alan Schelten
机构
*
University of Oxford(牛津大学)
;
ELLIS Institute Tübingen(图宾根ELLIS研究所)
;
Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所)
;
Meta Superintelligence Labs(元宇宙超级智能实验室)
机构
*
School of Computer Science and Technology, Tianjin University, China(天津大学计算机科学与技术学院)
;
School of Future Technology, Tianjin University, China(天津大学未来技术学院)
;
Georgia Tech Shenzhen Institute, Tianjin University, China(Georgia Tech深圳研究院)
;
School of Artificial Intelligence, Tianjin University, China(天津大学人工智能学院)
机构
*
School of Mathematics, Tianjin University, Tianjin, China(天津大学数学学院)
;
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China(中国科学院信息工程研究所)
;
Shanghai Advanced Institute of Finance (SAIFS), East China Normal University, Shanghai, China(上海先进金融研究所(SAIFS),东华大学)
;
ENN Group, Digital Technology Research Institute, China(ENN集团,数字技术研究院)
;
Nanyang Technological University, Singapore(南洋理工大学)
;
National University of Singapore, Singapore(新加坡国立大学)
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)
;
Nanjing University of Science(南京理工大学)
;
City University of Hong Kong(香港城市大学)
GUI Agents with Reinforcement Learning: Toward Digital Inhabitants
具有强化学习的图形用户界面代理:迈向数字居民
Junan Hu, Jian Liu, Jingxiang Lai, Jiarui Hu, Yiwei Sheng, Shuang Chen, Jian Li, Dazhao Du, Song Guo
机构
*
Shandong University(山东大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
The Hong Kong University(香港大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Tencent(腾讯)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
University of California San Diego(加州大学圣地亚哥分校)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Fudan University(复旦大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
推理时的验证扩展:通过测试时的评分指南验证实现自我进化深度研究代理
Yuxuan Wan, Tianqing Fang, Zaitang Li, Yintong Huo, Wenxuan Wang, Haitao Mi, Dong Yu, Michael R. Lyu
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Tencent AI Lab(腾讯人工智能实验室)
;
Singapore Management University(新加坡管理学院)
;
The Renmin University of China(中国人民大学)