Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning Model
在移动边缘训练:一种在线验证的提示选择方法用于大型推理模型的高效强化学习训练
Jiahao Wu, Ning Lu, Shengcai Liu, Kun Wang, Yanting Yang, Bailong Lin, Chen Jason Zhang, Li Qing, Ke Tang
机构
*
Southern University of Science and Technology(南方科技大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
The Hong Kong University of Science and Technology(香港科学理工大学)
;
Nanyang Technological University(南洋理工大学)
;
Rutgers University(罗格斯大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学理工大学(广州))
PAEC: Position-Aware Entropy Calibration for LLM Reasoning in RLVR
PAEC:面向RLVR中LLM推理的位置感知熵校准
Shumeng Yang, Yisu Liu, Jiayi Zheng, Zhaohui Yang, Linjing Li
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)
机构
*
School of Chemistry, Tel Aviv University(特拉维夫大学化学系)
;
The Center for Physics and Chemistry of Living Systems, Tel Aviv University(特拉维夫大学生命系统物理与化学中心)
;
School of Physics and Astronomy, Tel Aviv University(特拉维夫大学物理与天文学系)
;
The Center for Computational Molecular and Materials Science, Tel Aviv University(特拉维夫大学计算分子与材料科学中心)
Operationalising the Superficial Alignment Hypothesis via Task Complexity
通过任务复杂度操作化浅层对齐假设
Tomás Vergara-Browne, Darshan Patil, Ivan Titov, Siva Reddy, Tiago Pimentel, Marius Mosbach
机构
*
University of Maryland(马里兰大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
University of Washington(华盛顿大学)
;
University of Toronto(多伦多大学)
;
University of Edinburgh(爱丁堡大学)
AlphaOPT: Formulating Optimization Programs with Self-Improving LLM Experience Library
AlphaOPT: 利用自改进LLM经验库构建优化问题
Minwei Kong, Ao Qu, Xiaotong Guo, Wenbin Ouyang, Chonghe Jiang, Han Zheng, Yining Ma, Dingyi Zhuang, Yuhan Tang, Junyi Li, Shenhao Wang, Haris Koutsopoulos, Hai Wang, Cathy Wu, Jinhua Zhao
机构
*
Singapore-MIT Alliance for Research and Technology(新加坡-麻省理工联合研究技术联盟)
;
Massachusetts Institute of Technology(麻省理工学院)
;
University of Florida(佛罗里达大学)
;
Northeastern University(东北大学)
;
Singapore Management University(新加坡管理学院)
Sound and Complete Neurosymbolic Reasoning with LLM-Grounded Interpretations
基于LLM解释的完备且可靠的神经常识推理
Bradley P. Allen, Prateek Chhikara, Thomas Macaulay Ferguson, Filip Ilievski, Paul Groth
机构
*
University of Amsterdam(阿姆斯特丹大学)
;
University of Southern California(南加州大学)
;
Rensselaer Polytechnic Institute(拉特格斯理工学院)
;
Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)
Comments43 pages, 14 tables, 4 figures. Accepted to the 19th Conference on Neurosymbolic Learning and Reasoning (NeSy 2025); to appear Neurosymbolic Artifical Intelligence Special Issue on NeSy 2025 Extended Papers
Bridging the Agent-World Gap: Text World Models for LLM-based Agents
弥合智能体-世界鸿沟:面向基于LLM的智能体的文本世界模型
Yixia Li, Hongru Wang, Peng Lai, Zhiwen Ruan, He Zhu, Youxin Zhu, Ganlong Zhao, Minda Hu, Yun Chen, Sibei Yang, Peng Li, Jeff Z. Pan, Jia Pan, Guanhua Chen, Yang Liu, Guanbin Li
机构
*
Southern University of Science and Technology(南方科技大学)
;
University of Edinburgh(爱丁堡大学)
;
Peking University(北京大学)
;
Sun Yat-sen University(中山大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai University of Finance and Economics(上海财经大学)
;
Tsinghua University(清华大学)
;
The University of Hong Kong(香港大学)