Grad2Reward: From Sparse Judgment to Dense Rewards for Improving Open-Ended LLM Reasoning
Grad2Reward: 从稀疏判断到密集奖励以提升开放性大语言模型推理
Zheng Zhang, Ao Lu, Yuanhao Zeng, Ziwei Shan, Jinjin Guo, Lufei Li, Yexin Li, Kan Ren
机构
*
School of Information Science and Technology, ShanghaiTech University(信息科学与技术学院,上海科技大学)
;
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)
ALIGN: Aligned Delegation with Performance Guarantees for Multi-Agent LLM Reasoning
ALIGN: 多智能体LLM推理中的对齐委托与性能保证
Tong Zhu, Baiting Chen, Jin Zhou, Hua Zhou, Sriram Sankararaman, Xiaowu Dai
机构
*
Department of Biostatistics, UCLA(生物统计学系,加州大学洛杉矶分校)
;
Department of Statistics and Data Science, UCLA(统计学与数据科学系,加州大学洛杉矶分校)
;
Department of Computer Science, UCLA(计算机科学系,加州大学洛杉矶分校)
;
Departments of Statistics and Data Science, and of Biostatistics, UCLA(统计学与数据科学系和生物统计学系,加州大学洛杉矶分校)
Towards Reasoning for PDE Foundation Models: A Reward-Model-Driven Inference-Time-Scaling Algorithm
面向PDE基础模型的推理:一种基于奖励模型的推理时扩展算法
Siddharth Mansingh, James Amarel, Ragib Arnab, Arvind Mohan, Kamaljeet Singh, Gerd J. Kunde, Nicolas Hengartner, Benjamin Migliori, Emily Casleton, Nathan A. Debardeleben, Ayan Biswas, Diane Oyen, Earl Lawrence
机构
*
College of Computer Science, Sichuan University(四川大学计算机学院)
;
Columbia University(哥伦比亚大学)
;
School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)
;
Apon AI and Brain-Computer Engineering Research Institute(Apon人工智能与脑机工程研究院)
;
Faculty of Science and Technology, University of Macau(澳门大学科学与技术学院)
;
Faculty of Applied Science and Engineering, University of Toronto(多伦多大学应用科学与工程学院)
;
Faculty of Computer Science and Information Technology, University of Malaya(马来亚大学计算机科学与信息技术学院)
;
Zhaolong Technology(智龙科技)
;
Purdue University(普渡大学)
Uncertainty Reasoning with Photonic Bayesian Machines
基于光子贝叶斯机器的不确定性推理
F. Brückerhoff-Plückelmann, H. Borras, S. U. Hulyal, L. Meyer, X. Ji, J. Hu, J. Sun, B. Klein, F. Ebert, J. Dijkstra, L. McRae, P. Schmidt, T. J. Kippenberg, H. Fröning, W. Pernice
Adapting Like Humans: A Metacognitive Agent with Test-time Reasoning
像人类一样适应:具有测试时推理的元认知代理
Yang Li, Zhiyuan He, Yuxuan Huang, Zhuhanling Xiao, Chao Yu, Meng Fang, Kun Shao, Jun Wang
机构
*
Huawei Noah’s Ark Lab(华为诺亚实验室)
;
University of Oxford(牛津大学)
;
Tsinghua University(清华大学)
;
University of Liverpool(利物浦大学)
;
University College London(伦敦大学学院)
Do We Truly Need So Many Samples? Multi-LLM Repeated Sampling Efficiently Scales Test-Time Compute
Jianhao Chen, Zishuo Xun, Bocheng Zhou, Han Qi, Hangfan Zhang, Qiaosheng Zhang, Yang Chen, Wei Hu, Yuzhong Qu, Wanli Ouyang, Shuyue Hu
机构
*
State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
The University of Auckland(奥克兰大学)
;
The Pennsylvania State University(宾夕法尼亚州立大学)
ROC-n-reroll: How verifier imperfection affects test-time scaling
Florian E. Dorner, Yatong Chen, André F. Cruz, Fanny Yang
机构
*
ETH Zürich(苏黎世联邦理工学院)
;
Max Planck ETH Center for Learning Systems(马克斯·普朗克ETH学习系统中心)
;
Max Planck Institute for Intelligent Systems, Tübingen(图宾根马克斯·普朗克智能系统研究所)
;
Tübingen AI Center(图宾根人工智能中心)