Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
Nested-ReFT:通过离策略展开实现大语言模型微调的高效强化学习
Maxime Heuillet, Yufei Cui, Boxing Chen, Audrey Durand, Prasanna Parthasarathi
机构
*
Mila - Québec AI Institute, Canada(魁北克人工智能研究所)
;
Huawei Noah's Ark Lab (Montreal Research Center), Canada(华为诺亚实验室(蒙特利尔研究中心))
;
Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)
Enabling Agents to Communicate Entirely in Latent Space
使智能体能够在潜在空间中完全交流
Zhuoyun Du, Runze Wang, Huiyu Bai, Zouying Cao, Xiaoyong Zhu, Yu Cheng, Bo Zheng, Wei Chen, Haochao Ying
机构
*
State Key Lab of CAD&CG(CAD与CG国家重点实验室)
;
Future Living Lab of Alibaba(阿里巴巴未来生活实验室)
;
Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence(浙江医学影像人工智能重点实验室)
;
Shanghai Jiao Tong University(上海交通大学)
LLM-Driven Collaborative Model for Untangling Commits via Explicit and Implicit Dependency Reasoning
基于显式和隐式依赖推理的大语言模型驱动的提交解缠协作模型
Bo Hou, Xin Tan, Kai Zheng, Fang Liu, Yinghao Zhu, Li Zhang
机构
*
State Key Laboratory of Complex \& Critical Software Environment (SKLCCSE), School of Computer Science
;
Engineering, Beihang University Beijing China
;
SKLCCSE, School of Computer Science
;
School of Computing
;
Data Science, The University of Hong Kong Hong Kong SAR China
;
Engineering, Beihang University
;
Data Science, The University of Hong Kong
Measuring AI Ability to Complete Long Software Tasks
衡量AI完成长期软件任务的能力
Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkovic, Luke Harold Miles, Seraphina Nix, Tao Lin, Chris Painter, Neev Parikh, David Rein, Lucas Jun Koba Sato, Hjalmar Wijk, Daniel M. Ziegler, Elizabeth Barnes, Lawrence Chan
机构
*
Model Evaluation & Threat Research (METR)(模型评估与威胁研究(METR))
;
Ohm Chip
;
Anthropic
Mao-Lin Luo, Zhe-Xu Wang, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei
机构
*
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education(教育部计算机网络和信息集成重点实验室(东南大学))
;
Zhongguancun Academy(中关村科学城公司)
;
Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)
Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment
AI能像城市规划师一样推理吗?基于专业判断的大语言模型基准测试
Yijie Deng, He Zhu, Wen Wang, Junyou Su, Minxin Chen, Wenjia Zhang
机构
*
School of Architecture and Urban Planning, Shenzhen University(深圳大学建筑与城市规划学院)
;
Shenzhen Key Laboratory of Urban Spatial Information and Intelligent Modeling(深圳市城市空间信息与智能建模重点实验室)
;
Department of Urban Planning and Design, The University of Hong Kong(香港大学城市规划与设计系)
CommentsThis paper has been withdrawn by the authors because the current version requires substantial revision and further validation before it can be considered a reliable representation of the work
People use fast and flat simulation to reason about new games
人们使用快速且扁平的模拟来对新游戏进行推理
Katherine M. Collins, Cedegao E. Zhang, Lionel Wong, Mauricio Barba da Costa, Graham Todd, Adrian Weller, Samuel J. Cheyette, Thomas L. Griffiths, Joshua B. Tenenbaum
机构
*
Massachusetts Institute of Technology(麻省理工学院)
;
Princeton University(普林斯顿大学)
;
University of Cambridge(剑桥大学)
;
Stanford University(斯坦福大学)
;
New York University(纽约大学)
;
The Alan Turing Institute(艾伦·图灵研究所)
机构
*
Analemma
;
National University of Singapore(新加坡国立大学)
;
Fudan University(复旦大学)
;
University College Dublin(都柏林大学)
;
Washington University in St. Louis(圣路易斯华盛顿大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
ByteDance(字节跳动)
;
ShanghaiTech University(上海科技大学)
;
University of Warwick(华威大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
University of Michigan, Ann Arbor(密歇根大学安娜堡分校)
;
East China Normal University(华东师范大学)
;
Stanford University(斯坦福大学)
;
Tencent(腾讯)
;
The University of Hong Kong(香港大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
Nex-AGI Team(Nex-AGI团队)
Comments@2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
机构
*
Research Center for Space Computing System, Zhejiang Lab, Hangzhou(杭州浙大实验室空间计算系统研究中心)
;
Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, Hangzhou(中国科学院大学杭州高等研究院)
;
School of Cyber Science and Engineering, Zhengzhou University(郑州大学计算机科学与工程学院)
;
School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)