LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language Models
LearNAT: 基于AST引导任务分解的NL2SQL大语言模型学习
Weibin Liao, Xin Gao, Tianyu Jia, Rihong Qiu, Yifan Zhu, Yang Lin, Xinyu Ma, Junfeng Zhao, Yasha Wang
机构
*
School of Computer Science, Peking University(北京大学计算机科学系)
;
Key Laboratory of High Confidence Software Technologies, Ministry of Education(教育部高可信软件技术重点实验室)
;
Big Data Technology Research Center, Nanhu Laboratory(纳米实验室大数据技术研究中心)
;
National Engineering Research Center For Software Engineering, Peking University(北京大学软件工程国家工程研究中心)
;
Peking University Information Technology Institute (Tianjin Binhai)(北京大学信息技术研究院(天津滨海))
;
School of Computer Sciences, Beijing University of Posts and Telecommunications(北京邮电大学计算机科学系)
;
Huawei Technologies Co., Ltd(华为技术有限公司)
;
Seed, ByteDance Inc.(字节跳动公司)
专题命中
后训练与偏好优化
:LLM(summary_cn,abstract);large language model(title);language model(title);preference optimization(abstract)
CommentsWe found a critical flaw in the prompt complexity metric, which affects the 2D curriculum grid construction and leads to potentially invalid comparisons. Since this undermines our main conclusions, we are withdrawing the paper and will revise the methodology before resubmission
机构
*
Department of Data Science and Hong Kong Institute of AI for Science, City University of Hong Kong(数据科学系和香港人工智能科学研究所,香港城市大学)
;
Li Auto Inc., China(中国利汽车公司)
;
Department of Statistics, University of Oxford(统计系,牛津大学)
专题命中
后训练与偏好优化
:large language model(title,abstract_cn);language model(title,abstract_cn);LLM(abstract,abstract_cn);分类 cs.CL
Tail-Aware Information-Theoretic Bounds for LLM Alignment under Heavy-Tailed Rewards
尾感知信息论泛化用于RLHF和SGLD
Huiming Zhang, Binghan Li, Wan Tian, Qiang Sun
机构
*
Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)
;
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(北京未来区块链与隐私计算高精尖创新中心)
;
Advanced Institute of Information Technology, Peking University(北京大学信息技术高等研究院)
;
Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机技术研究所)
;
Computer and Mathematical Sciences, Computer Science, and Statistics, University of Toronto(多伦多大学计算机与数学科学、计算机科学和统计学系)
;
MBZUAI(穆罕默德·本·扎耶德人工智能大学)
专题命中
后训练与偏好优化
:RLHF(title_cn,abstract);LLM(title);large language model(abstract);language model(abstract)
PolyFact: Comparing Consistency-Driven Post-training Methods for Cross-Lingual Factual Recall
通过一致性驱动的强化学习改进跨语言事实回忆
Jonathan von Rad, Louis Arts, George Burgess, Eleftheria Kolokytha, Harry O'Donnell, Ektor Oikonomidis Doumpas, Eduardo Sanchez, Yao Lu, Pontus Stenetorp
机构
*
University College London(伦敦大学学院)
;
Centre for Artificial Intelligence(人工智能中心)
专题命中
后训练与偏好优化
:post-training(title,abstract);SFT(abstract,abstract_cn);large language model(abstract);language model(abstract)
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
从可验证奖励强化学习到自验证奖励强化学习:任务转换为开放式语言模型自我改进带来自验证奖励
Qinsi Wang, Jing Shi, Huazheng Wang, Kun Wan, Yiran Wu, Bo Liu, Qingyun Wu, Hai Helen Li, Yiran Chen, Handong Zhao, Wentian Zhao
机构
*
Duke University(杜克大学)
;
Adobe Inc.(奥多比公司)
;
Oregon State University(俄勒冈州立大学)
;
Pennsylvania State University(宾夕法尼亚州立大学)
;
National University of Singapore(新加坡国立大学)
;
Amazon(亚马逊)
专题命中
后训练与偏好优化
:LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI
Commentsv2: added sycophancy related work; Section 4 positioned against concurrent Bradley-Terry amplification results (Shapira et al. 2026); minor revisions
机构
*
College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息与电子工程学院)
;
Huawei Technologies Company Ltd.(华为技术有限公司)
;
Zhejiang Lab(之江实验室)
;
Zhejiang University(浙江大学)
;
Macau University of Science and Technology(澳门科技大学)
机构
*
Stanford University(斯坦福大学)
;
Georgia Institute of Technology(佐治亚理工学院)
;
The University of Tokyo(东京大学)
;
RIKEN AIP(日本理化学研究所智能系统研究中心)
;
Pennsylvania State University(宾夕法尼亚州立大学)
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
Harvard University(哈佛大学)
;
UNC–Chapel Hill(北卡罗来纳大学教堂山分校)
专题命中
后训练与偏好优化
:preference optimization(title,abstract);RLHF(abstract,abstract_cn);large language model(abstract);language model(abstract)