Improving Reasoning Capabilities in Small Models through Mixture-of-Layers Distillation with Stepwise Attention on Key Information
通过关键信息的分步注意力混合层蒸馏提升小模型的推理能力
Yao Chen, Jiawei Sheng, Wenyuan Zhang, Tingwen Liu
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
TabularMath: Understanding Math Reasoning over Tables with Large Language Models
TabularMath: 通过大规模语言模型理解表格上的数学推理
Shi-Yu Tian, Zhi Zhou, Wei Dong, Kun-Yang Yu, Ming Yang, Zi-Jian Cheng, Lan-Zhe Guo, Yu-Feng Li
机构
*
National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室)
;
School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)
;
School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)
When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning
何时信任工具?用于工具集成数学推理的自适应工具信任校准
Ruotao Xu, Yixin Ji, Yu Luo, Jinpeng Li, Dong Li, Peifeng Li, Juntao Li, Min Zhang
机构
*
School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院)
;
Department of Foundation Model, 2012 Labs, Huawei(华为基础模型部门)
;
Harbin Institute of Technology, Shenzhen (HITSZ)(哈尔滨工业大学深圳(HITSZ))
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
SelfBudgeter:面向高效大语言模型推理的自适应令牌分配
Zheng Li, Qingxiu Dong, Jingyuan Ma, Di Zhang, Kai Jia, Zhifang Sui
机构
*
State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)
;
BandAI, Bytedance(字节跳动BandAI)
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
University of Oxford(牛津大学)
;
Imperial College London(伦敦帝国学院)
;
University of Georgia(佐治亚大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Chinese Academy of Sciences(中国科学院)
;
Dalian University of Technology(大连理工大学)
;
National University of Singapore(新加坡国立大学)
;
Wuhan University(武汉大学)
Disentangling Mathematical Reasoning in LLMs: A Methodological Investigation of Internal Mechanisms
解构大语言模型中的数学推理:对内部机制的调查
Tanja Baeumel, Josef van Genabith, Simon Ostermann
机构
*
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
;
Saarland University(萨尔兰大学)
;
Center for European Research in Trusted AI (CERTAIN)(可信人工智能欧洲研究中心)
Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
用评分奖励模型治愈大语言模型数学推理中的奇迹步骤
Youliang Yuan, Qiuyang Mang, Jingbang Chen, Hong Wan, Xiaoyuan Liu, Junjielong Xu, Jen-tse Huang, Wenxuan Wang, Wenxiang Jiao, Pinjia He
机构
*
School of Data Science, The Chinese University of Hong Kong, Shenzhen, China(数据科学学院,香港中文大学(深圳))
;
UC Berkeley(加州大学伯克利分校)
;
Zhejiang University(浙江大学)
;
Johns Hopkins University(约翰霍普金斯大学)
;
Renmin University of China(中国人民大学)
;
Xiaohongshu Inc.(小红书公司)
Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning
重新审视熵正则化:自适应系数解锁其在大语言模型强化学习中的潜力
Xiaoyun Zhang, Xiaojian Yuan, Di Huang, Wang You, Chen Hu, Jingqing Ruan, Ai Jian, Kejiang Chen, Xing Hu
机构
*
State Key Lab of Processors, Institute of Computing Technology, CAS(处理器国家重点实验室,计算技术研究所,中国科学院)
;
University of Science and Technology of China(中国科学技术大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
StepFun Inc(StepFun公司)
机构
*
Department of Computer Science and Engineering(计算机科学与工程系)
;
University of Moratuwa(穆拉图瓦大学)
;
Department of Electrical and Information Engineering(电气与信息工程系)
;
University of Ruhuna(鲁胡纳大学)
;
Faculty of Information Technology and Communication Sciences(信息科技与通讯科学学院)
;
Tampere University(塔尔皮奥大学)
;
School of Computing(计算学院)
;
Australian National University(澳大利亚国立大学)