Improving Small-Scale Large Language Models Function Calling for Reasoning Tasks
专题命中 数学推理 :reasoning(title,abstract);logical reasoning(abstract);分类 cs.AI
AI 大模型
大模型数学、逻辑、规划、多步推理和测试时计算能力。
专题命中 数学推理 :reasoning(title,abstract);logical reasoning(abstract);分类 cs.AI
专题命中 数学推理 :verifier(title,abstract);reasoning(abstract);分类 cs.CL
Comments 15 pages
专题命中 数学推理 :reasoning(title,abstract);verifier(abstract);分类 cs.CL
Comments under review
专题命中 数学推理 :reasoning(title,abstract);CoT(abstract);分类 cs.CL
Comments EMNLP Findings 2024 Camera Ready
专题命中 数学推理 :reasoning(title,abstract);CoT(abstract);分类 cs.AI
Comments 13 pages, 3 figures
专题命中 数学推理 :reasoning(title,abstract);CoT(abstract);分类 cs.CL
专题命中 数学推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL
专题命中 数学推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL
Comments Accepted to Findings of ACL 2024
专题命中 数学推理 :reasoning(title,abstract);CoT(abstract);分类 cs.CL
Comments This paper has been accepted to the ACL 2024 main conference
专题命中 数学推理 :reasoning(title,abstract);math reasoning(abstract);分类 cs.CL
Comments ICML 2024 (Camera-ready); First two authors contributed equally; GitHub: https://github.com/dinobby/MAGDi
专题命中 数学推理 :reasoning(title,abstract);logical reasoning(abstract);分类 cs.CL
Comments ACL 2024 Main
专题命中 数学推理 :chain-of-thought(title,abstract);reasoning(abstract);分类 cs.CL
Comments In press at EAAI-24: The 14th Symposium on Educational Advances in Artificial Intelligence
专题命中 数学推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL
专题命中 数学推理 :reasoning(title,abstract);CoT(abstract);分类 cs.CL
Comments Natural language processing (NLP)
专题命中 数学推理 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL
专题命中 数学推理 :reasoning(title,abstract);verifier(abstract);分类 cs.CL
Comments Accepted to ACL 2023 Main Conference; Camera Ready
专题命中 数学推理 :CoT(title,abstract);reasoning(abstract);分类 cs.CL
Comments Accepted to ENMLP 2023
专题命中 数学推理 :chain-of-thought(title,abstract);reasoning(abstract);分类 cs.CL
Comments EMNLP 2023
专题命中 数学推理 :reasoning(title,abstract);CoT(abstract);分类 cs.CL
机构 * Open-Thought ; Scale AI ; University College London(伦敦大学学院)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments NeurIPS 2025 Spotlight. For code, see https://github.com/open-thought/reasoning-gym
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments The state-of-the-art open-source language models for mathematical reasoning
HiPO:用于大语言模型自适应推理的分层偏好优化
机构 * Vellore Institute of Technology(维洛雷理工学院) ; University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校) ; Northwestern University(西北大学) ; Yale University(耶鲁大学) ; Algoverse AI Research(Algoverse AI研究)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.AI、cs.LG;math reasoning(comments)
AI总结 HiPO通过分层优化提升大语言模型在复杂推理任务中的表现,结合偏好优化与结构化推理的优势,实现更高效的训练和更一致的输出。
Comments 12 pages, 4 figures, 6 tables. Includes ablation study across Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct on 5 math reasoning benchmarks (GSM8K, MATH500, Minerva, AIME24, Gaokao2023). GPT-4.1 used for structured evaluation of reasoning quality
Agent-GWO: 为大型语言模型的动态提示优化设计的协作代理
机构 * Kyung Hee University(韩国庆熙大学) ; University of Electronic Science and Technology of China(电子科技大学) ; Nota Inc.(Nota公司) ; Tongji University(同济大学)
专题命中 数学推理 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG
AI总结 本文提出Agent-GWO框架,通过统一提示模板和解码超参数作为可继承的代理配置,利用灰狼优化器的领导者-追随者机制,自动选择领导者代理以指导协作更新,提升复杂推理的准确性和稳定性。
Comments Accepted to ACL 2026. 9 pages, 5 figures
为何扩散语言模型在真正并行(非自回归)解码中表现不佳?
机构 * The Hong Kong Polytechnic University, Hong Kong, China(香港理工大学) ; ELLIS Institute Tübingen, Tübingen, Germany(图宾根ELLIS研究所) ; Max Planck Institute for Intelligent Systems, Tübingen, Germany(智能系统马克斯·普朗克研究所) ; Tübingen AI Center, Tübingen, Germany(图宾根人工智能中心) ; University of Surrey, Guildford, United Kingdom(萨里大学) ; The University of North Carolina at Chapel Hill, Chapel Hill, NC, USA(北卡罗来纳大学教堂山分校)
专题命中 数学推理 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);math reasoning(abstract)
AI总结 本文提出NAP方法,通过数据驱动策略改进扩散语言模型的非自回归并行解码性能。
专题命中 数学推理 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);math reasoning(abstract)
专题命中 数学推理 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);verifier(abstract)
AXIOM: 一种用于可验证数学推理的信任优先神经符号执行架构
机构 * Independent researcher(独立研究者)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 提出AXIOM架构,将语言模型限制为规范化器,通过确定性计算机代数系统管道实现可验证的数学推理,在4个MATH类别上达到94.36%的正确率和100%的信任度。
Comments Preprint. 16 pages, 2 figures. Live interactive demo: https://huggingface.co/spaces/Squagghy/moxia. Paper artifact and dataset on Zenodo (concept-DOI): 10.5281/zenodo.21906509
理解从预训练到训练后阶段的推理
机构 * New York University(纽约大学) ; Modal Labs(模态实验室) ; University of California, Los Angeles(加州大学洛杉矶分校) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Columbia University(哥伦比亚大学)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 研究强化学习在大语言模型从预训练到训练后阶段对推理的作用,以国际象棋为测试平台,按标准流程训练模型,发现预训练损失可预测RL后性能,RL奖励曲线斜率与预训练令牌有关,还揭示RL对SFT策略的影响,且在数学领域也有相同模式。
路径锁定专家:通过架构层面分离实现混合思维中的推理模式分离
机构 * Case Western Reserve University(凯斯西储大学) ; NII LLMC, Japan(日本NII LLMC) ; Michigan State University(密歇根州立大学)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出Path-Lock Expert,通过架构层面分离推理与非推理模式,减少推理泄漏,提升非推理模式的准确性和简洁性。
Comments 38 pages, 12 figures, 25 tables
在线推理校准:测试时训练使可泛化的符合性大语言模型推理
机构 * Department of Electrical Engineering and Computer Science (MIT EECS)(麻省理工学院电气工程与计算机科学系) ; Computer Science and Artificial Intelligence Laboratory (MIT CSAIL)(麻省理工学院计算机科学与人工智能实验室) ; Laboratory for Information and Decision Systems (MIT LIDS)(麻省理工学院信息与决策系统实验室) ; Computational Science and Engineering (MIT CSE)(麻省理工学院计算科学与工程) ; Massachusetts Institute of Technology(麻省理工学院)
专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出ORCA框架,通过测试时训练和符合性预测校准推理过程,提升大语言模型在分布变化下的效率和泛化能力,实验证明其在不同任务中具有更高的性能。
Comments Published as a conference paper at COLM 2026; 22 pages