arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2026-07-09 至 2026-07-09 共收录 9 信号源:cs.CL, cs.AI, cs.LG

1. 数学推理 9 篇

2607.06447 2026-07-09 cs.AI cs.CL cs.MA 新提交 86%

Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory

Danus:利用事实图内存编排数学推理智能体

Jihao Liu, Guoxiong Gao, Zeming Sun, Bin Wu, Shurui Liu, Jiedong Jiang, Haocheng Ju, Leheng Chen, Ronnie Cheng, Xiping Zhang, Bin Dong

机构 * School of Mathematical Sciences, Peking University(北京大学数学科学学院) Beijing International Center for Mathematical Research, Peking University(北京大学北京国际数学研究中心) Research Institute for Mathematical Sciences, Kyoto University(京都大学数理解析研究所) School of Mathematics, Tianjin University(天津大学数学学院) Zhongguancun Academy(中关村科学院) Department of Mathematics, Stanford University(斯坦福大学数学系) Westlake Institute for Advanced Study, Westlake University(西湖大学西湖高等研究院) School of Mathematical Sciences, Key Laboratory of Intelligent Computing and Applications (Ministry of Education), Tongji University(同济大学数学科学学院(教育部智能计算与应用重点实验室)) Beijing International Center for Mathematical Research and the New Cornerstone Science Laboratory, Peking University(北京大学北京国际数学研究中心与新基石科学实验室) Center for Machine Learning Research, Peking University(北京大学机器学习研究中心) Center for Intelligent Computing, Great Bay Institute for Advanced Study, Great Bay University(大湾区大学高等研究院智能计算中心)

专题命中 数学推理 :reasoning(title,abstract);planning(abstract);verifier(abstract);分类 cs.CL、cs.AI

AI总结 本文针对基于大语言模型的数学推理智能体扩展和编排难题,提出以共享事实图为全局内存管理机制的Danus系统,由主智能体、工作智能体和验证器构成,通过案例研究评估,展示其构建长证明的能力,为长期研究问题提供有效编排途径。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19723 2026-07-09 cs.CL cs.AI 版本更新 84%

Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges

大型语言模型中的数学推理:基准测试、架构、评估与开放挑战

Husnain Amjad, Raja Khurram Shahzad, Aamir Shahzad, Mehwish Fatima

机构 * organization= School of Electrical Engineering Computer Science, National University of Science organization= School of Computing, Data Mathematical Sciences, Western Sydney University, Indonesia organization= Department of Communication, Quality Management Information Systems, Mid Sweden University, Östersund Campus, Sweden

专题命中 数学推理 :reasoning(title,abstract);verifier(abstract);分类 cs.CL、cs.AI

AI总结 本文综述了大型语言模型在数学推理方面的最新进展,通过分析数据集、架构、训练策略和评估协议,探讨了数学推理的基准测试、架构设计、评估方法以及未来的研究挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06855 2026-07-09 cs.LG cs.CL 新提交 81%

Geometric Self-Distillation for Reasoning Generalization

用于推理泛化的几何自蒸馏

Josip Jukić, Ivan Titov

机构 * ILLC, University of Amsterdam(阿姆斯特丹大学伊利沙伯格学院) ILCC, University of Edinburgh(爱丁堡大学伊利沙伯格中心)

专题命中 数学推理 :reasoning(title,abstract);分类 cs.CL、cs.LG

AI总结 研究针对大语言模型特权上下文自蒸馏中监督难以信赖、导致分布外推理能力下降的问题,提出几何自蒸馏目标GeoSD,通过Hellinger损失和近端项对抗漂移,提升了模型分布外推理准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07391 2026-07-09 cs.AI 新提交 79%

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

MIRA-Math:最小信息请求与数学推理基准测试

Charbel Al Bateh, Samer Saab

机构 * Lebanese American University(黎巴嫩美国大学)

专题命中 数学推理 :reasoning(title,abstract);分类 cs.AI

AI总结 介绍MIRA-Math基准测试,用于解决特定数学问题,求解器需按预算请求缺失信息并整合到答案,包含多类实例,实验表明请求成功率和答案准确率可分离,还发布相关工具支持可重复评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07026 2026-07-09 cs.LG 新提交 77%

Constrained Decoding for Diffusion Language Models via Efficient Inference over Finite Automata

通过有限自动机上的高效推理实现扩散语言模型的约束解码

Meihua Dang, Stefano Ermon

机构 * Stanford University(斯坦福大学)

专题命中 数学推理 :reasoning(abstract);math reasoning(abstract);planning(abstract);分类 cs.LG

AI总结 研究扩散语言模型的约束解码问题,提出基于有限自动机高效推理的算法,能保证约束满足,支持多种解码方式,经实验验证在多种任务上提升准确率且推理开销小。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07066 2026-07-09 cs.LG cs.AI math.NT math.RT 新提交 62%

Multiplication Beyond Groups: Stratified Fourier Mechanisms in Transformer Circuits

超越群的乘法:Transformer电路中的分层傅里叶机制

Zitong Andrew Chen, Junaid Hasan, Akhil Srinivasan, Hemkesh Bandi, Jarod Alper

专题命中 数学推理 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 研究小型变压器学习复合模上模整数乘法的方法,提出幺半群扩展,将输入空间划分区域,发现训练的变压器中嵌入、注意力等呈现特定特性,表明相关表示理论机制可超越群扩展到更一般结构。

Comments 29 pages, 15 figures. Spotlight at the Mechanistic Interpretability Workshop at ICML 2026. First three authors contributed equally. Code at https://github.com/uw-math-ai/interpreting-monoids

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06820 2026-07-09 cs.AI 新提交 57%

Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

评估用于计算和实验数学的SageMath增强型大语言模型智能体

Pavel Snopov, German Magai

机构 * The Sage Developers(Sage开发者团队)

专题命中 数学推理 :reasoning(abstract);分类 cs.AI

AI总结 研究聚焦AI在数学领域进展,提出结合大语言模型推理与SageMath反馈的智能体设置,改进RealMath基准。实验表明该设置能提升模型性能,缩小模型差距,CAS增强型智能体在计算探索上有潜力,朝自动猜想发现迈进。

Comments 37 pages, 16 figures, accepted to 3rd AI for Math Workshop at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05835 2026-07-09 math.AG cs.AI math.CO 新提交 57%

Tangent classes of matroids and wonderful compactifications

拟阵的切线类与美妙紧化

Ronnie Cheng, Shurui Liu, Guoxiong Gao

机构 * Stanford University(斯坦福大学)

专题命中 数学推理 :reasoning(abstract);分类 cs.AI

AI总结 本文针对无环拟阵 \(M\) 及特定构建集 \(\mathcal{G}\) 构造整切线类 \(T_{M,\mathcal{G}}^{\mathbb{Z}}\),可恢复周环希尔伯特级数等。主体由人工智能达努斯自主生成,在相关论文公开前解决问题,展现了AI在数学研究中的潜力。

Comments v2: added reference to Danus report [Liu et al., arXiv:2607.06447] and minor edits

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21476 2026-07-09 cs.CL 版本更新 57%

Thinking Seeds: Leveraging Historical Diversity for Position-Aware RL in LLMs

思维种子:在语言模型中利用历史多样性进行位置感知强化学习

Lei Yang, Wei Bi, Chenxi Sun, Renren Jin, Deyi Xiong

机构 * TJUNLP Lab, College of Intelligence and Computing, Tianjin University(天津大学智能与计算学院 TJUNLP 实验室)

专题命中 数学推理 :reasoning(abstract);分类 cs.CL

AI总结 研究语言模型在线强化学习问题,提出思维种子这一令牌级混合策略框架,利用模型历史检查点作离线前缀,通过令牌级重要性比率有效利用历史多样性,实验表明其性能优于标准在线训练和现有离线扩展。

详情

展开后加载摘要…

URL PDF HTML 收藏