arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 1116 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1116 篇

2503.09730 2025-06-25 cs.AI cs.CL cs.LG cs.LO 86%

Local Look-Ahead Guidance via Verifier-in-the-Loop for Automated Theorem Proving

Sara Rajaee, Kumar Pratik, Gabriele Cesa, Arash Behboodi

机构 * Language Technology Lab, University of Amsterdam(阿姆斯特丹大学语言技术实验室) Qualcomm AI Research(高通人工智能研究)

专题命中 代码与定理证明 :verifier(title,abstract);reasoning(abstract,comments);分类 cs.CL、cs.AI、cs.LG;planning(comments)

Comments Accepted at the Findings of ACL 2025, Accepted at ICLR 2025 Workshop on Reasoning and Planning for Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07339 2026-08-13 cs.AI 版本更新 85%

Tools as Continuous Flow for Evolving Agentic Reasoning

工具作为连续流用于演进代理推理

Tairan Huang, Siyu Shang, Qiang Chen, Xiu Su, Yi Chen

机构 * Central South University(中南大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 代码与定理证明 :reasoning(title,abstract);planning(abstract);分类 cs.AI

AI总结 本文提出FlowAgent,通过连续轨迹生成改进代理推理,引入计划级闭环基准,理论证明其鲁棒性和泛化能力,实验证明在长周期推理任务中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13932 2026-06-30 cs.SE cs.AI 85%

Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action

代码推理用于软件工程任务:调查与呼吁行动

Saurabh Pujar, Ira Ceka, Irene Manotas, Gail Kaiser, Baishakhi Ray, Shyam Ramji

机构 * IBM Columbia University(哥伦比亚大学)

专题命中 代码与定理证明 :reasoning(title,abstract);planning(abstract);test-time compute(abstract);分类 cs.AI

AI总结 本文调查代码推理技术,探讨其在软件工程任务中的影响,提出未来研究方向。

Comments Published in Transactions on Machine Learning Research (06/2026) 40 pages, 8 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04120 2026-04-07 cs.CL 85%

Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression

更短,但依然可信?链式推理压缩的实证研究

Lingjie Zeng, Xiaofan Chen, Yanbo Wang, Xiuying Chen

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 代码与定理证明 :chain-of-thought(title,abstract);reasoning(abstract);CoT(abstract);分类 cs.CL

AI总结 本文实证研究了链式推理压缩对模型可信度的影响,评估了不同模型在安全、抗幻觉和多语言鲁棒性上的表现,提出标准化效率评分以揭示信任度权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18916 2026-02-24 cs.MA cs.AI cs.SC 85%

Adaptive Collaboration of Arena-Based Argumentative LLMs for Explainable and Contestable Legal Reasoning

面向可解释和可争议法律推理的 Arena 基础论证 LLM 自适应协作

Hoang-Loc Cao, Phuc Ho, Truong Thanh Hung Nguyen, Phuc Truong Loc Nguyen, Dinh Thien Loc Nguyen, Hung Cao

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.AI

AI总结 ACAL通过自适应多智能体协作和 Arena 基础论证框架,提升法律推理的可解释性和可争议性,实验证明其在效率和透明度上的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24230 2025-06-02 cs.AI 85%

ProofNet++: A Neuro-Symbolic System for Formal Proof Verification with Self-Correction

Murari Ambati

机构 * Austin, USA(美国奥斯汀)

专题命中 代码与定理证明 :self-correction(title,abstract);reasoning(abstract);verifier(abstract);分类 cs.AI

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.01240 2023-03-03 cs.CL 85%

Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought

Abulhair Saparov, He He

专题命中 代码与定理证明 :chain-of-thought(title,abstract);reasoning(abstract);planning(abstract);分类 cs.CL

Comments Published as a conference paper at ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03636 2026-03-20 cs.AI cs.CL cs.LG 85%

CausalARC: Abstract Reasoning with Causal World Models

CausalARC:基于因果世界模型的抽象推理

Jacqueline Maasch, John Kalantari, Kia Khezeli

专题命中 代码与定理证明 :reasoning(title,abstract);logical reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 CausalARC通过因果世界模型进行低数据和分布外推理,结合原则性数据增强提供少量示例反馈,用于评估语言模型在抽象推理、反事实推理、程序合成和因果发现中的表现。

Comments Peer-reviewed workshop paper

Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Bridging Language, Agent, and World Models (LAW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22849 2025-10-28 cs.CL cs.AI cs.LG 85%

Once Upon an Input: Reasoning via Per-Instance Program Synthesis

Adam Stein, Neelay Velingker, Mayur Naik, Eric Wong

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 代码与定理证明 :reasoning(title,abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at NeurIPS 2025. 34 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23145 2025-08-11 cs.PL cs.AI cs.CL cs.LG 85%

CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis

Anjiang Wei, Tarun Suresh, Jiannan Cao, Naveen Kannan, Yuheng Wu, Kai Yan, Thiago S. F. X. Teixeira, Ke Wang, Alex Aiken

机构 * Stanford University(斯坦福大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) MIT(麻省理工学院) Intel(英特尔) Nanjing University(南京大学)

专题命中 代码与定理证明 :reasoning(title,abstract);self-correction(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17352 2026-07-21 cs.AI cs.LG quant-ph 新提交 84%

Self-Modifying Lean Proof Agents with Verifier-Grounded Benchmark Coevolution

基于验证器的基准协同进化的自修改精益证明智能体

Yuqing Li, Zeguan Wu, Yu Gan, Junyu Liu

机构 * University of Pittsburgh(匹兹堡大学)

专题命中 代码与定理证明 :verifier(title,abstract);reasoning(abstract);分类 cs.AI、cs.LG

AI总结 研究设计有效精益证明智能体的挑战,提出自进化精益证明智能体,其工作区可变,与基准协同进化,通过特定更新机制保持分数可比,经实验验证该方法能有效改善精益证明工作流程。

Comments 22 pages, 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02080 2026-06-25 cs.AI cs.FL cs.LG cs.SE 版本更新 84%

The 4/$δ$ Bound: Designing Predictable LLM-Verifier Systems for Formal Method Guarantee

4/δ 界:设计具有形式方法保证的可预测 LLM-验证器系统

Pierre Dantas, Lucas Cordeiro, Youcheng Sun, Waldir Junior

专题命中 代码与定理证明 :verifier(title,abstract);planning(abstract);分类 cs.AI、cs.LG

AI总结 提出 LLM-验证器收敛定理,将多阶段验证建模为顺序吸收马尔可夫链,证明成功概率 δ>0 时系统几乎必然验证,并导出精确延迟上界 4/δ,经 90000+ 试验验证。

Comments 36 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03182 2026-03-20 cs.RO cs.AI cs.CL cs.SC 84%

Simulation to Rules: A Dual-VLM Framework for Formal Visual Planning

模拟到规则:一个双VLM框架用于正式视觉规划

Yilun Hao, Yongchao Chen, Chuchu Fan, Yang Zhang

机构 * MIT(麻省理工学院) Harvard University(哈佛大学) MIT-IBM Watson AI Lab(麻省理工-IBM沃森人工智能实验室)

专题命中 代码与定理证明 :planning(title,abstract);reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出VLMFP框架,结合SimVLM和GenVLM实现自动生成PDDL问题和领域文件,提升视觉规划的精确性和泛化能力。

Comments 40 pages, 6 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22819 2026-03-18 cs.AI cs.FL cs.LG 84%

Hilbert: Recursively Building Formal Proofs with Informal Reasoning

Hilbert:通过非正式推理递归构建形式证明

Sumanth Varambally, Thomas Voice, Yanchao Sun, Zhifeng Chen, Rose Yu, Ke Ye

机构 * UC San Diego(加州大学圣地亚哥分校) Apple(苹果公司)

专题命中 代码与定理证明 :reasoning(title,abstract);verifier(abstract);分类 cs.AI、cs.LG

AI总结 Hilbert结合非正式推理与形式验证,通过递归分解问题并利用反馈优化证明,显著提升在形式证明任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24111 2026-03-02 cs.CV cs.AI cs.CL cs.LO 84%

Toward Guarantees for Clinical Reasoning in Vision Language Models via Formal Verification

通过形式验证为视觉语言模型中的临床推理提供保证

Vikash Singh, Debargha Ganguly, Haotian Yu, Chengwei Zhou, Prerna Singh, Brandon Lee, Vipin Chaudhary, Gourav Datta

机构 * Case Western Reserve University(凯斯西储大学)

专题命中 代码与定理证明 :reasoning(title,abstract);verifier(abstract);分类 cs.CL、cs.AI

AI总结 通过形式验证框架,为视觉语言模型的临床推理提供保证,验证诊断主张的数学推导性,提升生成临床助手的正确性和精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13392 2026-01-21 cs.CL cs.AI cs.FL 84%

Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks

超越记忆:在未见的计算理论任务上测试LLM推理能力

Shlok Shelat, Jay Raval, Souvik Roy, Manas Gaur

机构 * Ahmedabad University Gujarat, India(古吉拉特邦阿赫迈德亚布大学) University of Maryland Baltimore County(马里兰大学巴尔的摩县分校)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 本文通过构建DFA基准测试,揭示LLM在未见计算理论任务上的推理缺陷,发现其在处理复杂约束和语义一致性时存在系统性不足。

Comments 30 pages, 11 figures, 6 tables, Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23726 2025-08-04 cs.AI cs.CL 84%

Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving

Luoxin Chen, Jinming Gu, Liankai Huang, Wenhao Huang, Zhicheng Jiang, Allan Jie, Xiaoran Jin, Xing Jin, Chenggang Li, Kaijing Ma, Cheng Ren, Jiawei Shen, Wenlei Shi, Tong Sun, He Sun, Jiahui Wang, Siran Wang, Zhihong Wang, Chenrui Wei, Shufa Wei, Yonghui Wu, Yuchen Wu, Yihang Xia, Huajian Xin, Fan Yang, Huaiyuan Ying, Hongyi Yuan, Zheng Yuan, Tianyang Zhan, Chi Zhang, Yue Zhang, Ge Zhang, Tianyun Zhao, Jianqiu Zhao, Yichi Zhou, Thomas Hanwen Zhu

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21801 2025-07-21 cs.CL cs.AI 84%

DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition

Z. Z. Ren, Zhihong Shao, Junxiao Song, Huajian Xin, Haocheng Wang, Wanjia Zhao, Liyue Zhang, Zhe Fu, Qihao Zhu, Dejian Yang, Z. F. Wu, Zhibin Gou, Shirong Ma, Hongxuan Tang, Yuxuan Liu, Wenjun Gao, Daya Guo, Chong Ruan

机构 * DeepSeek-AI

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04017 2025-06-23 cs.AI cs.LG cs.LO cs.NE cs.SC 84%

Learning Guided Automated Reasoning: A Brief Survey

Lasse Blaauwbroek, David Cerna, Thibault Gauthier, Jan Jakubův, Cezary Kaliszyk, Martin Suda, Josef Urban

机构 * Czech Technical University in Prague(捷克技术大学布拉格) Radboud University Nijmegen(拉德堡德大学奈杰姆) Czech Academy of Sciences Institute for Computer Science(捷克科学院计算机科学研究所) University of Innsbruck(因斯布鲁克大学)

专题命中 代码与定理证明 :reasoning(title,abstract);logical reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02735 2025-05-06 cs.AI cs.LG 84%

FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models

Zhouliang Yu, Ruotian Peng, Keyi Ding, Yizhe Li, Zhongyuan Peng, Minghao Liu, Yifan Zhang, Zheng Yuan, Huajian Xin, Wenhao Huang, Yandong Wen, Ge Zhang, Weiyang Liu

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

Comments Technical Report v1 (33 pages, 8 figures, project page: https://sphere-ai-lab.github.io/FormalMATH/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06586 2024-06-12 cs.CL cs.AI 84%

Bi-Chainer: Automated Large Language Models Reasoning with Bidirectional Chaining

Shuqi Liu, Bowei He, Linqi Song

专题命中 代码与定理证明 :reasoning(title,abstract);logical reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted by ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16870 2026-08-03 cs.CV cs.AI 版本更新 84%

Demystifying Video Reasoning

揭秘视频推理

Ruisi Wang, Zhongang Cai, Fanyi Pu, Junxiang Xu, Wanqi Yin, Maijunxian Wang, Ran Ji, Chenyang Gu, Bo Li, Ziqi Huang, Hokin Deng, Dahua Lin, Ziwei Liu, Lei Yang

机构 * SenseTime Research(商汤科技研究院) Nanyang Technological University(南洋理工大学) University of California, Berkeley(加州大学伯克利分校) University of California, San Diego(加州大学圣地亚哥分校) Carnegie Mellon University(卡内基梅隆大学)

专题命中 代码与定理证明 :reasoning(title,abstract);self-correction(abstract);分类 cs.AI

AI总结 本文通过实验揭示视频扩散模型中的推理主要发生在去噪步骤中,提出链式步骤(CoS)机制,并发现工作记忆、自我修正和感知先行等涌现行为,最后提出一种无需训练的集成策略来提升推理能力。

Comments Homepage: https://www.wruisi.com/demystifying_video_reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18831 2026-06-17 cs.CL 版本更新 83%

Adaptive Activation Steering for Efficient LLM Reasoning via Closed-Loop PID Control

自适应激活引导:通过闭环PID控制实现高效LLM推理

Aryasomayajula Ram Bharadwaj

机构 * Independent Researcher(独立研究者)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL

AI总结 提出PID-steering方法,利用PID控制器根据块级冗余分类器动态调整激活引导强度,在减少推理开销的同时提升准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11521 2026-06-11 cs.LG 新提交 83%

Counterexample Guided Learning in the Large using Reasoning Agents

使用推理代理的大规模反例引导学习

Hongyi Liu, Frederic Sala, Thomas Reps, Adithya Murali

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 代码与定理证明 :reasoning(title,abstract);verifier(abstract);分类 cs.LG

AI总结 提出反例引导的LLM正则表达式归纳框架,通过验证器反馈和代理策略(如反思与修复循环)显著提升样本效率和复杂任务成功率。

Comments Code, data, and resources are publicly available for research purposes: https://github.com/Lhtie/CEGML

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08011 2026-05-11 cs.AI stat.CO 83%

Abductive Reasoning with Probabilistic Commonsense

基于概率常识的归纳推理

Joseph Cotnareanu, Chiara Roverato, Han Zhou, Didier Chetelat, Yingxue Zhang, Mark Coates

机构 * International Laboratory on Learning Systems(学习系统国际实验室) McGill University(麦吉尔大学) Mila - Quebec Artificial Intelligence Institute(魁北克人工智能研究所) Huawei Noah's Ark Lab(华为诺亚实验室)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI

AI总结 本文提出PACS算法,通过LLM和形式求解器结合,建模常识认知的个体差异,提升大语言模型的归纳推理能力。

Journal ref Proceedings of the International Conference on Machine Learning, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16584 2026-04-21 cs.SE cs.AI cs.PL 83%

Certified Program Synthesis with a Multi-Modal Verifier

带有多模验证器的认证程序合成

Yueyang Feng, Dipesh Kafle, Vladimir Gladshtein, Vitaly Kurin, George Pîrlea, Qiyuan Zhao, Peter Müller, Ilya Sergey

机构 * National University of Singapore(新加坡国立大学) Neapolis University Pafos(纳皮奥斯大学帕福斯)

专题命中 代码与定理证明 :verifier(title,abstract);reasoning(abstract);分类 cs.AI

AI总结 本文提出LeetProof,通过多模验证器解决认证程序合成中的规范缺陷和验证模式碎片化问题,实现更高效的全认证解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21152 2026-03-27 physics.geo-ph cs.AI 83%

TRACE: A Multi-Agent System for Autonomous Physical Reasoning for Seismology

TRACE:一种用于地震学自主物理推理的多智能体系统

Feng Liu, Jian Xu, Xin Cui, Xinghao Wang, Zijie Guo, Jiong Wang, S. Mostafa Mousavi, Xinyu Gu, Hao Chen, Ben Fei, Lihua Fang, Fenghua Ling, Zefeng Li, Lei Bai

机构 * School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University(上海交通大学电子信息与电气工程学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) School of Earth and Space Sciences, University of Science and Technology of China(中国科学技术大学地球和空间科学学院) Department of Earth and Planetary Sciences, Harvard University(哈佛大学地球与行星科学系) Institute of Earthquake Forecasting, China Earthquake Administration(中国地震局地震预报研究所)

专题命中 代码与定理证明 :reasoning(title,abstract);planning(abstract);分类 cs.AI

AI总结 TRACE通过结合大语言模型规划与正式地震学约束,从原始观测中推导出可审计的物理机制,解决了地震序列物理机制推断的挑战,推动了地球科学从专家依赖分析向知识引导的自主发现发展。

Comments 25 pages for main text and 164 pages for appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13443 2026-03-17 cs.SE cs.AI cs.PL 83%

NormCode Canvas: Making LLM Agentic Workflows Development Sustainable via Case-Based Reasoning

NormCode Canvas:通过基于案例的推理使LLM代理工作流开发可持续

Xin Guan, Yunshan Li, Ze Wang

机构 * Psylens AI, China(Psylens AI,中国) Shenzhen University, China(深圳大学,中国) University College London, United Kingdom(伦敦大学学院,英国)

专题命中 代码与定理证明 :reasoning(title,abstract);planning(abstract);分类 cs.AI

AI总结 NormCode Canvas通过基于案例的推理实现多步骤LLM工作流的可持续开发,其核心是NormCode语言,确保每个执行检查点都是自包含的案例,从而消除隐含共享状态带来的问题。

Comments 16 pages, 5 figures. Submitted to ICCBR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01896 2026-03-05 cs.SE cs.AI cs.PL 83%

Agentic Code Reasoning

代理代码推理

Shubham Ugare, Satish Chandra

机构 * Meta, USA(Meta公司)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI

AI总结 本文提出半形式推理方法,通过结构化提示提升代理在代码补丁验证、故障定位和问答任务中的准确性,实现无需执行的代码语义分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17716 2026-01-27 cs.LG 83%

Do Reasoning Models Ask Better Questions? A Formal Information-Theoretic Analysis on Multi-Turn LLM Games

推理模型会提出更好的问题吗?多轮LLM游戏中的信息论分析

Daniel M. Pedrozo, Telma W. de L. Soares, Bryan L. M. de Oliveira

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.LG

AI总结 本文提出多轮对话框架,通过信息增益评估LLM在地理游戏中提问效率,发现具有推理能力的模型在部分可观察环境下表现更优。

Comments Presented at the NeusymBridge Workshop at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏