arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 1117 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1117 篇

2503.22738 2025-12-01 cs.LG cs.CR 83%

ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning

ShieldAgent: 通过可验证的安全策略推理实现安全代理

Zhaorun Chen, Mintong Kang, Bo Li

机构 * University of Chicago, Chicago IL, USA(芝加哥大学) University of Illinois at Urbana-Champaign, Champaign IL, USA(伊利诺伊大学厄巴纳-香槟分校)

专题命中 代码与定理证明 :reasoning(title,abstract);logical reasoning(abstract);分类 cs.LG

AI总结 ShieldAgent通过逻辑推理实现对代理动作轨迹的安全策略合规性,有效提升代理防护性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23686 2025-10-22 cs.CL cs.PL cs.SE 83%

Evaluating Program Semantics Reasoning with Type Inference in System F

Yifeng He, Luning Yang, Christopher Castro Gaw Gonzalo, Hao Chen

机构 * University of California, Davis(加州大学戴维斯分校) University of Hong Kong(香港大学)

专题命中 代码与定理证明 :reasoning(title,abstract);test-time compute(abstract);分类 cs.CL

Comments NeurIPS '25, package released at: https://github.com/SecurityLab-UCD/TF-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12164 2025-10-15 cs.CL 83%

A Survey on Parallel Reasoning

Ziqi Wang, Boye Niu, Zipeng Gao, Zhi Zheng, Tong Xu, Linghui Meng, Zhongli Li, Jing Liu, Yilong Chen, Chen Zhu, Hua Wu, Haifeng Wang, Enhong Chen

机构 * USTC(中国科学技术大学) Baidu(百度) USYD(悉尼大学)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14119 2025-10-14 cs.AI cs.SE 83%

CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning

Man Ho Lam, Chaozheng Wang, Jen-tse Huang, Michael R. Lyu

机构 * The Chinese University of Hong Kong(香港中文大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI

Comments NeurIPS 2025; 10 pages of main text; 25 pages of appendices. Website - https://cuhk-arise.github.io/CodeCrash/

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11461 2025-10-09 cs.AI 83%

Hint of Pseudo Code (HoPC): Zero-Shot Step by Step Pseudo Code Reasoning Prompting

Iok Tong Lei, Ziyu Zhu, Han Yu, Yige Yao, Zhidong Deng

机构 * Department of Computer Science and Technology(计算机科学与技术系) Institute for Artificial Intelligence at Tsinghua University(清华大学人工智能研究院) Beijing National Research Center for Information Science and Technology(北京信息科学国家研究中心) Tsinghua University(清华大学)

专题命中 代码与定理证明 :reasoning(title,abstract);CoT(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15739 2025-09-22 cs.CL 83%

Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics

Reza Sanayei, Srdjan Vesic, Eduardo Blanco, Mihai Surdeanu

机构 * Department of Computer Science, University of Arizona(亚利桑那大学计算机科学系) CRIL CNRS & University of Artois(CNRS CRIL与阿维尼昂大学)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00963 2025-06-12 cs.LG 83%

PDE-Controller: LLMs for Autoformalization and Reasoning of PDEs

Mauricio Soroco, Jialin Song, Mengzhou Xia, Kye Emond, Weiran Sun, Wuyang Chen

机构 * Princeton University(普林斯顿大学) Simon Fraser University(西蒙弗雷泽大学) School of Computing Science, Simon Fraser University(西蒙弗雷泽大学计算科学学院)

专题命中 代码与定理证明 :reasoning(title,abstract);planning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17481 2025-05-26 cs.CL 83%

MARCO: Meta-Reflection with Cross-Referencing for Code Reasoning

Yusheng Zhao, Xiao Luo, Weizhi Zhang, Wei Ju, Zhiping Xiao, Philip S. Yu, Ming Zhang

机构 * Peking University(北京大学) University of California, Los Angeles(加州大学洛杉矶分校) University of Illinois Chicago(伊利诺伊大学香槟分校) University of Washington(华盛顿大学)

专题命中 代码与定理证明 :reasoning(title,abstract);logical reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06176 2026-02-09 cs.AI cs.CL cs.LG 83%

Large Language Model Reasoning Failures

大语言模型推理失败

Peiyang Song, Pengrui Han, Noah Goodman

机构 * California Institute of Technology(加州理工学院) Stanford University(斯坦福大学) Carleton College(卡尔顿学院)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文首次系统调查大语言模型推理失败问题,提出分类框架并分析其根本原因,旨在提升模型的推理能力与鲁棒性。

Comments Repository: https://github.com/Peiyang-Song/Awesome-LLM-Reasoning-Failures. Published at TMLR 2026 with Survey Certification

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17270 2024-10-24 cs.AI cs.CL cs.LG cs.LO cs.NE 83%

Proof of Thought : Neurosymbolic Program Synthesis allows Robust and Interpretable Reasoning

Debargha Ganguly, Srinivasan Iyengar, Vipin Chaudhary, Shivkumar Kalyanaraman

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 38th Conference on Neural Information Processing Systems (NeurIPS 2024) System 2 Reasoning At Scale Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06690 2026-05-11 cs.AI cs.CL cs.LG 82%

State Representation and Termination for Recursive Reasoning Systems

递归推理系统的状态表示与终止

Debashis Guha, Amritendu Mukherjee, Sanjay Kukreja, Tarun Kumar

机构 * S P Jain School of Global Management(S P Jain 全球管理学院) Indian Statistical Institute(印度统计研究所) eClerx Services Ltd.(eClerx 服务有限公司)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出了一种递归推理系统的状态表示方法及终止条件,通过epistemic状态图编码提取的主张、证据关系、开放问题和置信度权重,并定义了顺序间隙以判断迭代的必要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12059 2026-04-30 cs.CL cs.AI cs.LG 82%

MeTHanol: Modularized Thinking Language Models with Intermediate Layer Thinking, Decoding and Bootstrapping Reasoning

MeTHanol:模块化思维语言模型与中间层思维、解码与推理bootstrap

Ningyuan Xi, Xiaoyu Wang, Yetao Wu, Teng Chen, Qingqing Gu, Yue Zhao, Jinxian Qu, Zhonglin Jiang, Yong Chen, Luo Ji

机构 * Beihang University(北航) Beijing Institute of Technology(北京理工大学) Geely AI Lab(吉利人工智能实验室)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出MeTHanol模块化思维语言模型,通过中间层思维解码与双阶段推理提升LLM的认知能力,实验表明其在理论思维和 vignette 任务中表现出色,能规划、反思并生成类人回答。

Comments 19 pages, 7 figures. IJCNN2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20917 2026-04-24 cs.LG cs.AI cs.CL cs.PL cs.SE 82%

The Path Not Taken: Duality in Reasoning about Program Execution

未选择的道路:程序执行推理中的对偶性

Eshgin Hasanov, Md Mahadi Hassan Sibat, Santu Karmaker, Aashish Yadavally

机构 * Department of Computer Science, University of Central Florida(中央佛罗里达大学计算机科学系)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出通过双重推理任务评估程序执行的因果理解,提出DexBench基准测试集,评估13种LLM,证明双路径推理能有效提升动态代码理解能力。

Comments Accepted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19558 2026-04-22 cs.SE 82%

On Reasoning-Centric LLM-based Automated Theorem Proving

基于推理的LLM自动定理证明方法

Yican Sun, Chengwei Shi, Hangzhou Lyu, Yingfei Xiong

专题命中 代码与定理证明 :reasoning(title,abstract);planning(abstract)

AI总结 本文提出ReCent-Prover,通过引入反思验证和规划检索技术,提升自动定理证明能力,在CoqStoq基准测试中实现22.58%的定理证明提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17010 2026-04-21 cs.CL cs.AI cs.LG cs.PL 82%

Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification

通过形式验证的语义等价自博弈提升大语言模型代码推理

Antonio Valerio Miceli Barone, Poon Tsz Nok

机构 * School of Informatics University of Edinburgh(信息学院爱丁堡大学)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出一种基于形式验证的语义等价自博弈框架,利用对抗训练提升代码推理能力,通过验证等价性与非等价性,释放OpInstruct-HSx数据集,实验显示在EquiBench上提升13.3个百分点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02709 2026-04-16 cs.CL cs.AI cs.LG cs.SE 82%

Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy

通过乔姆斯基层级评估大语言模型的正式推理能力

Yihong Dong, Jianha Xiao, Xue Jiang, Xuyuan Guo, Zhiyuan Fan, Jiaru Qian, Kechi Zhang, Jia Li, Zhi Jin, Ge Li

机构 * Peking University(北京大学) Wuhan University(武汉大学)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出ChomskyBench基准,通过乔姆斯基层级系统评估大语言模型的正式推理能力,揭示其在结构化语言处理中的性能分层及效率瓶颈。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13962 2026-02-17 cs.SE 82%

CodeGlance: Understanding Code Reasoning Challenges in LLMs through Multi-Dimensional Feature Analysis

CodeGlance:通过多维特征分析理解LLM中的代码推理挑战

Yunkun Wang, Xuanhe Zhang, Junxiao Han, Chen Zhi, Shuiguang Deng

专题命中 代码与定理证明 :reasoning(title,abstract);CoT(abstract)

AI总结 CodeGlance通过多维特征分析揭示LLM在代码推理中的挑战,发现未见过的函数推理对小型模型尤为困难,提出改进策略并提供开发指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12680 2025-10-21 cs.AI cs.CL cs.LG 82%

Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving on Inequalities

Haoyu Zhao, Yihan Geng, Shange Tang, Yong Lin, Bohan Lyu, Hongzhou Lin, Chi Jin, Sanjeev Arora

机构 * Princeton Language and Intelligence(普林斯顿语言与智能) Princeton University(普林斯顿大学) Peking University(北京大学) Tsinghua University(清华大学) Amazon(亚马逊)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments To appear in NeurIPS 2025 Track on Datasets and Benchmarks. 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15748 2025-07-29 cs.CL cs.AI cs.LG 82%

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models

Shamus Sim, Tyrone Chen

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 25 pages, 7 figures, 3 tables. Conceptualization, both authors. formal analysis, both authors. funding acquisition, both authors. investigation, both authors. resources, both authors. supervision, T.C.. validation, both authors. visualization, both authors. writing original draft, both authors. writing review and editing, both authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08590 2025-06-24 cs.CL cs.AI cs.LG 82%

Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference

Geonhee Kim, Marco Valentino, André Freitas

机构 * Department of Computer Science, University of Manchester, UK(英国曼彻斯特大学计算机科学系) Idiap Research Institute, Switzerland(瑞士Idiap研究所) School of Computer Science, University of Sheffield, UK(英国谢菲尔德大学计算机科学学院) National Biomarker Centre, CRUK-MI, University of Manchester, UK(英国曼彻斯特大学国家生物标志物中心)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15989 2025-06-02 cs.SE 82%

Optimizing Token Consumption in LLMs: A Nano Surge Approach for Code Reasoning Efficiency

Junwei Hu, Weicheng Zheng, Yihan Liu, Yan Liu

专题命中 代码与定理证明 :reasoning(title,abstract);CoT(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14838 2025-03-20 cs.SE 82%

Think Like Human Developers: Harnessing Community Knowledge for Structured Code Reasoning

Chengran Yang, Zhensu Sun, Hong Jin Kang, Jieke Shi, David Lo

专题命中 代码与定理证明 :reasoning(title,abstract);logical reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17190 2024-07-25 cs.CE 82%

Fusing LLMs and KGs for Formal Causal Reasoning behind Financial Risk Contagion

Guanyuan Yu, Xv Wang, Qing Li, Yu Zhao

专题命中 代码与定理证明 :reasoning(title,abstract);logical reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.18120 2024-03-28 cs.AI cs.CL cs.LG 82%

Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization

Jin Peng Zhou, Charles Staats, Wenda Li, Christian Szegedy, Kilian Q. Weinberger, Yuhuai Wu

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06634 2024-02-13 cs.AI cs.CL cs.LG 82%

SocraSynth: Multi-LLM Reasoning with Conditional Statistics

Edward Y. Chang

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 1 figure, 6 tables, 6 appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.02830 2020-10-07 cs.CL cs.AI cs.LG 82%

PRover: Proof Generation for Interpretable Reasoning over Rules

Swarnadeep Saha, Sayan Ghosh, Shashank Srivastava, Mohit Bansal

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP 2020 (15 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09751 2026-06-09 cs.AI cs.CL cs.LO 版本更新 82%

Sound and Complete Neurosymbolic Reasoning with LLM-Grounded Interpretations

基于LLM解释的完备且可靠的神经常识推理

Bradley P. Allen, Prateek Chhikara, Thomas Macaulay Ferguson, Filip Ilievski, Paul Groth

机构 * University of Amsterdam(阿姆斯特丹大学) University of Southern California(南加州大学) Rensselaer Polytechnic Institute(拉特格斯理工学院) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 提出将LLM直接集成到次协调逻辑的语义解释函数中,实现可靠且完备的神经常识推理,在GPQA和SimpleQA基准上宏F1提升约6个百分点,并成功检测药物安全知识库中的矛盾。

Comments 43 pages, 14 tables, 4 figures. Accepted to the 19th Conference on Neurosymbolic Learning and Reasoning (NeSy 2025); to appear Neurosymbolic Artifical Intelligence Special Issue on NeSy 2025 Extended Papers

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15674 2026-03-19 cs.AI cs.IT cs.LG math.IT stat.ML 82%

Theoretical Foundations of Latent Posterior Factors: Formal Guarantees for Multi-Evidence Reasoning

潜在后验因子的理论基础:多证据推理的正式保证

Aliyu Agboola Alege

机构 * Epalea

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种理论框架,通过变分自编码器将多异质证据转换为高斯潜在后验,并利用Sum-Product网络或神经聚合器进行聚合,提供多证据推理的正式保证。

Comments 30 pages, 8 figures, 10 tables. Theoretical characterization of the Latent Posterior Factors (LPF) framework for multi-evidence probabilistic reasoning, with formal guarantees and empirical validation

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16015 2025-06-23 cs.AI cs.CL cs.DB cs.LO math.LO 82%

Bayesian Epistemology with Weighted Authority: A Formal Architecture for Truth-Promoting Autonomous Scientific Reasoning

Craig S. Wright

机构 * Department of Computer Science University of Exeter(计算机科学系埃克塞特大学)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments 91 pages, 0 figures, includes mathematical appendix and formal proofs. Designed as a foundational submission for a modular autonomous epistemic reasoning system. Suitable for logic in computer science, AI epistemology, and scientific informatics

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0009019 2009-11-30 cs.AI cs.CL 82%

Computing Presuppositions by Contextual Reasoning

Christof Monz

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments 5 pages

Journal ref In: P. Brezillon, R. Turner, J-C. Pomerol and E. Turner (Eds.) Proceedings of the AAAI-99 Workshop on Reasoning in Context for AI Applications, AAAI Press, 1999, pp. 75-79

详情

展开后加载摘要…

URL PDF HTML 收藏