arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 45006 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1127 篇

2505.12680 2025-10-21 cs.AI cs.CL cs.LG 82%

Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving on Inequalities

Haoyu Zhao, Yihan Geng, Shange Tang, Yong Lin, Bohan Lyu, Hongzhou Lin, Chi Jin, Sanjeev Arora

机构 * Princeton Language and Intelligence(普林斯顿语言与智能) Princeton University(普林斯顿大学) Peking University(北京大学) Tsinghua University(清华大学) Amazon(亚马逊)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments To appear in NeurIPS 2025 Track on Datasets and Benchmarks. 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15748 2025-07-29 cs.CL cs.AI cs.LG 82%

Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models

Shamus Sim, Tyrone Chen

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 25 pages, 7 figures, 3 tables. Conceptualization, both authors. formal analysis, both authors. funding acquisition, both authors. investigation, both authors. resources, both authors. supervision, T.C.. validation, both authors. visualization, both authors. writing original draft, both authors. writing review and editing, both authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08590 2025-06-24 cs.CL cs.AI cs.LG 82%

Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference

Geonhee Kim, Marco Valentino, André Freitas

机构 * Department of Computer Science, University of Manchester, UK(英国曼彻斯特大学计算机科学系) Idiap Research Institute, Switzerland(瑞士Idiap研究所) School of Computer Science, University of Sheffield, UK(英国谢菲尔德大学计算机科学学院) National Biomarker Centre, CRUK-MI, University of Manchester, UK(英国曼彻斯特大学国家生物标志物中心)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15989 2025-06-02 cs.SE 82%

Optimizing Token Consumption in LLMs: A Nano Surge Approach for Code Reasoning Efficiency

Junwei Hu, Weicheng Zheng, Yihan Liu, Yan Liu

专题命中 代码与定理证明 :reasoning(title,abstract);CoT(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14838 2025-03-20 cs.SE 82%

Think Like Human Developers: Harnessing Community Knowledge for Structured Code Reasoning

Chengran Yang, Zhensu Sun, Hong Jin Kang, Jieke Shi, David Lo

专题命中 代码与定理证明 :reasoning(title,abstract);logical reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17190 2024-07-25 cs.CE 82%

Fusing LLMs and KGs for Formal Causal Reasoning behind Financial Risk Contagion

Guanyuan Yu, Xv Wang, Qing Li, Yu Zhao

专题命中 代码与定理证明 :reasoning(title,abstract);logical reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.18120 2024-03-28 cs.AI cs.CL cs.LG 82%

Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization

Jin Peng Zhou, Charles Staats, Wenda Li, Christian Szegedy, Kilian Q. Weinberger, Yuhuai Wu

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06634 2024-02-13 cs.AI cs.CL cs.LG 82%

SocraSynth: Multi-LLM Reasoning with Conditional Statistics

Edward Y. Chang

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 1 figure, 6 tables, 6 appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.02830 2020-10-07 cs.CL cs.AI cs.LG 82%

PRover: Proof Generation for Interpretable Reasoning over Rules

Swarnadeep Saha, Sayan Ghosh, Shashank Srivastava, Mohit Bansal

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP 2020 (15 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09751 2026-06-09 cs.AI cs.CL cs.LO 版本更新 82%

Sound and Complete Neurosymbolic Reasoning with LLM-Grounded Interpretations

基于LLM解释的完备且可靠的神经常识推理

Bradley P. Allen, Prateek Chhikara, Thomas Macaulay Ferguson, Filip Ilievski, Paul Groth

机构 * University of Amsterdam(阿姆斯特丹大学) University of Southern California(南加州大学) Rensselaer Polytechnic Institute(拉特格斯理工学院) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 提出将LLM直接集成到次协调逻辑的语义解释函数中,实现可靠且完备的神经常识推理,在GPQA和SimpleQA基准上宏F1提升约6个百分点,并成功检测药物安全知识库中的矛盾。

Comments 43 pages, 14 tables, 4 figures. Accepted to the 19th Conference on Neurosymbolic Learning and Reasoning (NeSy 2025); to appear Neurosymbolic Artifical Intelligence Special Issue on NeSy 2025 Extended Papers

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15674 2026-03-19 cs.AI cs.IT cs.LG math.IT stat.ML 82%

Theoretical Foundations of Latent Posterior Factors: Formal Guarantees for Multi-Evidence Reasoning

潜在后验因子的理论基础:多证据推理的正式保证

Aliyu Agboola Alege

机构 * Epalea

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种理论框架,通过变分自编码器将多异质证据转换为高斯潜在后验,并利用Sum-Product网络或神经聚合器进行聚合,提供多证据推理的正式保证。

Comments 30 pages, 8 figures, 10 tables. Theoretical characterization of the Latent Posterior Factors (LPF) framework for multi-evidence probabilistic reasoning, with formal guarantees and empirical validation

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16015 2025-06-23 cs.AI cs.CL cs.DB cs.LO math.LO 82%

Bayesian Epistemology with Weighted Authority: A Formal Architecture for Truth-Promoting Autonomous Scientific Reasoning

Craig S. Wright

机构 * Department of Computer Science University of Exeter(计算机科学系埃克塞特大学)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments 91 pages, 0 figures, includes mathematical appendix and formal proofs. Designed as a foundational submission for a modular autonomous epistemic reasoning system. Suitable for logic in computer science, AI epistemology, and scientific informatics

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0009019 2009-11-30 cs.AI cs.CL 82%

Computing Presuppositions by Contextual Reasoning

Christof Monz

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

Comments 5 pages

Journal ref In: P. Brezillon, R. Turner, J-C. Pomerol and E. Turner (Eds.) Proceedings of the AAAI-99 Workshop on Reasoning in Context for AI Applications, AAAI Press, 1999, pp. 75-79

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20895 2026-06-23 cs.AI 新提交 81%

Neurosymbolic Clinical Trial Matching via LLM-Driven Abduction and Logical Verification

基于LLM驱动溯因与逻辑验证的神经符号临床试验匹配

Baiyang Qu, Leonardo Ranaldi, Xi Wang, Marco Valentino

机构 * University of Leicester(莱斯特大学) University of Edinburgh(爱丁堡大学) University of Sheffield(谢菲尔德大学)

专题命中 代码与定理证明 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract);分类 cs.AI

AI总结 提出一种结合大语言模型与逻辑验证的神经符号框架αNeSy-CTM,通过溯因推理处理噪声和不完整临床文本,在临床试验匹配中相对零样本基线提升30%准确率。

Comments 21 pages (including appendix), 5 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01413 2026-08-17 cs.CL cs.AI 版本更新 81%

Adaptive Stopping for Multi-Turn LLM Reasoning

多轮LLM推理中的自适应停止

Xiaofan Zhou, Huy Nguyen, Bo Yu, Chenxi Liu, Lu Cheng

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Augustana College(奥古斯塔纳学院) University of Utah(犹他大学)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出MiCP框架,通过分配不同误差预算实现多轮推理中的自适应停止,同时保证覆盖率,减少轮次、推理成本和预测集规模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10916 2026-08-12 cs.CL cs.AI cs.LO 新提交 81%

FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation

FaithformBench:数学链式思维自动形式化的忠实性基准测试

Rob Cornish, Iacopo Ghinassi, Po-Hung Yeh, Shuqi Liu, Qiyuan Xu, Haoxuan Yin, Dominik Wagner, Wenda Li, Yee Whye Teh, Luke Ong

专题命中 代码与定理证明 :chain-of-thought(title);reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出新基准FaithformBench,通过自动生成扰动推理步骤评估AF系统的忠实性,发现多数AF存在“悄悄修正”无效输入的谄媚现象,当前AF在有效性与无效性保留间存在张力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18350 2026-07-28 cs.CL cs.AI 版本更新 81%

Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis

适配器合并重新激活潜在推理痕迹:一种机制分析

Junyi Zou

机构 * Zjydiary Group(Zjydiary小组)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 研究揭示适配器合并后推理痕迹的重新激活机制,通过几何方法减少泄漏并提升准确性。

Comments Withdrawn by the authors after identifying an implementation error in adapter coefficient scaling that materially affects the main empirical results and invalidates the current conclusions. A corrected reanalysis is in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12733 2026-07-15 cs.AI cs.LG 新提交 81%

LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos

大语言模型能看到烟雾却看不到火焰:用Elenchos评估溯因推理

Julius Steiglechner, Lucas Mahler, Gabriele Lohmann

机构 * Max-Planck-Institute for Biological Cybernetics(马克斯·普朗克生物控制论研究所) University Hospital Tübingen(图宾根大学医院)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 研究大语言模型的溯因推理能力,引入Elenchos评估框架,通过给定形式系统及变异对应物,让智能体判断变异并推断规则修改,发现模型存在检测-归因分离问题,相互作用变异下性能降,推理时间收益递减。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12650 2026-07-15 cs.LG cs.AI cs.CY cs.SE 新提交 81%

Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs

基于证据的可验证智能推理:通过工具验证内核证明消除经验推理中语言模型幻觉的途径

Junyu Ren

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 研究旨在消除语言模型经验推理中的幻觉,提出基于Lean 4的EG-VAR工具调用架构,通过工具验证公理等生成可验证声明,经实验在数值推理等测试中表现良好,定位为高风险经验声明的技术治理接口,可审计相关条件并转化错误为审计目标。

Comments Accepted at the ICML 2026 TAIGR workshop. System name: EG-VAR

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21867 2026-06-23 cs.AI cs.CL cs.SC 新提交 81%

ForEx: A Formal Verification Framework for Explainable Reasoning in Logical Fallacy Detection and Annotation

ForEx:逻辑谬误检测与可解释推理的形式验证框架

Pei-Cing Huang, Chienyu Liu, Chan Hsu, Ci-Siang Chen, Pei-Ju Lee, Yihuang Kang

机构 * Department of Information Management(信息管理系) National Sun Yat-Sen University(国立中山大学) National Chung Hsing University(中国科技大学)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 提出ForEx框架,将LLM生成的解释翻译为Lean4并验证其形式可推导性,通过论证验证矩阵区分标签一致性与形式验证状态,揭示形式可推导性与标签一致性之间的系统性差距。

Comments 2026 IEEE 27th International Conference on Information Reuse and Integration for Data Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19399 2026-06-19 cs.LG cs.AI cs.LO cs.PL 新提交 81%

VERITAS: Verifier-Guided Proof Search for Zero-Shot Formal Theorem Proving

VERITAS:验证器引导的零样本形式定理证明搜索

Manish Acharya, Zhenyu Liao, Yueke Zhang, Kevin Leach, Yu Huang, Yifan Zhang

机构 * Department of Computer Science, Vanderbilt University(范德堡大学计算机科学系) Amazon(亚马逊)

专题命中 代码与定理证明 :verifier(title,abstract);分类 cs.AI、cs.LG

AI总结 提出VERITAS框架,通过两阶段协议(Best-of-N采样+批评引导MCTS)利用验证器反馈进行零样本定理证明,在miniF2F上达40.6%准确率,并发布组合学基准VERITAS-CombiBench。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00555 2026-06-05 cs.AI cs.CL cs.SE 81%

Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems: A Neurosymbolic Architecture for Domain-Grounded AI Agents

企业智能体系统中的本体约束神经推理:一种面向领域 grounded AI 智能体的神经符号架构

Thanh Luong Tuan, Abhijit Sanyal

机构 * Golden Gate University, San Francisco Foundation(金门大学,旧金山基金会) AgenticOS (FAOS)(AgenticOS(FAOS)) Associate Director, Data, Digital & IT Novartis Healthcare Pvt. Ltd.(数据、数字与IT部门,诺华健康有限公司) Novartis Healthcare Pvt. Ltd., Hyderabad, India(诺华健康有限公司,海得拉巴,印度)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种神经符号架构,通过本体约束神经推理解决企业大语言模型在幻觉、领域漂移和无法在推理层面强制执行监管合规性方面的限制,展示了该架构在提升智能体的指标准确性和角色一致性方面的显著效果。

Comments 24 pages, 6 tables, 6 figures, 1 algorithm, 65 references. Replication study: 1,800 runs (600 per model) across 5 regulated industries (3 English, 2 Vietnamese) and 3 LLMs (Claude Sonnet 4, Qwen 2.5 72B, Gemma 4 26B). v3 changes: deep-review trim from 34pp. Code and data: https://github.com/frank-luongt/faos-research/tree/main/RA-3

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18987 2026-05-27 cs.CL cs.AI cs.PL 81%

LLMs versus the Halting Problem: Characterizing Program Termination Reasoning

LLMs 与停机问题:程序终止推理的特征化

Oren Sultan, Jordi Armengol-Estape, Pascal Kesseli, Julien Vanegue, Dafna Shahaf, Yossi Adi, Peter O'Hearn

机构 * FAIR Team, Meta AI(Meta AI FAIR 团队) The Hebrew University of Jerusalem, Israel(耶路撒冷希伯来大学) Bloomberg, New York, USA(彭博社,纽约,美国) Imperial College London, UK(伦敦帝国理工学院,英国) University College London, UK(伦敦大学学院,英国)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本文评估了前沿LLMs在程序终止推理上的能力,发现GPT-5和Claude Sonnet 4.5在C程序终止判断上达到顶级验证工具水平,但无法生成形式化证明,并引入分歧前置条件形式化描述非终止条件。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21851 2026-05-25 cs.LG cs.AI 81%

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning

OPPO: 用于LLM推理中令牌级信用分配的贝叶斯价值递归

Yu Li, Rui Miao, Tian Lan, Zhengling Qi

机构 * George Washington University(乔治华盛顿大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 提出OPPO方法,通过贝叶斯更新累积轨迹级成功概率,实现无需价值网络的令牌级优势估计,在数学、科学和代码推理基准上优于GRPO、DAPO和SDPO。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14049 2026-05-15 cs.AI cs.CL cs.CY 81%

Bridging Legal Interpretation and Formal Logic: Faithfulness, Assumption, and the Future of AI Legal Reasoning

弥合法律解释与形式逻辑:忠实性、假设与AI法律推理的未来

Olivia Peiyu Wang, Leilani H. Gilpin

机构 * University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出一种结合大语言模型与形式验证的神经符号方法,旨在提升AI辅助法律推理的可靠性与可信度,减少人工验证负担。

Comments 2 pages abstract accepted by Bloomberg LSLLAI 2026 Symposium

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02676 2026-05-12 cs.CL cs.AI 81%

ITLC at SemEval-2026 Task 11: Normalization and Deterministic Parsing for Formal Reasoning in LLMs

在SemEval-2026任务11中采用ITLC:为LLMs中的形式推理进行规范化和确定性解析

Wicaksono Leksono Muhamad, Joanito Agili Lopo, Tack Hwa Wong, Muhammad Ravi Shulthan Habibi, Samuel Cahyawijaya

机构 * SEACrowd Mantera Studio Universiti Teknologi PETRONAS(普特拉联合大学) Universitas Indonesia(印度尼西亚大学) Cohere

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出通过显式结构抽象和确定性解析减少LLMs推理中的内容偏差,方法在SemEval-2026任务11中取得前五名,有效降低偏差并提供替代复杂微调的方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21654 2026-05-11 cs.LG cs.AI cs.CC 81%

Limitations on Accurate, Trusted, Human-level Reasoning

对准确、可信、人类水平推理的限制

Rina Panigrahy, Vatsal Sharan

机构 * Google Research(谷歌研究) University of Southern California(南加州大学)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 本文研究了准确、可信和人类水平推理在AI系统中的根本矛盾,证明了准确且可信的系统无法实现人类水平推理。

Comments 19 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05485 2026-05-08 cs.CL cs.AI 81%

ReaComp: Compiling LLM Reasoning into Symbolic Solvers for Efficient Program Synthesis

ReaComp:将LLM推理编译为符号求解器以实现高效的程序合成

Atharva Naik, Yash Mathur, Prakam, Carolyn Rose, David Mortensen

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 通过编译推理轨迹生成符号求解器,提升程序合成效率与准确性,同时减少对LLM的依赖,适用于多个基准测试任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20055 2026-05-05 cs.CL cs.AI 81%

VERGE: Formal Refinement and Guidance Engine for Verifiable LLM Reasoning

VERGE:可验证LLM推理的正式细化和指导引擎

Vikash Singh, Darion Cassel, Nathaniel Weir, Nick Feng, Sam Bayless

机构 * Case Western Reserve University(凯斯西储大学) Amazon Web Services(亚马逊网络服务)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 VERGE结合LLM与SMT求解器,通过迭代细化生成验证引导答案。其通过分解LLM输出为原子声明,自动形式化为一阶逻辑,并利用自动定理证明验证逻辑一致性。引入多模型共识、语义路由和精确逻辑错误定位等创新,提升推理可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00270 2026-05-04 cs.CL cs.AI cs.CY cs.HC 81%

Are You the A-hole? A Fair, Multi-Perspective Ethical Reasoning Framework

你真的是混蛋吗?一种公平、多视角的伦理推理框架

Sheza Munir, Ahanaf Rodoshi, Sumin Lee, Feiran Chang, Xujie Si, Syed Ishtiaque Ahmed

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出一种结合神经语义提取与形式求解器的多视角伦理推理框架,通过MaxSAT解决高冲突领域中的逻辑一致性问题,实验显示其在Reddit论坛中生成的裁决在62%的情况下偏离流行标签,且与人类评估者一致率高达86%。

详情

展开后加载摘要…

URL PDF HTML 收藏