arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 45006 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1127 篇

2604.04120 2026-08-18 cs.CL 版本更新 87%

Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression

更短,但依然可信?链式推理压缩的实证研究

Lingjie Zeng, Xiaofan Chen, Yanbo Wang, Xiuying Chen

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 代码与定理证明 :chain-of-thought(title,abstract);CoT(abstract,abstract_cn);reasoning(abstract);分类 cs.CL

AI总结 本文实证研究了链式推理压缩对模型可信度的影响,评估了不同模型在安全、抗幻觉和多语言鲁棒性上的表现,提出标准化效率评分以揭示信任度权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06269 2026-05-08 q-bio.QM cs.AI 87%

MAT-Cell: A Multi-Agent Tree-Structured Reasoning Framework for Batch-Level Single-Cell Annotation

MAT-Cell: 一种多智能体树状推理框架用于批量单细胞注释

Yehui Yang, Zelin Zang, Xienan Zheng, Yuzhe Jia, Changxi Chi, Jingbo Zhou, Chang Yu, Jinlin Wu, Fuji Yang, Jiebo Luo, Zhen Lei, Stan Z. Li

机构 * Westlake University(西湖大学) Shenzhen University of Advanced Technology(深圳先进技术大学) Center for Artificial Intelligence and Robotics(人工智能与机器人中心) Hong Kong Institute of Science and Innovation(香港科学与创新研究院) Chinese Academy of Sciences(中国科学院) University Key Laboratory of Information and Communication Security Backup and Recovery(信息与通信安全备份与恢复大学重点实验室)

专题命中 代码与定理证明 :reasoning(title,abstract);CoT(abstract,abstract_cn);verifier(abstract);分类 cs.AI

AI总结 MAT-Cell通过分离证据基础与标签决策,结合反向验证查询和多轮辩论,提升批量单细胞注释的准确性与可追溯性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13069 2026-07-16 cs.AI cs.CL cs.LO 新提交 87%

Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution

介入式基础审计:通过谓词替换对大语言模型思维链进行黑盒前提依赖性测试

Hironao Nakamura

专题命中 代码与定理证明 :chain-of-thought(title,abstract);reasoning(abstract,comments);CoT(abstract);分类 cs.CL、cs.AI

AI总结 研究大语言模型思维链前提依赖性问题,提出介入式基础审计方法,通过谓词替换干预前提并重新运行模型,在ProntoQA基准测试中检测前提依赖性表现优异,还发现了“正确答案,错误推理”信号。

Comments Accepted at the ICLR 2026 Workshop on Logical Reasoning of Large Language Models (https://iclr.cc/virtual/2026/10017466)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05318 2025-12-08 cs.CL cs.AI cs.LG 87%

To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples

思考还是不思考:过度使用Co-T示例的元训练隐藏成本

Vignesh Kothapalli, Ata Fatahibaarzi, Hamed Firooz, Maziar Sanjabi

机构 * Stanford University(斯坦福大学) LinkedIn AI

专题命中 代码与定理证明 :CoT(title,abstract);reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出CoT-Recipe方法,通过调节元训练序列中CoT和非CoT示例的比例,提升大型语言模型在新任务上的推理准确性,实验显示在无CoT示例时准确率可提升300%。

Comments 26 pages, 45 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04592 2025-06-06 cs.CL cs.AI cs.LG 87%

Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification

Chengwu Liu, Ye Yuan, Yichun Yin, Yan Xu, Xin Xu, Zaoyu Chen, Yasheng Wang, Lifeng Shang, Qun Liu, Ming Zhang

机构 * School of Computer Science, National Key Laboratory for Multimedia Information Processing, PKU-Anker LLM Lab, Peking University(计算机学院、多媒体信息处理国家重点实验室、PKU-Anker LLM实验室、北京大学) Huawei Noah’s Ark Lab(华为诺亚实验室) The Hong Kong University of Science and Technology(香港科技大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted in ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17053 2026-03-13 cs.CL cs.AI cs.DB 86%

Knowledge Distillation with Structured Chain-of-Thought for Text-to-SQL

基于结构化思维链的知识蒸馏用于文本到SQL

Khushboo Thaker, Yony Bresler

专题命中 代码与定理证明 :chain-of-thought(title,abstract);reasoning(abstract);CoT(abstract);分类 cs.CL、cs.AI

AI总结 本文提出Struct-SQL框架,通过结构化思维链知识蒸馏提升小型语言模型在文本到SQL任务中的性能,实现8.1%的准确率提升。

Comments Accepted at the 39th Canadian Conference on Artificial Intelligence (Canadian AI 2026). This is the extended version containing additional details and appendices omitted from the camera-ready proceedings due to space constraints

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03613 2025-08-06 cs.LG cs.AI 86%

Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction

Yong Lin, Shange Tang, Bohan Lyu, Ziran Yang, Jui-Hui Chung, Haoyu Zhao, Lai Jiang, Yihan Geng, Jiawei Ge, Jingruo Sun, Jiayun Wu, Jiri Gesi, Ximing Lu, David Acuna, Kaiyu Yang, Hongzhou Lin, Yejin Choi, Danqi Chen, Sanjeev Arora, Chi Jin

专题命中 代码与定理证明 :self-correction(title,abstract);test-time compute(abstract);verifier(abstract);分类 cs.AI、cs.LG

Comments 24 pages, 10 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07106 2025-08-01 cs.CL cs.AI 86%

Theorem-of-Thought: A Multi-Agent Framework for Abductive, Deductive, and Inductive Reasoning in Language Models

Samir Abdaljalil, Hasan Kurban, Khalid Qaraqe, Erchin Serpedin

机构 * Texas A&M University(德克萨斯大学) Hamad Bin Khalifa University(哈马德·本·卡伊夫大学)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments ACL 2025 KnowFM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14431 2025-05-29 cs.CL cs.LG 86%

Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains

Xu Chu, Zhijie Tan, Hanlin Xue, Guanyu Wang, Tong Mo, Weiping Li

机构 * School of Software and Microelectronics, Peking University(软件与微电子学院,北京大学)

专题命中 代码与定理证明 :reasoning(title,abstract);CoT(abstract);self-correction(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15120 2025-02-24 cs.CL cs.AI 86%

Unveiling Reasoning Thresholds in Language Models: Scaling, Fine-Tuning, and Interpretability through Attention Maps

Yen-Che Hsiao, Abhishek Dutta

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10007 2024-09-17 cs.CL cs.AI 86%

SelECT-SQL: Self-correcting ensemble Chain-of-Thought for Text-to-SQL

Ke Shen, Mayank Kejriwal

专题命中 代码与定理证明 :chain-of-thought(title,abstract);CoT(abstract);self-correction(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.13894 2023-05-30 cs.AI cs.LG 86%

LAMBADA: Backward Chaining for Automated Reasoning in Natural Language

Mehran Kazemi, Najoung Kim, Deepti Bhatia, Xin Xu, Deepak Ramachandran

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);logical reasoning(abstract);分类 cs.AI、cs.LG

Comments Accepted at ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19834 2026-01-28 cs.AI 86%

Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models

视觉生成通过多模态世界模型解锁类人推理

Jialong Wu, Xiaoying Zhang, Hongyi Yuan, Xiangcheng Zhang, Tianhao Huang, Changjing He, Chaoyi Deng, Renrui Zhang, Youbin Wu, Mingsheng Long

机构 * Tsinghua University(清华大学)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.AI

AI总结 本文提出视觉生成在特定任务中优于纯语言推理,通过构建VisWorld-Eval评估套件验证了多模态世界模型提升类人推理的能力。

Comments Project page: https://thuml.github.io/Reasoning-Visual-World

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09730 2025-06-25 cs.AI cs.CL cs.LG cs.LO 86%

Local Look-Ahead Guidance via Verifier-in-the-Loop for Automated Theorem Proving

Sara Rajaee, Kumar Pratik, Gabriele Cesa, Arash Behboodi

机构 * Language Technology Lab, University of Amsterdam(阿姆斯特丹大学语言技术实验室) Qualcomm AI Research(高通人工智能研究)

专题命中 代码与定理证明 :verifier(title,abstract);reasoning(abstract,comments);分类 cs.CL、cs.AI、cs.LG;planning(comments)

Comments Accepted at the Findings of ACL 2025, Accepted at ICLR 2025 Workshop on Reasoning and Planning for Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07339 2026-08-13 cs.AI 版本更新 85%

Tools as Continuous Flow for Evolving Agentic Reasoning

工具作为连续流用于演进代理推理

Tairan Huang, Siyu Shang, Qiang Chen, Xiu Su, Yi Chen

机构 * Central South University(中南大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 代码与定理证明 :reasoning(title,abstract);planning(abstract);分类 cs.AI

AI总结 本文提出FlowAgent,通过连续轨迹生成改进代理推理,引入计划级闭环基准,理论证明其鲁棒性和泛化能力,实验证明在长周期推理任务中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13932 2026-06-30 cs.SE cs.AI 85%

Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action

代码推理用于软件工程任务:调查与呼吁行动

Saurabh Pujar, Ira Ceka, Irene Manotas, Gail Kaiser, Baishakhi Ray, Shyam Ramji

机构 * IBM Columbia University(哥伦比亚大学)

专题命中 代码与定理证明 :reasoning(title,abstract);planning(abstract);test-time compute(abstract);分类 cs.AI

AI总结 本文调查代码推理技术,探讨其在软件工程任务中的影响,提出未来研究方向。

Comments Published in Transactions on Machine Learning Research (06/2026) 40 pages, 8 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18916 2026-02-24 cs.MA cs.AI cs.SC 85%

Adaptive Collaboration of Arena-Based Argumentative LLMs for Explainable and Contestable Legal Reasoning

面向可解释和可争议法律推理的 Arena 基础论证 LLM 自适应协作

Hoang-Loc Cao, Phuc Ho, Truong Thanh Hung Nguyen, Phuc Truong Loc Nguyen, Dinh Thien Loc Nguyen, Hung Cao

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.AI

AI总结 ACAL通过自适应多智能体协作和 Arena 基础论证框架,提升法律推理的可解释性和可争议性,实验证明其在效率和透明度上的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24230 2025-06-02 cs.AI 85%

ProofNet++: A Neuro-Symbolic System for Formal Proof Verification with Self-Correction

Murari Ambati

机构 * Austin, USA(美国奥斯汀)

专题命中 代码与定理证明 :self-correction(title,abstract);reasoning(abstract);verifier(abstract);分类 cs.AI

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.01240 2023-03-03 cs.CL 85%

Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought

Abulhair Saparov, He He

专题命中 代码与定理证明 :chain-of-thought(title,abstract);reasoning(abstract);planning(abstract);分类 cs.CL

Comments Published as a conference paper at ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03636 2026-03-20 cs.AI cs.CL cs.LG 85%

CausalARC: Abstract Reasoning with Causal World Models

CausalARC:基于因果世界模型的抽象推理

Jacqueline Maasch, John Kalantari, Kia Khezeli

专题命中 代码与定理证明 :reasoning(title,abstract);logical reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 CausalARC通过因果世界模型进行低数据和分布外推理,结合原则性数据增强提供少量示例反馈,用于评估语言模型在抽象推理、反事实推理、程序合成和因果发现中的表现。

Comments Peer-reviewed workshop paper

Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Bridging Language, Agent, and World Models (LAW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22849 2025-10-28 cs.CL cs.AI cs.LG 85%

Once Upon an Input: Reasoning via Per-Instance Program Synthesis

Adam Stein, Neelay Velingker, Mayur Naik, Eric Wong

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 代码与定理证明 :reasoning(title,abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at NeurIPS 2025. 34 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23145 2025-08-11 cs.PL cs.AI cs.CL cs.LG 85%

CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis

Anjiang Wei, Tarun Suresh, Jiannan Cao, Naveen Kannan, Yuheng Wu, Kai Yan, Thiago S. F. X. Teixeira, Ke Wang, Alex Aiken

机构 * Stanford University(斯坦福大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) MIT(麻省理工学院) Intel(英特尔) Nanjing University(南京大学)

专题命中 代码与定理证明 :reasoning(title,abstract);self-correction(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17352 2026-07-21 cs.AI cs.LG quant-ph 新提交 84%

Self-Modifying Lean Proof Agents with Verifier-Grounded Benchmark Coevolution

基于验证器的基准协同进化的自修改精益证明智能体

Yuqing Li, Zeguan Wu, Yu Gan, Junyu Liu

机构 * University of Pittsburgh(匹兹堡大学)

专题命中 代码与定理证明 :verifier(title,abstract);reasoning(abstract);分类 cs.AI、cs.LG

AI总结 研究设计有效精益证明智能体的挑战,提出自进化精益证明智能体,其工作区可变,与基准协同进化,通过特定更新机制保持分数可比,经实验验证该方法能有效改善精益证明工作流程。

Comments 22 pages, 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02080 2026-06-25 cs.AI cs.FL cs.LG cs.SE 版本更新 84%

The 4/$δ$ Bound: Designing Predictable LLM-Verifier Systems for Formal Method Guarantee

4/δ 界:设计具有形式方法保证的可预测 LLM-验证器系统

Pierre Dantas, Lucas Cordeiro, Youcheng Sun, Waldir Junior

专题命中 代码与定理证明 :verifier(title,abstract);planning(abstract);分类 cs.AI、cs.LG

AI总结 提出 LLM-验证器收敛定理,将多阶段验证建模为顺序吸收马尔可夫链,证明成功概率 δ>0 时系统几乎必然验证,并导出精确延迟上界 4/δ,经 90000+ 试验验证。

Comments 36 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03182 2026-03-20 cs.RO cs.AI cs.CL cs.SC 84%

Simulation to Rules: A Dual-VLM Framework for Formal Visual Planning

模拟到规则:一个双VLM框架用于正式视觉规划

Yilun Hao, Yongchao Chen, Chuchu Fan, Yang Zhang

机构 * MIT(麻省理工学院) Harvard University(哈佛大学) MIT-IBM Watson AI Lab(麻省理工-IBM沃森人工智能实验室)

专题命中 代码与定理证明 :planning(title,abstract);reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出VLMFP框架,结合SimVLM和GenVLM实现自动生成PDDL问题和领域文件,提升视觉规划的精确性和泛化能力。

Comments 40 pages, 6 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22819 2026-03-18 cs.AI cs.FL cs.LG 84%

Hilbert: Recursively Building Formal Proofs with Informal Reasoning

Hilbert:通过非正式推理递归构建形式证明

Sumanth Varambally, Thomas Voice, Yanchao Sun, Zhifeng Chen, Rose Yu, Ke Ye

机构 * UC San Diego(加州大学圣地亚哥分校) Apple(苹果公司)

专题命中 代码与定理证明 :reasoning(title,abstract);verifier(abstract);分类 cs.AI、cs.LG

AI总结 Hilbert结合非正式推理与形式验证,通过递归分解问题并利用反馈优化证明,显著提升在形式证明任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24111 2026-03-02 cs.CV cs.AI cs.CL cs.LO 84%

Toward Guarantees for Clinical Reasoning in Vision Language Models via Formal Verification

通过形式验证为视觉语言模型中的临床推理提供保证

Vikash Singh, Debargha Ganguly, Haotian Yu, Chengwei Zhou, Prerna Singh, Brandon Lee, Vipin Chaudhary, Gourav Datta

机构 * Case Western Reserve University(凯斯西储大学)

专题命中 代码与定理证明 :reasoning(title,abstract);verifier(abstract);分类 cs.CL、cs.AI

AI总结 通过形式验证框架,为视觉语言模型的临床推理提供保证,验证诊断主张的数学推导性,提升生成临床助手的正确性和精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13392 2026-01-21 cs.CL cs.AI cs.FL 84%

Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks

超越记忆:在未见的计算理论任务上测试LLM推理能力

Shlok Shelat, Jay Raval, Souvik Roy, Manas Gaur

机构 * Ahmedabad University Gujarat, India(古吉拉特邦阿赫迈德亚布大学) University of Maryland Baltimore County(马里兰大学巴尔的摩县分校)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 本文通过构建DFA基准测试,揭示LLM在未见计算理论任务上的推理缺陷,发现其在处理复杂约束和语义一致性时存在系统性不足。

Comments 30 pages, 11 figures, 6 tables, Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23726 2025-08-04 cs.AI cs.CL 84%

Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving

Luoxin Chen, Jinming Gu, Liankai Huang, Wenhao Huang, Zhicheng Jiang, Allan Jie, Xiaoran Jin, Xing Jin, Chenggang Li, Kaijing Ma, Cheng Ren, Jiawei Shen, Wenlei Shi, Tong Sun, He Sun, Jiahui Wang, Siran Wang, Zhihong Wang, Chenrui Wei, Shufa Wei, Yonghui Wu, Yuchen Wu, Yihang Xia, Huajian Xin, Fan Yang, Huaiyuan Ying, Hongyi Yuan, Zheng Yuan, Tianyang Zhan, Chi Zhang, Yue Zhang, Ge Zhang, Tianyun Zhao, Jianqiu Zhao, Yichi Zhou, Thomas Hanwen Zhu

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21801 2025-07-21 cs.CL cs.AI 84%

DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition

Z. Z. Ren, Zhihong Shao, Junxiao Song, Huajian Xin, Haocheng Wang, Wanjia Zhao, Liyue Zhang, Zhe Fu, Qihao Zhu, Dejian Yang, Z. F. Wu, Zhibin Gou, Shirong Ma, Hongxuan Tang, Yuxuan Liu, Wenjun Gao, Daya Guo, Chong Ruan

机构 * DeepSeek-AI

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏