arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 1116 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1116 篇

2608.03291 2026-08-05 cs.LG cs.AI cs.CL 新提交 93%

The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics

可察觉的轨迹:利用思维链动态检测大语言模型中的推理失败

Shashwat Sourav, Aishwarya Balwani

专题命中 代码与定理证明 :chain-of-thought(title,abstract);CoT(summary_cn,abstract);reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究利用思维链动态特性,在不假设语言化CoT语义忠实性的情况下,检测大语言模型在布尔可满足性任务中的分布式推理失败,并通过针对性提示干预提升了Llama3-70B的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14089 2025-09-16 cs.CL cs.AI cs.LG 91%

LogicTree: Structured Proof Exploration for Coherent and Rigorous Logical Reasoning with Large Language Models

Kang He, Kaushik Roy

机构 * Electrical and Computer Engineering, Purdue University(电子与计算机工程系,普渡大学)

专题命中 代码与定理证明 :reasoning(title,abstract);logical reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract)

Comments EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10480 2026-07-14 cs.CL 新提交 90%

When Reasoning Hurts Legal Drafting: The Verbalization Bottleneck in Patent Claim Generation

当推理对法律起草造成阻碍时:专利权利要求生成中的语言表达瓶颈

Lekang Jiang, Wenjun Sun, Stephan Goetz

专题命中 代码与定理证明 :CoT(summary_cn,abstract);reasoning(title,abstract);chain-of-thought(abstract);分类 cs.CL

AI总结 研究专利权利要求生成中CoT提示是否有益,提出特定任务CoT方法并评估。结果显示推理增强提示可提升权利要求质量,且隐式CoT优于显式CoT,显式CoT会引入信息瓶颈,为法律任务和CoT应用提供新见解。

Comments Accepted to AI for Law Workshop @ ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03205 2025-05-28 cs.CL cs.AI 90%

MA-LoT: Model-Collaboration Lean-based Long Chain-of-Thought Reasoning enhances Formal Theorem Proving

Ruida Wang, Rui Pan, Yuxin Li, Jipeng Zhang, Yizhen Jia, Shizhe Diao, Renjie Pi, Junjie Hu, Tong Zhang

专题命中 代码与定理证明 :chain-of-thought(title,abstract);reasoning(title);CoT(abstract);verifier(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24940 2026-01-28 cs.CL 89%

SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens

SemCoT:通过语义对齐的隐式令牌加速链式推理

Yinhan He, Wendy Zheng, Yaochen Zhu, Zaiyi Zheng, Lin Su, Sriram Vasudevan, Qi Guo, Liangjie Hong, Jundong Li

机构 * University of Virginia(弗吉尼亚大学) LinkedIn Inc.(领英公司)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL

AI总结 SemCoT通过语义对齐的隐式令牌优化,提升链式推理的效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01069 2025-10-02 cs.AI 89%

Typed Chain-of-Thought: A Curry-Howard Framework for Verifying LLM Reasoning

Elija Perrier

机构 * Centre for Quantum Software and Information(量子软件与信息中心) University of Technology Sydney(悉尼技术大学)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08649 2025-07-14 cs.AI 89%

Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning

Xingguang Ji, Yahui Liu, Qi Wang, Jingyuan Zhang, Yang Yue, Rui Shi, Chenxi Sun, Fuzheng Zhang, Guorui Zhou, Kun Gai

机构 * Kuaishou Technology(快手科技)

专题命中 代码与定理证明 :reasoning(title,abstract);verifier(title,abstract);CoT(abstract);分类 cs.AI

Comments 23 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01632 2025-03-04 cs.AI 89%

CoT-VLM4Tar: Chain-of-Thought Guided Vision-Language Models for Traffic Anomaly Resolution

Tianchi Ren, Haibo Hu, Jiacheng Zuo, Xinhong Chen, Jianping Wang, Chun Jason Xue, Jen-Ming Wu, Nan Guan

专题命中 代码与定理证明 :chain-of-thought(title,abstract);CoT(title,abstract);reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21286 2026-04-07 cs.HC 89%

When the Chain Breaks: Interactive Diagnosis of LLM Chain-of-Thought Reasoning Errors

当链条断裂时:LLM推理错误的交互诊断

Shiwei Chen, Niruthikka Sritharan, Xiaolin Wen, Chenxi Zhang, Xingbo Wang, Yong Wang

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract)

AI总结 本文提出ReasonDiag系统,通过结合外部事实核查与符号逻辑验证,帮助用户高效识别和诊断LLM推理链条中的错误步骤及根本原因。

Comments Accepted to EuroVis 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23402 2025-02-11 cs.SE 89%

VisualCoder: Guiding Large Language Models in Code Execution with Fine-grained Multimodal Chain-of-Thought Reasoning

Cuong Chi Le, Hoang-Chau Truong-Vinh, Huy Nhat Phan, Dung Duy Le, Tien N. Nguyen, Nghi D. Q. Bui

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract)

Comments NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08589 2023-11-09 cs.LG cs.AI cs.CL 89%

Chain-of-Thought Reasoning is a Policy Improvement Operator

Hugh Zhang, David C. Parkes

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16998 2025-05-23 cs.CL cs.AI 88%

Do Large Language Models Excel in Complex Logical Reasoning with Formal Language?

Jin Jiang, Jianing Wang, Yuchen Yan, Yang Liu, Jianhua Zhu, Mengdi Zhang, Xunliang Cai, Liangcai Gao

专题命中 代码与定理证明 :reasoning(title,abstract);logical reasoning(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19143 2026-05-01 cs.AI cs.CR cs.CV 88%

Imitation Game for Adversarial Disillusion with Chain-of-Thought Reasoning in Generative AI

对抗性欺骗的模仿游戏:生成AI中的链式推理

Ching-Chun Chang, Fan-Yun Chen, Shih-Hong Gu, Kai Gao, Hanrui Wang, Isao Echizen

机构 * Information and Society Research Division, National Institute of Informatics(信息与社会研究部,国立信息研究所) Department of Information Engineering and Computer Science, Feng Chia University(信息工程与计算机科学系,逢甲大学)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(title,abstract);分类 cs.AI

AI总结 本文提出基于模仿游戏的对抗性欺骗防御框架,利用链式推理生成多模态代理,有效对抗生成AI中的演绎和归纳欺骗攻击。

Journal ref in IEEE Access, vol. 13, pp. 95085-95093, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02984 2025-07-29 cs.CL 88%

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought

Wentao Tan, Qiong Cao, Yibing Zhan, Chao Xue, Changxing Ding

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13312 2024-03-21 cs.CL 88%

LeanReasoner: Boosting Complex Logical Reasoning with Lean

Dongwei Jiang, Marcio Fonseca, Shay B. Cohen

专题命中 代码与定理证明 :reasoning(title,abstract);logical reasoning(title,abstract);分类 cs.CL

Comments Accepted to NAACL 2024 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04285 2026-08-06 cs.AI cs.LG 新提交 88%

The RAIL Principles for Neurosymbolic AI: Reasoning, Assurances, Interfacing and Learning

神经符号AI的RAIL原则:推理(Reasoning)、保证(Assurances)、接口(Interfacing)与学习(Learning)

Agnese Chiatti, Michael Cochez, Cristina Cornelio, Sebastijan Dumancic, Artur d'Avila Garcez, Luis C. Lamb, Lia Morra, Mathias Niepert, Robert Peharz, Alberto Speranzon, Maarten Stol, Annette Ten Teije, Thiviyan Thanapalasingam, Frank Van Harmelen, Emile Van Krieken, Antonio Vergari, Benjie Wang

机构 * Politecnico di Milano(米兰理工大学) ELLIS Institute Finland(芬兰ELLIS研究所) Åbo Akademi University(奥博 Akademi 大学) Samsung AI(三星人工智能研究院) Delft University of Technology(代尔夫特理工大学) Stony Brook University(石溪大学) Politecnico di Torino(都灵理工大学) University of Stuttgart(斯图加特大学) Graz University of Technology(格拉茨工业大学) Lockheed Martin, Advanced Technology Labs(洛克希德·马丁公司先进技术实验室) BrainCreators(BrainCreators公司) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) University of Amsterdam(阿姆斯特丹大学) University of Edinburgh(爱丁堡大学) UCLA(加利福尼亚大学洛杉矶分校)

专题命中 代码与定理证明 :reasoning(title,title_cn);分类 cs.AI、cs.LG

AI总结 该文提出神经符号AI的RAIL四项原则,可统一分析多类AI系统,助力工程师更科学地设计部署生产级AI,指导整合神经符号方法到下一代AI技术。

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03667 2025-01-28 cs.CL cs.AI 88%

Large Language Models as an Indirect Reasoner: Contrapositive and Contradiction for Automated Reasoning

Yanfang Zhang, Yiliu Sun, Yibing Zhan, Dapeng Tao, Dacheng Tao, Chen Gong

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);logical reasoning(abstract)

Comments Accepted by COLING 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15972 2026-06-16 cs.CL cs.AI cs.LG 新提交 88%

Formalize Once, Edit the Rest: Efficient Lean-Based Answer Selection for Math Reasoning

一次形式化,其余编辑:基于Lean的高效数学推理答案选择

Ji Feng, Zhouxing Shi

机构 * University of California, Riverside(加州大学河滨分校)

专题命中 代码与定理证明 :reasoning(title,abstract);math reasoning(title);分类 cs.CL、cs.AI、cs.LG

AI总结 提出BASE流水线,通过形式化一个候选答案并编辑其余答案,减少自动形式化调用约5倍,同时提升选择准确性。

Comments 15 pages, 1 figure. Code available at https://github.com/ucr-rai/base-and-edit

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06269 2026-05-08 q-bio.QM cs.AI 87%

MAT-Cell: A Multi-Agent Tree-Structured Reasoning Framework for Batch-Level Single-Cell Annotation

MAT-Cell: 一种多智能体树状推理框架用于批量单细胞注释

Yehui Yang, Zelin Zang, Xienan Zheng, Yuzhe Jia, Changxi Chi, Jingbo Zhou, Chang Yu, Jinlin Wu, Fuji Yang, Jiebo Luo, Zhen Lei, Stan Z. Li

机构 * Westlake University(西湖大学) Shenzhen University of Advanced Technology(深圳先进技术大学) Center for Artificial Intelligence and Robotics(人工智能与机器人中心) Hong Kong Institute of Science and Innovation(香港科学与创新研究院) Chinese Academy of Sciences(中国科学院) University Key Laboratory of Information and Communication Security Backup and Recovery(信息与通信安全备份与恢复大学重点实验室)

专题命中 代码与定理证明 :reasoning(title,abstract);CoT(abstract,abstract_cn);verifier(abstract);分类 cs.AI

AI总结 MAT-Cell通过分离证据基础与标签决策,结合反向验证查询和多轮辩论,提升批量单细胞注释的准确性与可追溯性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13069 2026-07-16 cs.AI cs.CL cs.LO 新提交 87%

Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution

介入式基础审计:通过谓词替换对大语言模型思维链进行黑盒前提依赖性测试

Hironao Nakamura

专题命中 代码与定理证明 :chain-of-thought(title,abstract);reasoning(abstract,comments);CoT(abstract);分类 cs.CL、cs.AI

AI总结 研究大语言模型思维链前提依赖性问题,提出介入式基础审计方法,通过谓词替换干预前提并重新运行模型,在ProntoQA基准测试中检测前提依赖性表现优异,还发现了“正确答案,错误推理”信号。

Comments Accepted at the ICLR 2026 Workshop on Logical Reasoning of Large Language Models (https://iclr.cc/virtual/2026/10017466)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05318 2025-12-08 cs.CL cs.AI cs.LG 87%

To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples

思考还是不思考:过度使用Co-T示例的元训练隐藏成本

Vignesh Kothapalli, Ata Fatahibaarzi, Hamed Firooz, Maziar Sanjabi

机构 * Stanford University(斯坦福大学) LinkedIn AI

专题命中 代码与定理证明 :CoT(title,abstract);reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出CoT-Recipe方法,通过调节元训练序列中CoT和非CoT示例的比例,提升大型语言模型在新任务上的推理准确性,实验显示在无CoT示例时准确率可提升300%。

Comments 26 pages, 45 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04592 2025-06-06 cs.CL cs.AI cs.LG 87%

Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification

Chengwu Liu, Ye Yuan, Yichun Yin, Yan Xu, Xin Xu, Zaoyu Chen, Yasheng Wang, Lifeng Shang, Qun Liu, Ming Zhang

机构 * School of Computer Science, National Key Laboratory for Multimedia Information Processing, PKU-Anker LLM Lab, Peking University(计算机学院、多媒体信息处理国家重点实验室、PKU-Anker LLM实验室、北京大学) Huawei Noah’s Ark Lab(华为诺亚实验室) The Hong Kong University of Science and Technology(香港科技大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted in ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17053 2026-03-13 cs.CL cs.AI cs.DB 86%

Knowledge Distillation with Structured Chain-of-Thought for Text-to-SQL

基于结构化思维链的知识蒸馏用于文本到SQL

Khushboo Thaker, Yony Bresler

专题命中 代码与定理证明 :chain-of-thought(title,abstract);reasoning(abstract);CoT(abstract);分类 cs.CL、cs.AI

AI总结 本文提出Struct-SQL框架,通过结构化思维链知识蒸馏提升小型语言模型在文本到SQL任务中的性能,实现8.1%的准确率提升。

Comments Accepted at the 39th Canadian Conference on Artificial Intelligence (Canadian AI 2026). This is the extended version containing additional details and appendices omitted from the camera-ready proceedings due to space constraints

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03613 2025-08-06 cs.LG cs.AI 86%

Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction

Yong Lin, Shange Tang, Bohan Lyu, Ziran Yang, Jui-Hui Chung, Haoyu Zhao, Lai Jiang, Yihan Geng, Jiawei Ge, Jingruo Sun, Jiayun Wu, Jiri Gesi, Ximing Lu, David Acuna, Kaiyu Yang, Hongzhou Lin, Yejin Choi, Danqi Chen, Sanjeev Arora, Chi Jin

专题命中 代码与定理证明 :self-correction(title,abstract);test-time compute(abstract);verifier(abstract);分类 cs.AI、cs.LG

Comments 24 pages, 10 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07106 2025-08-01 cs.CL cs.AI 86%

Theorem-of-Thought: A Multi-Agent Framework for Abductive, Deductive, and Inductive Reasoning in Language Models

Samir Abdaljalil, Hasan Kurban, Khalid Qaraqe, Erchin Serpedin

机构 * Texas A&M University(德克萨斯大学) Hamad Bin Khalifa University(哈马德·本·卡伊夫大学)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

Comments ACL 2025 KnowFM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14431 2025-05-29 cs.CL cs.LG 86%

Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains

Xu Chu, Zhijie Tan, Hanlin Xue, Guanyu Wang, Tong Mo, Weiping Li

机构 * School of Software and Microelectronics, Peking University(软件与微电子学院,北京大学)

专题命中 代码与定理证明 :reasoning(title,abstract);CoT(abstract);self-correction(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15120 2025-02-24 cs.CL cs.AI 86%

Unveiling Reasoning Thresholds in Language Models: Scaling, Fine-Tuning, and Interpretability through Attention Maps

Yen-Che Hsiao, Abhishek Dutta

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10007 2024-09-17 cs.CL cs.AI 86%

SelECT-SQL: Self-correcting ensemble Chain-of-Thought for Text-to-SQL

Ke Shen, Mayank Kejriwal

专题命中 代码与定理证明 :chain-of-thought(title,abstract);CoT(abstract);self-correction(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.13894 2023-05-30 cs.AI cs.LG 86%

LAMBADA: Backward Chaining for Automated Reasoning in Natural Language

Mehran Kazemi, Najoung Kim, Deepti Bhatia, Xin Xu, Deepak Ramachandran

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);logical reasoning(abstract);分类 cs.AI、cs.LG

Comments Accepted at ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19834 2026-01-28 cs.AI 86%

Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models

视觉生成通过多模态世界模型解锁类人推理

Jialong Wu, Xiaoying Zhang, Hongyi Yuan, Xiangcheng Zhang, Tianhao Huang, Changjing He, Chaoyi Deng, Renrui Zhang, Youbin Wu, Mingsheng Long

机构 * Tsinghua University(清华大学)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.AI

AI总结 本文提出视觉生成在特定任务中优于纯语言推理,通过构建VisWorld-Eval评估套件验证了多模态世界模型提升类人推理的能力。

Comments Project page: https://thuml.github.io/Reasoning-Visual-World

详情

展开后加载摘要…

URL PDF HTML 收藏