arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 1117 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1117 篇

2510.06478 2026-01-06 cs.LG cs.AI 62%

Anytime-Valid Answer Sufficiency Certificates for LLM Generation via Sequential Information Lift

通过顺序信息提升实现LLM生成的任何时间有效性答案充分性证书

Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma

机构 * Department of Computer Science, Iowa State University(计算机科学系,爱荷华州立大学) Department of Civil, Construction & Environmental Engineering, Iowa State University(土木、建设与环境工程系,爱荷华州立大学)

专题命中 代码与定理证明 :verifier(abstract);分类 cs.AI、cs.LG

AI总结 通过顺序信息提升实现LLM生成的任何时间有效性答案充分性证书,利用经验动态形式提升方法,在减少生成长度的同时保持delta级误差控制,并通过轻量级正确性门提升端任务正确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00816 2026-01-06 cs.AI cs.CR cs.LG 62%

MathLedger: A Verifiable Learning Substrate with Ledger-Attested Feedback

MathLedger: 一种具有账本证明反馈的可验证学习基础

Ismail Ahmad Abdullah

机构 * CNU(中国矿业大学)

专题命中 代码与定理证明 :verifier(abstract);分类 cs.AI、cs.LG

AI总结 MathLedger通过整合形式验证、密码学证明和学习动态,提供一种可验证学习的基础,实现可审计的机器认知系统。

Comments 14 pages, 1 figure, 2 tables, 2 appendices with full proofs. Documents v0.9.4-pilot-audit-hardened audit surface with fail-closed governance, canonical JSON hashing, and artifact classification. Phase I infrastructure validation; no capability claims

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17904 2025-12-16 cs.CR cs.AI cs.CL 62%

BreakFun: Jailbreaking LLMs via Schema Exploitation

BreakFun: 通过模式利用对LLM进行劫持

Amirkia Rafiei Oskooei, Mehmet S. Aktas

机构 * Department of Computer Engineering, Yildiz Technical University(计算机工程系,伊兹密尔技术大学)

专题命中 代码与定理证明 :chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 BreakFun通过利用LLM对结构模式的遵守能力,揭示了其在劫持攻击中的脆弱性,并提出对抗性提示解构作为缓解策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15809 2025-12-02 cs.CL cs.AI cs.DB 62%

Chain-of-Query: Unleashing the Power of LLMs in SQL-Aided Table Understanding via Multi-Agent Collaboration

链式查询:通过多智能体协作释放LLM在SQL辅助表格理解中的潜力

Songyuan Sui, Hongyi Liu, Serena Liu, Li Li, Soo-Hyun Choi, Rui Chen, Xia Hu

机构 * Rice University(里士大学) Samsung Electronics America(三星电子美国分公司) Warner Bros. Discovery(华纳兄弟发现)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 链式查询通过多智能体协作提升SQL辅助表格理解的准确性与有效性

Comments AACL 2025 Main Conference (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21050 2025-11-27 cs.LG cs.AI stat.ML 62%

Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs

打破安全性与能力的权衡:具有可验证奖励的强化学习在LLMs中维持安全护栏

Dongkyu Derek Cho, Huan Song, Arijit Ghosh Chowdhury, Haotian An, Yawei Wang, Rohit Thekkanal, Negin Sokhandan, Sharlina Keshava, Hannah Marlowe

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过可验证奖励的强化学习方法,在提升LLM推理能力的同时维持安全性,挑战了传统安全性与能力权衡的假设。

Comments AAAI-26 Workshop on Post-AI Formal Methods

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11770 2025-11-18 cs.AI cs.LG 62%

Learning to Refine: An Agentic RL Approach for Iterative SPARQL Query Construction

Floris Vossebeld, Shenghui Wang

机构 * Faculty of Electrical Engineering, Mathematics and Computer Science(电气工程、数学与计算机科学学院) University of Twente(特文特大学) Microsoft Netherlands(微软荷兰)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08274 2025-11-12 cs.AI cs.CL 62%

Multi-Agent GraphRAG: A Text-to-Cypher Framework for Labeled Property Graphs

Anton Gusarov, Anastasia Volkova, Valentin Khrulkov, Andrey Kuznetsov, Evgenii Maslov, Ivan Oseledets

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Code to be released

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25427 2025-10-30 cs.CL cs.AI 62%

RLMEval: Evaluating Research-Level Neural Theorem Proving

Auguste Poiroux, Antoine Bosselut, Viktor Kunčak

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings. RLMEval benchmark released: https://github.com/augustepoiroux/RLMEval

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24115 2025-10-29 cs.AI cs.LG 62%

HistoLens: An Interactive XAI Toolkit for Verifying and Mitigating Flaws in Vision-Language Models for Histopathology

Sandeep Vissapragada, Vikrant Sahu, Gagan Raj Gupta, Vandita Singh

机构 * Indian Institute of Technology(印度理工学院) All India Institute of Medical Sciences(全印度医学科学研究所)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17873 2025-10-28 cs.CL cs.AI cs.CE 62%

MOOSE-Chem3: Toward Experiment-Guided Hypothesis Ranking via Simulated Experimental Feedback

Wanhao Liu, Zonglin Yang, Jue Wang, Lidong Bing, Di Zhang, Dongzhan Zhou, Yuqiang Li, Houqiang Li, Erik Cambria, Wanli Ouyang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Nanyang Technological University(南洋理工大学) MiroMind

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21894 2025-10-28 cs.CL cs.AI 62%

Understanding Network Behaviors through Natural Language Question-Answering

Mingzhe Xing, Chang Tian, Jianan Zhang, Lichen Pan, Peipei Liu, Zhaoteng Yan, Yinliang Yue

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01346 2025-10-14 cs.AI cs.CL 62%

Aristotle: IMO-level Automated Theorem Proving

Tudor Achim, Alex Best, Alberto Bietti, Kevin Der, Mathïs Fédérico, Sergei Gukov, Daniel Halpern-Leistner, Kirsten Henningsgard, Yury Kudryashov, Alexander Meiburg, Martin Michelsen, Riley Patterson, Eric Rodriguez, Laura Scharff, Vikram Shanker, Vladmir Sicca, Hari Sowrirajan, Aidan Swope, Matyas Tamas, Vlad Tenev, Jonathan Thomm, Harold Williams, Lawrence Wu

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14756 2025-10-10 cs.LG cs.AI 62%

LLINBO: Trustworthy LLM-in-the-Loop Bayesian Optimization

Chih-Yu Chang, Milad Azvar, Chinedum Okwudire, Raed Al Kontar

机构 * University of Michigan(密歇根大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01272 2025-10-03 cs.AI cs.LG 62%

Modeling Others' Minds as Code

Kunal Jha, Aydan Yuenan Huang, Eric Ye, Natasha Jaques, Max Kleiman-Weiner

机构 * Department of Computer Science, University of Washington(华盛顿大学计算机科学系) Department of Computer Science, Johns Hopkins University(约翰霍普金斯大学计算机科学系)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25343 2025-10-01 cs.AI cs.CL 62%

Spontaneous High-Order Generalization in Neural Theory-of-Mind Networks

Yiming Wang, Rui Wang

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09871 2025-09-22 cs.CL cs.AI 62%

Emulating Public Opinion: A Proof-of-Concept of AI-Generated Synthetic Survey Responses for the Chilean Case

Bastián González-Bustamante, Nando Verelst, Carla Cisternas

机构 * Universidad Diego Portales(迪亚戈·波特莱斯大学) Leiden University(莱顿大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments Working paper: 18 pages, 4 tables, 2 figures

Journal ref Empiria Lab Method Series (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12233 2025-09-17 cs.CR cs.AI cs.ET cs.LG cs.NI 62%

Towards Trustworthy Agentic IoEV: AI Agents for Explainable Cyberthreat Mitigation and State Analytics

Meryem Malak Dif, Mouhamed Amine Bouchiha, Abdelaziz Amara Korba, Yacine Ghamri-Doudane

机构 * L3i - La Rochelle University, La Rochelle, France(L3i - 拉罗谢尔大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 7 figures, Accepted at LCN'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06902 2025-09-09 cs.CL cs.CR cs.DB cs.LG 62%

Proof-Carrying Numbers (PCN): A Protocol for Trustworthy Numeric Answers from LLMs via Claim Verification

Aivin V. Solatorio

专题命中 代码与定理证明 :verifier(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01716 2025-09-03 cs.AI cs.CL 62%

An LLM-enabled semantic-centric framework to consume privacy policies

Rui Zhao, Vladyslav Melnychuk, Jun Zhao, Jesse Wright, Nigel Shadbolt

机构 * University of Oxford(牛津大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14870 2025-08-29 cs.LO cs.AI cs.LG 62%

Application of AI to formal methods - an analysis of current trends

Sebastian Stock, Jannik Dunkelau, Atif Mashkoor

机构 * Institute of Software System Engineering(软件系统工程研究所) Johannes Kepler University Linz(约翰·凯撒大学林茨分校) Faculty of Mathematics and Natural Sciences(数学与自然科学学院) Heinrich Heine University Düsseldorf(海因里希·海涅大学杜塞尔多夫分校)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14275 2025-08-21 cs.CL cs.AI 62%

Disentangling concept semantics via multilingual averaging in Sparse Autoencoders

Cliff O'Reilly, Ernesto Jimenez-Ruiz, Tillman Weyde

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06017 2025-08-11 cs.SE cs.CL cs.LG 62%

Position: Intelligent Coding Systems Should Write Programs with Justifications

Xiangzhe Xu, Shiwei Feng, Zian Su, Chengpeng Wang, Xiangyu Zhang

机构 * Purdue University(普渡大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.LG

Comments The first two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02584 2025-08-05 cs.CL cs.AI 62%

MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification

Ming Pok Ng, Junqi Jiang, Gabriel Freedman, Antonio Rago, Francesca Toni

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14722 2025-07-22 cs.LG cs.AI 62%

LeanTree: Accelerating White-Box Proof Search with Factorized States in Lean 4

Matěj Kripner, Michal Šustr, Milan Straka

机构 * Charles University, Faculty of Mathematics and Physics(查尔斯大学数学与物理系) Czech Technical University, Faculty of Electrical Engineering(捷克技术大学电气工程系)

专题命中 代码与定理证明 :verifier(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02726 2025-07-04 cs.AI cs.LG 62%

Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving

Matthieu Zimmer, Xiaotong Ji, Rasul Tutunov, Anthony Bordg, Jun Wang, Haitham Bou Ammar

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室) Imperial College London(伦敦帝国学院) Huawei Lagrange Center(华为拉格朗日中心) UCL Centre for AI(大学学院人工智能中心)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17161 2025-06-19 cs.CL cs.LG cs.LO 62%

Interchangeable Token Embeddings for Extendable Vocabulary and Alpha-Equivalence

İlker Işık, Ramazan Gokberk Cinbis, Ebru Aydin Gol

机构 * Department of Computer Engineering, Middle East Technical University, Ankara, Turkey(中欧技术大学计算机工程系) Microsoft, İstanbul, Turkey(微软公司)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.LG

Comments ICML 2025 Poster Paper, Camera Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13172 2025-06-18 cs.CL cs.AI 62%

AI-Facilitated Analysis of Abstracts and Conclusions: Flagging Unsubstantiated Claims and Ambiguous Pronouns

Evgeny Markhasin

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16207 2025-06-10 cs.AI cs.CL cs.PL 62%

From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs

Jialun Cao, Yaojie Lu, Meiziniu Li, Haoyang Ma, Haokun Li, Mengda He, Cheng Wen, Le Sun, Hongyu Zhang, Shengchao Qin, Shing-Chi Cheung, Cong Tian

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) Guangzhou Institute of Technology, Xidian University(西安电子科技大学广州研究院) Chongqing University(重庆大学) ICTT and ISN Laboratory, Xidian University(西安电子科技大学ICTT和ISN实验室)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06046 2025-05-30 cs.CR cs.AI cs.CL 62%

LLM for SoC Security: A Paradigm Shift

Dipayan Saha, Shams Tarek, Katayoon Yahyaei, Sujan Kumar Saha, Jingbo Zhou, Mark Tehranipoor, Farimah Farahmandi

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

Comments 42 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13484 2025-05-21 cs.AI cs.CL 62%

Evaluating Large Language Models for Real-World Engineering Tasks

Rene Heesch, Sebastian Eilermann, Alexander Windmann, Alexander Diedrich, Philipp Rosenthal, Oliver Niggemann

机构 * Helmut Schmidt University(海德堡-哈雷大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏