arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 1124 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 1124 篇

2509.22834 2025-09-30 cs.NI cs.AI 57%

Bridging Language Models and Formal Methods for Intent-Driven Optical Network Design

Anis Bekri, Amar Abane, Abdella Battou, Saddek Bensalem

机构 * National Institute of Standards and Technology(美国国家标准技术研究院)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Accepted at AICCSA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20182 2025-09-25 cs.AR cs.AI 57%

Automated Multi-Agent Workflows for RTL Design

Amulya Bhattaram, Janani Ramamoorthy, Ranit Gupta, Diana Marculescu, Dimitrios Stamoulis

机构 * Chandra Family Department of Electrical and Computer Engineering(查克拉家族电子与计算机工程系)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Accepted: ML for Systems Workshop NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01830 2025-09-23 cs.CL 57%

From Language to Cognition: How LLMs Outgrow the Human Language Network

Badr AlKhamissi, Greta Tuckute, Yingtian Tang, Taha Binhuraib, Antoine Bosselut, Martin Schrimpf

机构 * EPFL(苏黎世联邦理工学院) MIT(麻省理工学院) Georgia Institute of Technology(佐治亚理工学院)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

Comments EMNLP 2025. Project Page at https://language-to-cognition.epfl.ch

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16533 2025-09-23 cs.CL 57%

Challenging the Evaluator: LLM Sycophancy Under User Rebuttal

Sungwon Kim, Daniel Khashabi

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15195 2025-09-19 cs.SE cs.AI cs.CR 57%

Orion: Fuzzing Workflow Automation

Max Bazalii, Marius Fleischer

机构 * NVIDIA

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 11 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13597 2025-09-18 cs.CR cs.AI 57%

Agentic JWT: A Secure Delegation Protocol for Autonomous AI Agents

Abhishek Goswami

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 17 pages, 6 figures, 2 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11131 2025-09-16 cs.AI cs.MA q-bio.OT 57%

Neural cellular automata: applications to biology and beyond classical AI

Benedikt Hartl, Michael Levin, Léo Pio-Lopez

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07026 2025-09-10 cs.LO cs.AI 57%

Contradictions

Yang Xu, Shuwei Chen, Xiaomei Zhong, Jun Liu, Xingxing He

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 37 Pages,9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00110 2025-09-03 cs.CY cs.AI 57%

The Application of Virtual Environments and Artificial Intelligence in Higher Education: Experimental Findings in Philosophy Teaching

Adel Vehrer, Zsolt Palfalusi

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11302 2025-08-29 cs.CL 57%

Are formal and functional linguistic mechanisms dissociated in language models?

Michael Hanna, Yonatan Belinkov, Sandro Pezzelle

机构 * Institute for Logic, Language and Computation University of Amsterdam(逻辑、语言与计算研究所 阿姆斯特丹大学) Technion – Israel Institute of Technology(技术ion-以色列理工学院)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

Comments To appear in Computational Linguistics. Pre-MIT Press publication version. 40 pages, 14 figures, 3 tables. Code available at https://github.com/hannamw/formal-functional-dissociation

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16245 2025-08-25 cs.GT cs.LG cs.MA econ.TH 57%

Limit-Computable Grains of Truth for Arbitrary Computable Extensive-Form (Un)Known Games

Cole Wyeth, Marcus Hutter, Jan Leike, Jessica Taylor

机构 * David R. Cheriton School of Computer Science, University of Waterloo(多伦多大学大卫·R·切里顿计算机科学学院) Google DeepMind and Australian National University(谷歌DeepMind和澳大利亚国立大学) Anthropic(Anthropic公司) Median Group(Median集团)

专题命中 代码与定理证明 :planning(abstract);分类 cs.LG

Comments 42 pages; 2 figures; 7 algorithms

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14927 2025-08-22 cs.GT cs.AI 57%

AI Testing Should Account for Sophisticated Strategic Behaviour

Vojtech Kovarik, Eric Olav Chen, Sami Petersen, Alexis Ghersengorin, Vincent Conitzer

机构 * Department of Computer Science(计算机科学系) Czech Technical University Prague(捷克技术大学布拉格) Global Priorities Institute(全球优先研究所) University of Oxford(牛津大学) Foundations of Cooperative AI Lab(合作人工智能基础实验室) Carnegie Mellon University(卡内基梅隆大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14644 2025-08-21 cs.AI 57%

LeanGeo: Formalizing Competitional Geometry problems in Lean

Chendong Song, Zihan Wang, Frederick Pu, Haiming Wang, Xiaohan Lin, Junqi Liu, Jia Li, Zhengying Liu

机构 * Moonshot AI Numina Peking University(北京大学) University of Toronto(多伦多大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12416 2025-08-19 cs.HC cs.AI 57%

fCrit: A Visual Explanation System for Furniture Design Creative Support

Vuong Nguyen, Gabriel Vigliensoni

机构 * Concordia University(康科德大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments In Proceedings of Explainable AI for the Arts Workshop 2025 (XAIxArts 2025) arXiv:2406.14485

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20199 2025-08-14 cs.AI 57%

StepFun-Prover Preview: Let's Think and Verify Step by Step

Shijie Shang, Ruosi Wan, Yue Peng, Yutong Wu, Xiong-hui Chen, Jie Yan, Xiangyu Zhang

机构 * StepFun University of Chinese Academy of Sciences(中国科学院大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Added links to GitHub and Hugging Face

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08171 2025-08-12 cs.SE cs.AI 57%

PyVeritas: On Verifying Python via LLM-Based Transpilation and Bounded Model Checking for C

Pedro Orvalho, Marta Kwiatkowska

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 14 pages, 6 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18337 2025-08-06 cs.AI 57%

The AlphaPhysics Term Rewriting System for Marking Algebraic Expressions in Physics Exams

Peter Baumgartner, Lachlan McGinness

机构 * CSIRO/Data61 and Australian National University(CSIRO/Data61和澳大利亚国立大学) Australian National University and CSIRO/Data61(澳大利亚国立大学和CSIRO/Data61)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23694 2025-08-01 cs.MA cs.AI 57%

A survey of multi-agent geosimulation methodologies: from ABM to LLM

Virginia Padilla, Jacinto Dávila

机构 * Departamento de Ciencia y Tecnología, Universidad Nacional Experimental de Guayana(科学与技术系,圭亚那国家实验大学) CeSiMo, Facultad de Ingeniería, Universidad de los Andes(CeSiMo,工程学院,安第斯大学)

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI

Comments 20 pages, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22009 2025-07-30 cs.AI 57%

PHAX: A Structured Argumentation Framework for User-Centered Explainable AI in Public Health and Biomedical Sciences

Bahar İlgen, Akshat Dubey, Georges Hattab

机构 * Center for Artificial Intelligence in Public Health Research (ZKI-PH), Robert Koch-Institut, Nordufer 20, Berlin, 13353, Germany(公共健康人工智能研究中心(ZKI-PH)、罗伯特· Koch研究所) Department of Mathematics and Computer Science, Freie Universität Berlin, Arnimallee 14, Berlin, 14195, Germany(数学与计算机科学系、柏林自由大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Preprint. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21067 2025-07-30 cs.AI cs.CY cs.HC 57%

SynLang and Symbiotic Epistemology: A Manifesto for Conscious Human-AI Collaboration

Jan Kapusta

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 32 pages, 4 figures. Includes 2 Appendices containing SynLang v1.2.0 protocol specification, and formal BNF grammar

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18290 2025-07-25 cs.AI 57%

Foundations for Risk Assessment of AI in Protecting Fundamental Rights

Antonino Rotolo, Beatrice Ferrigno, Jose Miguel Angel Garcia Godinez, Claudio Novelli, Giovanni Sartor

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 24 pages, 1 figure. To be published in: The Philosophical Foundations of Information Technology Law. Oxford University Press, Oxford

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17134 2025-07-24 cs.MA cs.AI 57%

Resilient Multi-Agent Negotiation for Medical Supply Chains:Integrating LLMs and Blockchain for Transparent Coordination

Mariam ALMutairi, Hyungmin Kim

机构 * Virginia Tech(弗吉尼亚理工大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 11 pages, 6 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13334 2025-07-22 cs.CL 57%

A Survey of Context Engineering for Large Language Models

Lingrui Mei, Jiayu Yao, Yuyao Ge, Yiwei Wang, Baolong Bi, Yujun Cai, Jiazhi Liu, Mingyu Li, Zhong-Zhi Li, Duzhen Zhang, Chenlin Zhou, Jiayi Mao, Tianze Xia, Jiafeng Guo, Shenghua Liu

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of California, Merced(加州大学默塞德分校) The University of Queensland(昆士兰大学) Peking University(北京大学) Tsinghua University(清华大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL

Comments ongoing work; 166 pages, 1411 citations

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14335 2025-07-22 cs.AI 57%

ProofCompass: Enhancing Specialized Provers with LLM Guidance

Nicolas Wischermann, Claudio Mayrink Verdun, Gabriel Poesia, Francesco Noseda

机构 * Federal University of Rio de Janeiro(里约热内卢联邦大学) Harvard John A. Paulson School of Engineering and Applied Sciences(哈佛大学约翰·A·保罗森工程与应用科学学院) Stanford University(斯坦福大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments 19 pages, 7 figures. Accepted at the 2nd AI for MATH Workshop at the 42nd International Conference on Machine Learning (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01789 2025-07-22 cs.SE cs.AI 57%

Generating executable oracles to check conformance of client code to requirements of JDK Javadocs using LLMs

Shan Jiang, Chenguang Zhu, Sarfraz Khurshid

机构 * The University of Texas at Austin(德克萨斯大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11479 2025-07-16 cs.AI cs.GR cs.HC 57%

Perspective-Aware AI in Extended Reality

Daniel Platnick, Matti Gruener, Marjan Alirezaie, Kent Larson, Dava J. Newman, Hossein Rahnama

机构 * Flybits Labs(Flybits实验室) Creative Ai Hub(创意人工智能中心) Toronto Metropolitan University(多伦多 Metropolitan 大学) MIT Media Lab(MIT媒体实验室)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

Comments Accepted to the International Conference on eXtended Reality (2025), 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11127 2025-07-16 cs.AI 57%

Defining neurosymbolic AI

Lennert De Smet, Luc De Raedt

机构 * Department of Computer Science, KU Leuven(计算机科学系,卢万大学) Örebro University(奥雷布罗大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09083 2025-07-15 cs.GT cs.AI 57%

Learning from Synthetic Labs: Language Models as Auction Participants

Anand Shah, Kehang Zhu, Yanchen Jiang, Jeffrey G. Wang, Arif K. Dayi, John J. Horton, David C. Parkes

机构 * Dropbox Inc.(Dropbox公司) Expected Parrot Inc.(Expected Parrot公司)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08362 2025-07-14 cs.LG 57%

Leveraging Machine Learning and Enhanced Parallelism Detection for BPMN Model Generation from Text

Phuong Nam Lê, Charlotte Schneider-Depré, Alexandre Goossens, Alexander Stevens, Aurélie Leribaux, Johannes De Smedt

机构 * Research Centre for Information Systems Engineering, KU Leuven, Leuven, Belgium(信息系统工程研究中心,鲁汶大学,比利时)

专题命中 代码与定理证明 :planning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02639 2025-07-04 cs.LG 57%

On Efficient Bayesian Exploration in Model-Based Reinforcement Learning

Alberto Caron, Chris Hicks, Vasilios Mavroudis

机构 * The Alan Turing Institute(艾伦·图灵研究所)

专题命中 代码与定理证明 :planning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏