arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2026-05-11 至 2026-05-11 共收录 7 信号源:cs.CL, cs.AI, cs.LG

1. 代码与定理证明 7 篇

2605.08011 2026-05-11 cs.AI stat.CO 83%

Abductive Reasoning with Probabilistic Commonsense

基于概率常识的归纳推理

Joseph Cotnareanu, Chiara Roverato, Han Zhou, Didier Chetelat, Yingxue Zhang, Mark Coates

机构 * International Laboratory on Learning Systems(学习系统国际实验室) McGill University(麦吉尔大学) Mila - Quebec Artificial Intelligence Institute(魁北克人工智能研究所) Huawei Noah's Ark Lab(华为诺亚实验室)

专题命中 代码与定理证明 :reasoning(title,abstract);chain-of-thought(abstract);分类 cs.AI

AI总结 本文提出PACS算法,通过LLM和形式求解器结合,建模常识认知的个体差异,提升大语言模型的归纳推理能力。

Journal ref Proceedings of the International Conference on Machine Learning, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06690 2026-05-11 cs.AI cs.CL cs.LG 82%

State Representation and Termination for Recursive Reasoning Systems

递归推理系统的状态表示与终止

Debashis Guha, Amritendu Mukherjee, Sanjay Kukreja, Tarun Kumar

机构 * S P Jain School of Global Management(S P Jain 全球管理学院) Indian Statistical Institute(印度统计研究所) eClerx Services Ltd.(eClerx 服务有限公司)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出了一种递归推理系统的状态表示方法及终止条件,通过epistemic状态图编码提取的主张、证据关系、开放问题和置信度权重,并定义了顺序间隙以判断迭代的必要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21654 2026-05-11 cs.LG cs.AI cs.CC 81%

Limitations on Accurate, Trusted, Human-level Reasoning

对准确、可信、人类水平推理的限制

Rina Panigrahy, Vatsal Sharan

机构 * Google Research(谷歌研究) University of Southern California(南加州大学)

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI、cs.LG

AI总结 本文研究了准确、可信和人类水平推理在AI系统中的根本矛盾,证明了准确且可信的系统无法实现人类水平推理。

Comments 19 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06673 2026-05-11 cs.CL cs.AI cs.LG 67%

Domain-level metacognitive monitoring in frontier LLMs: A 33-model atlas

领域级元认知监控在前沿大语言模型中的应用:一个33模型图谱

Jon-Paul Cacioli

机构 * Independent Researcher(独立研究员)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文通过33个前沿LLM在MMLU基准领域中的表现,揭示了元认知评分掩盖的领域级差异,发现应用/专业知识领域监控效果最佳,而形式推理和自然科学领域最难,且中等难度领域无显著差异。

Comments 25 pages, 7 figures, 1 supplementary table. Code and data: https://github.com/synthiumjp/metacognitive-profile-atlas

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08037 2026-05-11 cs.LG cs.AI 62%

Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph

超越成对:你的语言模型其实是在优化一个偏好图

Ning Liu, Chuanneng Sun, Kristina Klinkner, Shervin Malmasi

机构 * Amazon(亚马逊)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出GraphDPO,通过构建偏好图结构优化语言模型,解决传统成对优化在处理多轮次数据时的局限性,提升模型在推理和编程任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23947 2026-05-11 cs.AI 57%

GamED.AI: A Hierarchical Multi-Agent Framework for Automated Educational Game Generation

GamED.AI:一种用于自动教育游戏生成的分层多智能体框架

Shiven Agarwal, Yash Shah, Ashish Raj Shekhar, Priyanuj Bordoloi, Vivek Gupta

机构 * Arizona State University(亚利桑那州立大学)

专题命中 代码与定理证明 :reasoning(abstract);分类 cs.AI

AI总结 本文提出GamEDAI框架,通过分层多智能体方法将教师提供的问题转化为可玩的教育游戏,验证通过形式化机制合同。系统在200个问题上实现90%验证通过率和98.3%的模式合规性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26576 2026-05-11 cs.SE 50%

"Show Me You Comply... Without Showing Me Anything": Zero-Knowledge Software Auditing for AI-Enabled Systems

展示你合规...而不展示任何东西:面向AI赋能系统的零知识软件审计

Filippo Scaramuzza, Renato Cordeiro Ferreira, Giovanni Quattrocchi, Damian Andrew Tamburri, Willem-Jan van den Heuvel

专题命中 代码与定理证明 :verifier(abstract)

AI总结 本文提出ZKMLOps框架,通过将零知识证明整合到机器学习运维生命周期中,解决AI系统审计中的透明性与保密性冲突问题,提供可验证的加密证据以满足监管要求。

Comments This work has been submitted to the IEEE Transactions on Software Engineering for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏