arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10440 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10440 篇

2407.16221 2024-09-25 cs.CL 77%

Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Nishanth Madhusudhan, Sathwik Tejaswi Madhusudhan, Vikas Yadav, Masoud Hashemi

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

Comments 8 pages (excluding limitations, references and appendix) and 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11160 2024-08-30 cs.AI 77%

Low-Cost Language Models: Survey and Performance Evaluation on Python Code Generation

Jessica López Espejel, Mahaman Sanoussi Yahaya Alassan, Merieme Bouhandi, Walid Dahhane, El Hassane Ettifouri

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.AI

Comments Under review at Elsevier's Engineering Applications of Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18961 2024-08-19 cs.AI 77%

MMAU: A Holistic Benchmark of Agent Capabilities Across Diverse Domains

Guoli Yin, Haoping Bai, Shuang Ma, Feng Nan, Yanchao Sun, Zhaoyang Xu, Shen Ma, Jiarui Lu, Xiang Kong, Aonan Zhang, Dian Ang Yap, Yizhe zhang, Karsten Ahnert, Vik Kamath, Mathias Berglund, Dominic Walsh, Tobias Gindele, Juergen Wiest, Zhengfeng Lai, Xiaoming Wang, Jiulong Shan, Meng Cao, Ruoming Pang, Zirui Wang

专题命中 推理评测 :reasoning(abstract);planning(abstract);self-correction(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18968 2024-07-30 cs.AI 77%

Intelligence Analysis of Language Models

Liane Galanti, Ethan Baron

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.17447 2024-06-28 cs.CL 77%

A Large Language Model Approach to Educational Survey Feedback Analysis

Michael J. Parker, Caitlin Anderson, Claire Stone, YeaRim Oh

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

Journal ref Int J Artif Intell Educ (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17304 2024-06-26 cs.CL 77%

Leveraging LLMs for Dialogue Quality Measurement

Jinghan Jia, Abi Komma, Timothy Leffel, Xujun Peng, Ajay Nagesh, Tamer Soliman, Aram Galstyan, Anoop Kumar

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16441 2024-06-25 cs.CL 77%

UniCoder: Scaling Code Large Language Model via Universal Code

Tao Sun, Linzheng Chai, Jian Yang, Yuwei Yin, Hongcheng Guo, Jiaheng Liu, Bing Wang, Liqun Yang, Zhoujun Li

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

Comments Accepted by ACL 2024 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15862 2024-06-19 cs.CL 77%

SportQA: A Benchmark for Sports Understanding in Large Language Models

Haotian Xia, Zhengbang Yang, Yuqing Wang, Rhys Tracy, Yun Zhao, Dongdong Huang, Zezhi Chen, Yan Zhu, Yuan-fang Wang, Weining Shen

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

Comments NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12342 2024-04-19 cs.CL 77%

Large Language Models in Targeted Sentiment Analysis

Nicolay Rusnachenko, Anton Golubev, Natalia Loukachevitch

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

Comments Fine-tuned Flan-T5-xl outperforms the top #1 results of transformer-based classifier in RuSentNE-2023 competition, to appear in Lobachevskii Journal of Mathematics No.8/2024 proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09214 2024-04-09 cs.CL 77%

Mind's Mirror: Distilling Self-Evaluation Capability and Comprehensive Thinking from Large Language Models

Weize Liu, Guocong Li, Kai Zhang, Bang Du, Qiyuan Chen, Xuming Hu, Hongxia Xu, Jintai Chen, Jian Wu

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

Comments Accepted to NAACL 2024 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12582 2024-03-20 cs.CL 77%

AlphaFin: Benchmarking Financial Analysis with Retrieval-Augmented Stock-Chain Framework

Xiang Li, Zhenyu Li, Chen Shi, Yong Xu, Qing Du, Mingkui Tan, Jun Huang, Wei Lin

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

Comments COLING 2024. The first three authors contributed equally. Project website: https://github.com/AlphaFin-proj/AlphaFin

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15630 2023-10-20 cs.CL 77%

NLPBench: Evaluating Large Language Models on Solving NLP Problems

Linxin Song, Jieyu Zhang, Lechao Cheng, Pengyuan Zhou, Tianyi Zhou, Irene Li

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07486 2023-06-14 cs.CL 77%

Knowledge-Prompted Estimator: A Novel Approach to Explainable Machine Translation Assessment

Hao Yang, Min Zhang, Shimin Tao, Minghan Wang, Daimeng Wei, Yanfei Jiang

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14896 2026-08-18 cs.CL cs.LG 新提交 76%

Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs

小语言模型中的可解释跨语言对齐:探究日英双语大语言模型的文化与语用推理

Florian Braun

机构 * Sakaguchi–Inui Laboratory (Tohoku University)(东北大学坂口–乾实验室) Natural Language Understanding Team at RIKEN AIP(理化学研究所先进智能项目中心自然语言理解团队) Swallow Project at the Institute of Science Tokyo(东京科学大学Swallow项目) Sakana AI Tohoku University(东北大学) RIKEN AIP(理化学研究所先进智能项目中心) Institute of Science Tokyo(东京科学大学)

专题命中 推理评测 :reasoning(title);分类 cs.CL、cs.LG

AI总结 本研究构建J-PragEval-v0基准,结合线性探针与教师强制评估探究TinySwallow-1.5B的日英语用表征,提出语用表征引导方法,下一步将扩展至Llama-3.1-Swallow-8B。

Comments 15 pages, no figures. Introduces the J-PragEval-v0 minimal-pair benchmark

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10453 2026-08-11 cs.CL cs.AI 版本更新 76%

Reasoning about Intent for Ambiguous Requests

意图推理以应对歧义请求

Irina Saparina, Mirella Lapata

机构 * School of Informatics, University of Edinburgh(爱丁堡大学信息学院)

专题命中 推理评测 :reasoning(title);分类 cs.CL、cs.AI

AI总结 本文提出生成结构化响应以枚举歧义请求的不同解释,通过强化学习训练模型提升覆盖有效解释的召回率和抑制虚假解释的精确率,实验表明方法在覆盖有效答案方面优于基线方法。

Comments COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29738 2026-08-10 cs.CL cs.AI 版本更新 76%

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions

Multi-Legal-Bench: 跨司法管辖区、语言和法律传统的法律推理评估LLM

Volodymyr Ovcharov

机构 * SecondLayer

专题命中 推理评测 :reasoning(title);分类 cs.CL、cs.AI

AI总结 提出Multi-Legal-Bench,首个跨司法管辖区法律基准,在6个国家、4个语系和1.34亿份法院判决上评估LLM,发现少样本效果跨辖区复制、无单一模型主导所有语言、跨语言迁移不遵循语言邻近性、分词器效率不显著预测跨语言准确率。

Comments 17 pages, 5 figures, 9 tables. v2 corrects scorer and taxonomy defects, adds no-model baselines showing label leakage, re-runs the Lithuanian cells on de-leaked text, and withdraws the claim that few-shot helps on judgment-form classification everywhere; all tables and figures regenerated. Dataset: https://huggingface.co/datasets/overthelex/multi-legal-bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02615 2026-08-05 cs.CL cs.AI 新提交 76%

OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning

OncoTriad-QA:用于泛癌推理的患者级放射学-病理学-基因组学基准

Ahnaf Munir, Dannong Wang, Michael W. McDonald, Mubarak Shah, Pegah Khosravi, Yu Tian

专题命中 推理评测 :reasoning(title);分类 cs.CL、cs.AI

AI总结 本文提出OncoTriad-QA多模态癌症问答基准及OncoVLM模型,实验显示OncoVLM经微调后在泛癌问答任务上优于现有模型,该基准可用于相关模型的训练与评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07699 2026-08-03 cs.CL cs.AI 版本更新 76%

DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain

DRIP-R:零售领域决策与推理在现实政策模糊性下的基准测试

Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang, Bei Chen, Yufang Hou

机构 * Interdisciplinary Transformation University Austria(跨学科转型大学奥地利) Amazon Berlin(亚马逊柏林)

专题命中 推理评测 :reasoning(title);分类 cs.CL、cs.AI

AI总结 DRIP-R基准测试通过现实零售政策模糊性构建无唯一正确解决方案的场景,评估LLM在模糊政策下的决策与推理能力,揭示其固有的系统性挑战。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25600 2026-07-30 cs.IR cs.AI cs.CL 版本更新 76%

Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs

超越自我认知:在语言模型中跨推理与检索传播不确定性

Chandan Kumar Sah, Li Zhang, Xiaoli Lian

专题命中 推理评测 :reasoning(title);分类 cs.CL、cs.AI

AI总结 研究在知识密集型问答中,利用黑盒语言模型的置信度指导检索路由。核心方法是BeyondUncertainty,先得临时答案和置信度,依阈值决定检索策略。该方法提升了平均词元级F1,减少检索段落,在多数设置中优于随机分配,揭示了证据获取与词元效率的权衡。

Comments 9 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08775 2026-07-13 cs.CL cs.LG 新提交 76%

HALO: Hybrid Adaptive Latent Reasoning for Language Models

HALO:语言模型的混合自适应潜在推理

Micah Zhang

机构 * Lockheed Martin(洛克希德·马丁公司)

专题命中 推理评测 :reasoning(title);分类 cs.CL、cs.LG

AI总结 研究如何用少量自适应计算改进预训练语言模型,提出HALO混合自适应潜在细化方法,结合粗略与选择性第二阶段细化。在基准比较中,HALO成绩最佳,使用细化步骤少,兼具更好的细化分配,计算量也少。

Comments 15 pages, 4 figures, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02175 2026-07-03 cs.AI cs.LG 新提交 76%

A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks

基于评分量表的专家级临床推理任务中前沿语言模型的受控比较

Samiha A. Ismail, Fan X. Chen, Ali Merali

机构 * Prime Analytics Consulting Limited(普林特分析咨询有限公司) Talbert House(塔尔伯特大楼;托马斯·莫尔大学) Thomas More University(MAKZ) MAKZ

专题命中 推理评测 :reasoning(title);分类 cs.AI、cs.LG

AI总结 通过5个专家编写的临床场景和加权评分量表,评估GPT、Claude和Gemini三个前沿模型,发现关键标准通过率低(32-42%),而低权重标准通过率高(80-90%),52%的关键标准无模型通过。

Comments 13 pages, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24950 2026-06-25 cs.LG cs.AI 新提交 76%

MacroLens: A Multi-Task Benchmark for Contextual Financial Reasoning under Macroeconomic Scenarios

MacroLens:宏观经济情景下的多任务上下文金融推理基准

Patara Trirat, Jin Myung Kwak, Jay Heo, Heejun Lee, Sung Ju Hwang

机构 * KAIST(韩国科学技术院)

专题命中 推理评测 :reasoning(title);分类 cs.AI、cs.LG

AI总结 针对金融决策中价格、基本面、宏观和文本四信号联合评估的缺失,构建了涵盖4426只美股、七个任务、1130个宏观事件的基准,评估19种方法,揭示上下文特征的重要性。

Comments 25 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09995 2026-06-09 cs.CL cs.AI 版本更新 76%

Context Over Compute Human-in-the-Loop Outperforms Iterative Chain-of-Thought Prompting in Interview Answer Quality

上下文胜过计算 人类在环优于迭代思维链提示在面试回答质量上的表现

Kewen Zhu, Zixi Liu, Yanjing Li, Jing Chen

专题命中 推理评测 :chain-of-thought(title);分类 cs.CL、cs.AI

AI总结 本文通过对比人类在环和自动思维链提示方法,发现人类在环在面试回答质量评估中表现更优,且迭代次数更少,同时具有更高的训练效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04704 2026-05-29 cond-mat.mtrl-sci cs.AI cs.CL 76%

AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials

AtomWorld: 评估大型语言模型在晶体材料空间推理能力的基准

Taoyuze Lv, Alexander Chen, Fengyu Xie, Chu Wu, Jeffrey Meng, Dongzhan Zhou, Yingheng Wang, Bram Hoex, Zhicheng Zhong, Tong Xie

机构 * University of New South Wales, NSW, Sydney, Australia(新南威尔士大学,新州,悉尼,澳大利亚) Suzhou Institute for Advanced Research, University of Science(苏州先进研究院,科学大学) Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室,上海,中国) Cornell University(康奈尔大学)

专题命中 推理评测 :reasoning(title);分类 cs.CL、cs.AI

AI总结 提出AtomWorld基准,通过十种基本原子结构操作评估LLM在材料科学中的空间推理能力,发现Claude Opus 4.6表现最佳但复杂空间关系操作成功率低,表明LLM更适合作为辅助工具而非完全自主的科研代理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29001 2026-05-29 cs.LG cs.AI 76%

FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks

FormInv:数学推理基准中语义不变性的测量协议

Nishal Thomas, Noel Thomas

机构 * Mohamed Bin Zayed University of Artificial Intelligence(Mohamed Bin Zayed人工智能大学)

专题命中 推理评测 :reasoning(title);分类 cs.AI、cs.LG

AI总结 提出FormInv协议,通过跨模型一致性审计检测语义错误,并引入语义一致性率(SCR)和Cochran's Q指标,揭示标准基准无法捕捉的排名变化和模型不一致性。

Comments 18 pages, 3 figures. Under review for the 3rd AI for Math Workshop (AI4Math), ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07231 2026-05-27 cs.CL cs.AI 76%

EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models

EconCausal: 面向大语言模型的上下文感知经济推理基准

Donggyu Lee, Hyeok Yun, Meeyoung Cha, Sungwon Park, Sangyoon Park, Jihee Kim

机构 * Graduate School of Data Science, KAIST(韩国科学技术院数据科学研究生院) College of Business, KAIST(韩国科学技术院商学院) Data Science for Humanity Group, MPI-SP(马克斯·普朗克所际数据科学为人类集团) School of Computing, KAIST(韩国科学技术院计算学院) Division of Social Science, HKUST(香港科技大学社会科学系)

专题命中 推理评测 :reasoning(title);分类 cs.CL、cs.AI

AI总结 提出EconCausal基准,包含从顶级经济金融期刊提取的10,490个上下文标注因果三元组,评估大语言模型在指定上下文中推断因果方向及随上下文变化调整判断的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11453 2026-05-15 cs.MA cs.AI cs.LG cs.SI math.SP 76%

Predictive Maps of Multi-Agent Reasoning: A Successor-Representation Spectrum for LLM Communication Topologies

多智能体推理的预测地图:一种用于LLM通信拓扑的后继表示谱

Ethan Parks, Dalal Alharthi

机构 * University of Arizona(亚利桑那大学)

专题命中 推理评测 :reasoning(title);分类 cs.AI、cs.LG

AI总结 本文提出基于后继表示谱的结构诊断方法,用于评估多智能体LLM通信拓扑的性能,通过分析谱半径、谱间隙和条件数等指标,预测拓扑在漂移、共识和扰动鲁棒性方面的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04825 2026-04-21 cs.CL cs.AI 76%

Plausibility as Commonsense Reasoning: Humans Succeed, Large Language Models Do not

可解释性作为常识推理:人类成功,大语言模型却不成功

Sercan Karakaş

机构 * University of Chicago(芝加哥大学)

专题命中 推理评测 :reasoning(title);分类 cs.CL、cs.AI

AI总结 研究探讨大语言模型是否能像人类一样在歧义解析中结合世界知识与句法结构。通过土耳其前置定语从句附着歧义,发现人类表现出显著的可解释性效应,而大语言模型则表现不一致。

Comments Accepted to The Workshop on Cognitive Modeling and Computational Linguistics co-located with LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20586 2026-03-25 cs.LG cs.AI 76%

MKA: Memory-Keyed Attention for Efficient Long-Context Reasoning

MKA:用于高效长上下文推理的内存键注意

Dong Liu, Yanxuan Yu, Ben Lengerich, Ying Nian Wu

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Columbia University(哥伦比亚大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 推理评测 :reasoning(title);分类 cs.AI、cs.LG

AI总结 本文提出MKA机制,通过多级KV缓存和动态路由提升长上下文推理效率,FastMKA在保持准确率的同时显著提升训练速度和评估效率。

Comments Accepted to the ACM Computing Frontiers 2026 Conference (Oral Presentation) and the ICML 2025 Long Context Modeling Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20969 2026-03-24 cs.LG cs.CL 76%

Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge

理解Transformer中的上下文回忆:微调如何使模型在预训练知识上进行上下文推理

Bhavya Vasudeva, Puneesh Deora, Alberto Bietti, Vatsal Sharan, Christos Thrampoulidis

机构 * University of Southern California(南加州大学) University of British Columbia(不列颠哥伦比亚大学) Flatiron Institute(Flatiron研究所)

专题命中 推理评测 :reasoning(title);分类 cs.CL、cs.LG

AI总结 本文研究了Transformer模型在上下文学习中如何通过微调实现对预训练知识的推理,探讨了预训练和微调对上下文回忆能力的影响及机制。

Comments 28 pages, 26 figures

详情

展开后加载摘要…

URL PDF HTML 收藏