arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10440 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10440 篇

2410.12878 2024-10-18 cs.CL cs.AI cs.LG 75%

Towards More Effective Table-to-Text Generation: Assessing In-Context Learning and Self-Evaluation with Open-Source Models

Sahar Iravani, Tim . O . F Conrad

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14867 2024-10-17 cs.LG cs.AI cs.CL 75%

Investigating the Transferability of Code Repair for Low-Resource Programming Languages

Kyle Wong, Alfonso Amayuelas, Liangming Pan, William Yang Wang

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19131 2024-06-28 cs.CV 75%

CELLO: Causal Evaluation of Large Vision-Language Models

Meiqi Chen, Bo Peng, Yan Zhang, Chaochao Lu

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);CoT(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04205 2024-04-22 cs.CL cs.AI cs.LG 75%

Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves

Yihe Deng, Weitong Zhang, Zixiang Chen, Quanquan Gu

专题命中 推理评测 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 25 pages, 7 figures, 22 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18225 2024-02-29 cs.CL cs.AI cs.LG 75%

CogBench: a large language model walks into a psychology lab

Julian Coda-Forno, Marcel Binz, Jane X. Wang, Eric Schulz

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07492 2023-12-29 cs.CL cs.AI cs.CY cs.LG 75%

SocialStigmaQA: A Benchmark to Uncover Stigma Amplification in Generative Language Models

Manish Nagireddy, Lamogha Chiazor, Moninder Singh, Ioana Baldini

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

Comments AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16542 2023-11-29 cs.CV 75%

Agents meet OKR: An Object and Key Results Driven Agent System with Hierarchical Self-Collaboration and Self-Evaluation

Yi Zheng, Chongyang Ma, Kanle Shi, Haibin Huang

专题命中 推理评测 :reasoning(abstract);planning(abstract);self-correction(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04533 2023-08-17 cs.CL cs.AI cs.LG 75%

Prompted LLMs as Chatbot Modules for Long Open-domain Conversation

Gibbeum Lee, Volker Hartmann, Jongho Park, Dimitris Papailiopoulos, Kangwook Lee

专题命中 推理评测 :chain-of-thought(abstract);CoT(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to the Findings of ACL2023. The camera-ready version with additional experimental results will be uploaded

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15173 2022-04-01 cs.CL cs.AI cs.LG 75%

An Evaluation Dataset for Legal Word Embedding: A Case Study On Chinese Codex

Chun-Hsien Lin, Pu-Jen Cheng

专题命中 推理评测 :reasoning(abstract);logical reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 16 pages, 9 figures, 3rd International Conference on Natural Language Computing and AI (NLCAI 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13191 2025-02-11 cs.CL cs.AI 74%

MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback

Zonghai Yao, Aditya Parashar, Huixue Zhou, Won Seok Jang, Feiyun Ouyang, Zhichao Yang, Hong Yu

专题命中 推理评测 :self-correction(abstract,comments);reasoning(abstract);分类 cs.CL、cs.AI

Comments Equal contribution for the first two authors. To appear in proceedings of the Main Conference on 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL). Keywords: Question Generation, USMLE, Self-Refine, Self-Critique, and Self-Correction, LLM-as-Judge, AI for Medical Education

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14778 2026-08-20 cs.CV cs.LG 版本更新 74%

AMPLIFAI: A Multiphase CT Dataset for Benchmarking Clinical Reasoning in LI-RADS Assessment of Liver Lesions

AMPLIFAI:用于基准测试LI-RADS肝脏病变评估临床推理的多期CT数据集

Pranav Kulkarni, Nikhil Shah, Amritansh Suryavanshi, Jana G. Delfino, James Tonascia, Jade Wong-You-Cheong, Barton Lane, Joseph Chirico, Jeffrey D. Hirsch, Ang Li, Heng Huang, Florence X. Doo

机构 * University of Maryland, College Park(马里兰大学帕克分校) University of Maryland School of Medicine(马里兰大学医学院) University of Maryland Institute for Health Computing(马里兰大学健康计算研究所) University of Maryland Medical System(马里兰大学医学系统)

专题命中 推理评测 :reasoning(title);分类 cs.LG

AI总结 研究针对LI-RADS肝脏病变评估AI模型缺乏公开高质量数据集的问题,推出首个公开多期腹部CT扫描的AMPLIFAI数据集,标注LI-RADS类别并分割关键特征,助力相关透明可复现研究。

Comments 15 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14552 2026-08-18 cs.AI 新提交 74%

Large Language Models Show Metacognitive Sensitivity in Medical Reasoning

大型语言模型在医学推理中表现出元认知敏感性

Ahmad Nazzal

专题命中 推理评测 :reasoning(title);分类 cs.AI

AI总结 本研究开发临床基准测试医学LLM,发现其在AT-NCD与DRCI鉴别中具部分元认知敏感性,错误集中于冲突病例,建立了医学LLM相关评估框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13326 2026-08-14 cs.CL 新提交 74%

Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation

超越局部准确率:面向受控大语言模型推理评估的协议级可识别性审计

Junhao Luo, Ning Huang, Ziqi Sha, Wenxuan Tang, Wei Deng

机构 * School of Statistics and Data Science, Southwestern University of Finance and Economics(西南财经大学统计与数据科学学院)

专题命中 推理评测 :reasoning(title);分类 cs.CL

AI总结 该研究提出协议级可识别性审计方法,无需调用模型即可验证LLM评估协议的有效性,发现基础准确率与选择性响应保真度存在显著差异,还确定了冻结策略类的最小识别支持。

Comments 15 pages, 9 figures. Ning Huang, Ziqi Sha, and Wenxuan Tang contributed equally as second authors. Wei Deng is the corresponding author

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01548 2026-08-11 cs.CV cs.AI 74%

mChartQA: A universal benchmark for multimodal Chart Question Answer based on Vision-Language Alignment and Reasoning

Jingxuan Wei, Nan Xu, Guiyong Chang, Yin Luo, BiHui Yu, Ruifeng Guo

专题命中 推理评测 :reasoning(title);分类 cs.AI

Journal ref Pattern Recognition 172 (2026) 112348

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24235 2026-08-07 cs.AI 版本更新 74%

SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis

SP-Mind: 用于空间蛋白质组学分析的自主推理智能体

Yucheng Yuan, Yuanfeng Ji, Zhongxiao Li, Ruijiang Li

机构 * Department of Computer Science, Stanford University, Stanford, USA(计算机科学系,斯坦福大学,斯坦福,美国) Department of Radiation Oncology, Stanford University, Stanford, USA(放射肿瘤学系,斯坦福大学,斯坦福,美国)

专题命中 推理评测 :reasoning(title);分类 cs.AI

AI总结 提出首个自主AI智能体SP-Mind,统一空间蛋白质组学分析流程,通过自然语言查询实现端到端分析,在SP-Bench基准上达到最优性能。

Comments 24 pages, 6 figures. Accepted to ICML 2026. Equal contribution by Yucheng Yuan and Yuanfeng Ji

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02027 2026-08-06 cs.AI cs.ET 版本更新 74%

Zero-shot reasoning for simulating scholarly peer-review

模拟学术同行评审的零样本推理

Khalid M. Saqr

机构 * College of Engineering and Technology(工程与技术学院) Arab Academy for Science, Technology, and Maritime Transport(阿拉伯科学、技术与海运学院)

专题命中 推理评测 :reasoning(title);分类 cs.AI

AI总结 本研究构建同行评审模拟基准 xPeer,对比其与人类评审的特征差异,发现 xPeer 报告针对性等更优,人类报告显式推理等更强,成果存档于 Zenodo。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29007 2026-07-21 cs.CL 版本更新 74%

Taxonomy-Targeted Error Generation for Quantitative Reasoning

错误作为透镜:通过合成误解生成探究LLM推理

Xinming Yang, Jun Li

机构 * CUNY Graduate Center(纽约大学研究生中心) CUNY Queens College(纽约市立大学皇后学院)

专题命中 推理评测 :reasoning(title);分类 cs.CL

AI总结 提出一个框架,通过生成针对Bloom分类学五类错误的合成误解,以诊断LLM推理能力,并发现目标错误生成比自由形式错误生成更难。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14948 2026-07-08 cs.SE cs.AI 新提交 74%

Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment

超越正确性:通过可扩展的智能体判断标注增强代码大模型的架构推理能力

Kirill Vasilevski, Ximing Dong, Benjamin Rombaut, Milad Soltany, Ruochen Deng, Jiahuei Lin, Arthur Leung, Dayi Lin, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan

机构 * Centre for Software Excellence, Huawei Canada(华为加拿大软件卓越中心) Department of Computer Science, University of Manitoba, Canada(曼尼托巴大学计算机科学系) School of Computing, Queen’s University, Canada(皇后大学计算科学学院)

专题命中 推理评测 :reasoning(title);分类 cs.AI

AI总结 针对代码大模型缺乏架构理解的问题,提出智能体判断流水线,利用强LLM作为专家架构评估的代理,通过两个判断器(ACJ和AQJ)实现可扩展标注,微调模型在SWE-bench上提升高达540%,并展现跨语言泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30686 2026-07-01 cs.RO cs.AI 新提交 74%

Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning

立场:视觉-语言-动作模型无法被验证执行物理推理

Taozhao Chen, Ian Manchester, Huaming Chen

专题命中 推理评测 :reasoning(title);分类 cs.AI

AI总结 本文指出基于预训练视觉-语言模型的VLA系统在机器人操作基准上的性能提升无法区分语义匹配与物理泛化,现有评估指标不能验证物理推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29689 2026-06-30 cs.CL 74%

Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models

MLLM 能否像人类一样进行批评?评估多模态大语言模型中的开放式审美推理

Sajjad Ghiasvand, Maryam Amirizaniani, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh, Ramtin Pedarsani

机构 * UCSB(加州大学圣塔芭芭拉分校) University of Washington(华盛顿大学) UCLA(加州大学洛杉矶分校)

专题命中 推理评测 :reasoning(title);分类 cs.CL

AI总结 本文评估多模态大语言模型在开放式审美批评中的表现,发现基于参考的相似度指标存在误导,模型在选择性、特异性和多样性方面与人类批评存在系统性差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15507 2026-06-16 cs.AI 新提交 74%

Frame-Conditioned Moral Computation in LLaMA 3.1-8B-Instruct: A Mechanistic Interpretability Audit of Ethical Reasoning

LLaMA 3.1-8B-Instruct中的框架条件化道德计算:伦理推理的机械可解释性审计

Ali Dasdan, Manan Shah, W. Russell Neuman, Chad Coleman, Kund Meghani, Safinah Ali

机构 * KD Consulting, CA, USA(KD咨询公司,美国加利福尼亚州) New York University, NY, USA(纽约大学,美国纽约州)

专题命中 推理评测 :reasoning(title);分类 cs.AI

AI总结 通过机械可解释性平台分析LLaMA 3.1-8B-Instruct在54个道德提示上的内部计算,发现情境锚定效应:领域特定表示主导激活列表顶部,模型道德能力恒定但显著性高度依赖于提示选择的解释框架。

Comments 47 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15107 2026-06-16 cs.AI 新提交 74%

Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning

迈向可验证的自主数据科学:通过基于工具的推理解决不规则时间序列问答

Sanhorn Chen, Xiaoyang Chen, Boyu Liu, Roy Zhao

机构 * University of Illinois Urbana Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 推理评测 :reasoning(title);分类 cs.AI

AI总结 针对现实世界时间序列数据的不规则性,提出IRTS-ToolBench基准(1700个问题,10种任务类型,13个领域),通过标准化输入和可复现评估协议,研究LLM和AI代理在不规则条件下的表现。

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04806 2026-06-04 cs.CV cs.AI 74%

NoRA: Evaluating Grounded Reasonableness in Visual First-person Normative Action Reasoning

NoRA: 评估视觉第一人称规范性动作推理中的基于事实的合理性

Sichao Li, Sai Ma, Daniel Kilov, Secil Yanik Guyot, Zhuang Li, Seth Lazar

机构 * The University of Sydney(悉尼大学) Australian National University(澳大利亚国立大学) RMIT University(皇家墨尔本理工大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 推理评测 :reasoning(title);分类 cs.AI

AI总结 提出NoRA基准,通过事实-理由-动作支持图评估多模态模型生成合理动作并基于可见事实进行推理的能力,发现当前VLM在构建完整动作空间和绑定正确支持方面存在不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30344 2026-05-29 cs.AI 74%

Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection

小巧但可信:面向时间序列异常检测的高效视觉-语言推理

Xiaona Zhou, Muntasir Wahed, Tianjiao Yu, Constantin Brif, Ismini Lourentzou

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Sandia National Laboratories(桑迪亚国家实验室)

专题命中 推理评测 :reasoning(title);分类 cs.AI

AI总结 针对时间序列异常检测中缺乏自然语言解释的问题,构建VisAnomBench基准并微调参数高效的视觉-语言模型VisAnomReasoner,在准确性和泛化性上显著超越基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22047 2026-05-22 cs.AI 74%

Active Evidence-Seeking and Diagnostic Reasoning in Large Language Models for Clinical Decision Support

大语言模型在临床决策支持中的主动证据获取与诊断推理

Chen Zhan, Xihe Qiu, Xiaoyu Tan, Xibing Zhuang, Gengchen Ma, Yue Zhang, Shuo Li, Peifeng Liu, Xiaoxiao Ge, Liang Liu, Lu Gan

机构 * Tencent Youtu Lab(腾讯优图实验室) Case Western Reserve University(凯斯西储大学)

专题命中 推理评测 :reasoning(title);分类 cs.AI

AI总结 研究探讨了大语言模型在临床决策支持中的主动证据获取与诊断推理问题,提出了一种基于OSCE的标准化患者模拟器和可控可复现的基准测试,发现多轮证据获取会降低诊断准确性并降低支持证据质量,表明静态全上下文基准可能高估交互证据获取场景中的性能,需引入互补的交互评估以提高临床决策安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11039 2026-05-13 cs.CR cs.AI 74%

The Granularity Mismatch in Agent Security: Argument-Level Provenance Solves Enforcement and Isolates the LLM Reasoning Bottleneck

代理安全中的粒度不匹配:论证层面的溯源解决了执行问题并隔离了LLM推理瓶颈

Linfeng Fan, Ziwei Li, Yuan Tian, Yichen Wang, Rongsheng Li, Xiong Wang

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院 Gallagher 学院) King Abdullah University of Science and Technology(国王 Abdullah 科学技术大学) Dongbei University of Finance and Economics(东北财经大学) University of Science and Technology of China(中国科学技术大学)

专题命中 推理评测 :reasoning(title);分类 cs.AI

AI总结 本文提出PACT框架,通过论证层面的溯源解决代理安全中的粒度不匹配问题,隔离LLM推理瓶颈,提升安全性和实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10032 2026-05-13 cs.CL 74%

PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoning

PlantMarkerBench: 一个多物种证据导向植物标记推理基准

Sajib Acharjee Dip, Song Li, Liqing Zhang

机构 * Department of Computer Science, Virginia Tech(弗吉尼亚理工学院计算机科学系) School of Plant and Environmental Sciences, Virginia Tech(弗吉尼亚理工学院植物与环境科学学院) Fralin Biomedical Research Institute, Virginia Tech(弗吉尼亚理工学院弗拉林生物医学研究学院) FBRI Cancer Research Center, Washington, DC(华盛顿特区FBRI癌症研究中心)

专题命中 推理评测 :reasoning(title);分类 cs.CL

AI总结 本文提出PlantMarkerBench,通过整合文献检索与生物 grounding,构建了包含4种植物的基准,评估文献支持的植物标记证据解释,发现模型在功能、间接和弱支持证据上表现欠佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09618 2026-05-12 cs.CL cs.CY 74%

Statistical Scouting Finds Debate-Safe but Not Debate-Useful Cases: A Matched-Ceiling Study of Open-Weight LLM Reasoning Protocols

统计 scouting 找到辩论安全但不实用的案例:开放权重 LLM 推理协议的匹配天花板研究

Julia Hu, Alfred Shen, Kumar Lakshmipathi

机构 * Amazon Web Services(亚马逊网络服务)

专题命中 推理评测 :reasoning(title);分类 cs.CL

AI总结 研究探讨在生成 token 数限制下,不同推理协议的路由头寸及恢复潜力,发现辩论安全性和实用性不一致,需更复杂的信号来优化。

Comments 14 pages, 5 figures. Technical report / preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08815 2026-05-12 cs.LG q-bio.BM q-bio.GN q-bio.QM 74%

MicroFuse: Protein-to-Genome Expert Fusion for Microbial Operon Reasoning

MicroFuse:用于微生物启动子推理的蛋白质到基因组专家融合

Seungik Cho

机构 * Department of Physics and Astronomy, Rice University, Texas, USA(物理与天文学系,里士满大学,德克萨斯州,美国)

专题命中 推理评测 :reasoning(title);分类 cs.LG

AI总结 本文提出MicroFuse框架,通过融合蛋白质结构表示和基因组上下文表示,提升微生物启动子预测的准确性。在OG-Operon100K数据集上,MicroFuse在多个评估指标上表现最优,尤其在生物上模糊的情况中表现突出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07251 2026-05-11 cs.AI 74%

Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning

智能体能否定价反应?评估大语言模型在化学成本推理上的能力

Yuyang Wu, Yue Huang, Shuaike Shen, Xujian Wang, Shuhao Zhang, Qiyao Xue, Weichen Liu, Runtian Gao, Jian Ma, Xiangliang Zhang, Olexandr Isayev

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Notre Dame(Notre Dame 大学) University of North Carolina, Chapel Hill(北卡罗来纳大学教堂山分校) University of Pittsburgh(匹兹堡大学)

专题命中 推理评测 :reasoning(title);分类 cs.AI

AI总结 本文提出ChemCost基准,评估大语言模型在化学成本推理任务中的能力,发现工具访问并非解决问题的充分条件,且在噪声环境下表现下降。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏