arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 10451 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 10451 篇

2605.11334 2026-05-13 cs.LG cs.CL cs.IR 62%

VERDI: Single-Call Confidence Estimation for Verification-Based LLM Judges via Decomposed Inference

VERDI:基于验证的LLM裁判的单次调用置信度估计方法

Jasmine Qi, Danylo Dantsev, Muyang Sun

机构 * Indeed Inc(Indeed公司)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 VERDI通过分解推理过程提取置信度信号,无需额外调用,提升验证型LLM裁判的可信度评估。

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11260 2026-05-13 cs.LG cs.AI 62%

Curriculum Learning-Guided Progressive Distillation in Large Language Models

基于课程学习的渐进蒸馏:大型语言模型中的知识蒸馏

Jincheng Cao, Fanzhi Zeng, Leqi Liu, Aryan Mokhtari

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Google Research(谷歌研究)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出CLPD框架,通过结合数据难度与教师强度,改进知识蒸馏方法,实验证明其在推理任务中优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11143 2026-05-13 cs.CL cs.AI cs.IR 62%

ClinicalBench: Stress-Testing Assertion-Aware Retrieval for Cross-Admission Clinical QA on MIMIC-IV

ClinicalBench: 用于MIMIC-IV跨入院临床问答的应力测试型断言感知检索

Alex Stinard

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 ClinicalBench通过400个问题测试43名MIMIC-IV患者,评估断言敏感类别下的检索性能,采用EpiKG知识图谱检索方法,通过六种LLM模型测试,发现断言感知检索在临床问答中提升22个百分点。

Comments 46 pages including appendices (two-column preprint format). Under review at JAMIA. Code, frozen evaluator, and benchmark released at https://huggingface.co/datasets/alexstinard/epikg-clinicalbench. ClinicalBench v2 is a 400-question MIMIC-IV stress test for assertion-aware retrieval

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11086 2026-05-13 cs.CR cs.AI cs.LG 62%

ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?

ExploitGym:AI代理能否将安全漏洞转化为实际攻击?

Zhun Wang, Nico Schiller, Hongwei Li, Srijiith Sesha Narayana, Milad Nasr, Nicholas Carlini, Xiangyu Qi, Eric Wallace, Elie Bursztein, Luca Invernizzi, Kurt Thomas, Yan Shoshitaishvili, Wenbo Guo, Jingxuan He, Thorsten Holz, Dawn Song

机构 * UC Berkeley(加州大学伯克利分校) Max Planck Institute for Security and Privacy(马克斯·普朗克安全与隐私研究所) UC Santa Barbara(加州大学圣巴巴拉分校) Arizona State University(亚利桑那州立大学) Anthropic(Anthropic公司) OpenAI Google(谷歌)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 ExploitGym通过大规模真实漏洞数据评估AI代理的漏洞利用能力,揭示前沿模型在复杂漏洞利用中的表现,凸显AI发展对网络安全的潜在风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15664 2026-05-13 cs.LG cs.AI 62%

Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints

Stargazer:一种可扩展的AI代理在天体物理约束下的模型拟合基准环境

Xinge Liu, Terry Jingchen Zhang, Bernhard Schölkopf, Zhijing Jin, Kristen Menou

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所) Max Planck Institute for Intelligent Systems(智能系统马克斯·普朗克研究所) ELLIS Institute Tübingen(图宾根ELLIS研究所)

专题命中 推理评测 :test-time compute(abstract);分类 cs.AI、cs.LG

AI总结 Stargazer通过120个任务评估AI代理在动态物理约束下的模型拟合能力,揭示了数值优化与物理约束间的差距,并提供了可扩展的评估框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22933 2026-05-13 cs.AI cs.CL 62%

RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild

RW-Post:可审计的证据导向多模态事实核查

Danni Xu, Shaojing Fan, Harry Cheng, Mohan Kankanhalli

机构 * School of Computing (SoC), National University of Singapore (NUS)(新加坡国立大学计算机学院(SoC)) National University of Singapore (NUS)(新加坡国立大学) Department of Electrical and Computer Engineering (ECE), National University of Singapore (NUS)(新加坡国立大学电子与计算机工程系(ECE))

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 RW-Post提出一种可审计的多模态事实核查基准,通过人类事实核查文章提取证据,支持不同场景下的可控评估,提升视觉 grounding 和证据利用能力。

Comments Code and dataset will be released at https://github.com/xudanni0927/AgentFact

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06371 2026-05-13 cs.CL cs.AI 62%

OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA

OASIS:一个多语言多模态数据集用于文化导向的语音视觉问答

Firoj Alam, Ali Ezzat Shahroor, Md. Arid Hasan, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Mohamed Bayan Kmainasi, Shammur Absar Chowdhury, Basel Mousi, Fahim Dalvi, Nadir Durrani, Natasa Milic-Frayling

机构 * Qatar Computing Research Institute(卡塔尔计算研究 institute)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 OASIS是一个多语言多模态数据集,涵盖图像、文本和语音,旨在评估模型在现实场景中的文化导向推理能力,包含0.92M真实图像和14.8M问答对。

Comments Multimodal Foundation Models, Large Language Models, Native, Multilingual, Language Diversity, Contextual Understanding, Culturally Informed

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04265 2026-05-13 cs.AI cs.CL math.ST stat.ML stat.TH 62%

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation

不要Pass@k:大规模语言模型评估的贝叶斯框架

Mohsen Hariri, Amirhossein Samandar, Michael Hinczewski, Vipin Chaudhary

机构 * Department of Computer and Data Sciences, Case Western Reserve University(计算机与数据科学系,凯斯西储大学) Department of Physics, Case Western Reserve University(物理系,凯斯西储大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出基于贝叶斯的评估框架,用后验估计替代Pass@k和avg@N,提升排名稳定性与透明度,适用于二元和非二元评估。

Comments OpenReview (ICLR 2026): https://openreview.net/forum?id=PTXi3Ef4sT

Journal ref The Fourteenth International Conference on Learning Representations (ICLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10889 2026-05-12 cs.LG cs.AI 62%

Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why

揭示在线策略蒸馏:它何时有益,何时有害,以及原因

Mohammadreza Armandpour, Fatih Ilhan, David Harrison, Ajay Jaiswal, Duc N. M Hoang, Fartash Faghri, Yizhe Zhang, Minsik Cho, Mehrdad Farajtabar

机构 * Apple(苹果公司)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 研究探讨在线策略蒸馏在不同场景下的有效性,提出基于梯度对齐度的诊断方法,发现蒸馏信号在学生表现不佳时更有效,且最优上下文依赖学生模型能力和任务需求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10186 2026-05-12 cs.CL cs.AI 62%

LegalCiteBench: Evaluating Citation Reliability in Legal Language Models

LegalCiteBench: 评估法律语言模型中的引文可靠性

Sijia Chen, Hang Yin, Shunfan Zhou

机构 * Northeastern University(东北大学) Phala

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出LegalCiteBench,用于评估法律语言模型在闭书环境下引文恢复、验证和案例匹配的能力,发现即使最强模型在引文检索和完成任务上得分低于7/100,且多数模型存在误导性引文问题。

Comments Preprint. 23 pages including references and appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10171 2026-05-12 cs.CL cs.AI 62%

When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews

当评论存在分歧时:科学同行评审中的细粒度矛盾分析

Sandeep Kumar, Yash Kamdar, Abid Hossain, Bharti Kumari, Tanik Saikh, Asif Ekbal

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Patna, India(印度理工学院帕纳瓦分校计算机科学与工程系) School of Computer Engineering, KIIT Deemed to be University, Bhubaneswar, India(比哈尔邦布尔萨大学计算机工程学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出细粒度的同行评审矛盾分析方法,通过识别矛盾证据跨度并分配矛盾强度评分,改进了现有二元矛盾检测方法,引入RevCI基准和IMPACT框架,并通过TIDE模型实现高效部署。

Comments accepted at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08031 2026-05-12 cs.SD cs.AI cs.LG eess.AS 62%

AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs

AU-Harness:音频大语言模型的开放式工具包

Hoang Nguyen, Sidharth Surapaneni, Akshay Kalkunte, Jash Mehta, Aman Tiwari, Oluwanifemi Bamgbose, Khyati Mahajan, Jash Shah, Shruthan Radhakrishna, Sathwik Tejaswi Madhusudhan, Vikas Yadav, Sai Rajeswar

机构 * ServiceNow University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 AU-Harness通过优化处理流程和并行执行,实现151%的速度提升,提供标准化提示协议和灵活配置,促进音频大语言模型的系统性发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10027 2026-05-12 cs.CL cs.AI 62%

Speech-based Psychological Crisis Assessment using LLMs

基于语音的心理危机评估使用大语言模型

Terumi Chiba, Yang Luo, Ziyun Cui, Yongsheng Tong, Chao Zhang

机构 * Tsinghua University(清华大学) Peking University Huilongguan Clinical Medical School(北京大学回龙guan临床医学院) WHO Collaborating Centre for Research and Training in Suicide Prevention(世界卫生组织自杀预防研究与培训协作中心)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出利用大语言模型自动分类危机等级,通过引入语音学注入方法和增强推理训练策略,提升热线服务质量。

Comments 5 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09918 2026-05-12 cs.LG cs.AI cs.CY 62%

NaiAD: Initiate Data-Driven Research for LLM Advertising

NaiAD:启动基于数据的研究以进行大语言模型广告

Yihang Zhang, Zimeng Huang, Ren Zhai, Yipeng Kang, Tonghan Wang

机构 * Tsinghua University(清华大学) College of AI(人工智能学院) Department of Literature, Arts and Communication(文学、艺术与传播系) Anhui International Studies University(安徽国际关系大学) State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出NaiAD数据集,用于研究大语言模型广告,包含58999个精心构建的广告嵌入响应与用户查询,通过理论指导的评估指标提升用户体验和商业价值。

Comments 37 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09852 2026-05-12 cs.AI cs.CE cs.CY cs.LG 62%

Fairness of Explanations in Artificial Intelligence (AI): A Unifying Framework, Axioms, and Future Direction toward Responsible AI

人工智能中解释的公平性:统一框架、公理与迈向负责任人工智能的未来方向

Gideon Popoola, John Sheppard

机构 * Montana State University(蒙大拿州立大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文探讨了人工智能中解释公平性的问题,提出统一框架和公理,指出解释过程中的程序性偏见,并提出评估流程以实现公平性审计。

Comments 53 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09661 2026-05-12 cs.CL cs.AI 62%

MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies

MedMeta:用于从医学研究中合成元分析结论的LLM基准测试

Huy Hoang Ha, Benoit Favre, Francois Portet

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MedMeta基准测试,评估LLM从医学元分析摘要中生成结论的能力,发现检索增强生成方法显著优于纯参数方法,揭示当前RAG系统在识别否定证据方面的缺陷。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09348 2026-05-12 cs.CL cs.AI cs.DB cs.MM 62%

HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities

HOME-KGQA:一个多模态知识图谱问答基准数据集用于家庭日常活动

Shusaku Egami, Aoi Ohta, Tomoki Tsujimura, Masaki Asada, Tatsuya Ishigaki, Ken Fukuda, Masahiro Hamasaki, Hiroya Takamura

机构 * National Institute of Advanced Industrial Science(国家工业科学与技术研究院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出HOME-KGQA基准数据集,用于多模态知识图谱问答,包含多跳自然语言问题和图数据库查询语言,挑战多级时空推理和多模态 grounding。实验表明基于LLM的KGQA方法在该数据集上表现不佳,凸显实际部署中KGQA系统的挑战。

Comments 12 pages, 4 figures, 7 tables, accepted at LREC2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09330 2026-05-12 cs.LG cs.AI 62%

The Trap of Trajectory: Towards Understanding and Mitigating Spurious Correlations in Agentic Memory

轨迹的陷阱:朝着理解并缓解代理记忆中的虚假相关性

Luoxi Tang, Rupali Rajendra Vaje, Yuqiao Meng, Sakshi Sunil Narkar, Weicheng Ma, Zeyu Ding, Dazheng Zhang, Zhaohan Xi

机构 * Binghamton University, State University of New York(宾夕法尼亚州立大学) Oakland University(奥克兰大学) University of Pennsylvania(宾夕法尼亚大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文研究代理记忆中的虚假相关性问题,通过基准测试揭示记忆在清洁输入下提升推理但放大虚假模式依赖性,并提出CAMEL方法在写入和检索时减少对虚假模式的依赖,同时保持性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09272 2026-05-12 cs.AI cs.CL cs.CV 62%

Towards Conversational Medical AI with Eyes, Ears and a Voice

面向有眼睛、耳朵和声音的对话式医疗AI

Meet Shah, Jason Gusdorf, Anil Palepu, Chunjong Park, Jack W. O'Sullivan, Vishnu Ravi, Tim Strother, Pavel Dubov, Aliya Rysbek, Toshiyuki Fukuzawa, Yana Lunts, Jan Freyberg, Michael B. Chang, Aniruddh Raghu, David Stutz, Devora Berlowitz, Eliseo Papa, Taylan Cemgil, JD Velasquez, Jack Chen, Arthur Chen, Doug Fritz, Charlie Taylor, Katya Tregubova, Jing Rong Lim, Richard Green, Sara Mahdavi, Mahvish Nagda, Jihyeon Lee, Craig Schiff, Liviu Panait, Sukhdeep Singh, Valentin Liévin, David G. T. Barrett, Hannah Gladman, Anna Cupani, Francesca Pietra, Uchechi Okereke, Katherine Tong, Clemens Meyer, Erwan Rolland, Mili Sanwalka, Michael D. Howell, Shixiang Shane Gu, Bibo Xu, Euan A. Ashley, S. M. Ali Eslami, Gregory Wayne, Pushmeet Kohli, Vivek Natarajan, Adam Rodman, Alan Karthikesalingam, Ryutaro Tanno

机构 * Google DeepMind(谷歌DeepMind) Google Research(谷歌研究) Beth Israel Deaconess Medical Center, Harvard Medical School(贝塞斯达医院, 哈佛医学院) Stanford University(斯坦福大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出AI co-clinician系统,利用音频视频数据实现实时临床决策,通过TelePACES评估标准显示其在管理计划和诊断差异方面接近医生,但在体格检查和疾病特异性推理上仍有不足。

Comments Video examples are available on Youtube: https://youtu.be/y5Vaa_SN1t0, https://youtu.be/dC4icb75vLQ, and https://youtu.be/E7iEvWo-E6c

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08838 2026-05-12 cs.CL cs.AI 62%

Generating Leakage-Free Benchmarks for Robust RAG Evaluation

生成无泄漏的基准以评估鲁棒的RAG

Jiayi Liu, Jiaxing Zhang, Bowen Jin, Jennifer Neville

机构 * Department of Computer Science, Purdue University(普渡大学计算机科学系) New Jersey Institute of Technology(新泽西理工学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Microsoft Research(微软研究院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出SeedRG方法,通过半合成方式生成无知识泄漏的RAG评估基准,解决基准老化问题,确保评估的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08462 2026-05-12 cs.CL cs.AI 62%

Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment

基准是否低估了大语言模型的性能?通过大语言模型优先的人类仲裁评估来评估幻觉检测

I. F. Atasoy, B. Mutlu, E. A. Sezer, A. Wahdan

机构 * Department of Computer Engineering, Hacettepe University(哈切塔佩大学计算机工程系) Department of Computer Engineering, Ankara University(安卡拉大学计算机工程系) Zephlen AI and Information Technologies Inc.(泽夫伦人工智能与信息技术公司)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文通过比较原始基准标注与Gemini 2.5 Flash和GPT-5 Mini的推理和跨度预测,评估了摘要任务中上下文幻觉检测的准确性,并发现人类仲裁提高了三重一致性和模型准确性。

Comments Presented at the ROMCIR Workshop at ECIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08437 2026-05-12 cs.CL cs.AI 62%

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

Magis-Bench:评估大语言模型在法官级别法律任务上的能力

Ramon Pires, Thales Sales Almeida, Celio Larcher Junior, Giovana Bonás, Hugo Abonizio, Marcos Piau, Roseval Malaquias Junior, Thiago Laitz, Rodrigo Nogueira

机构 * Maritaca AI Jusbrasil Campinas São Paulo Brazil(巴西坎皮纳斯圣保罗)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 Magis-Bench评估大语言模型在法官级别法律任务上的能力,包含74道来自巴西司法职位竞争考试的问题,采用LLM-as-a-judge方法评估23种先进模型,结果显示法官级法律推理和写作仍具挑战性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08220 2026-05-12 cs.AI cs.CE cs.CL cs.CV cs.SE 62%

Spatial Priming Outperforms Semantic Prompting: A Grid-Based Approach to Improving LLM Accuracy on Chart Data Extraction

空间提示优于语义提示:一种基于网格的方法以提高LLM在图表数据提取上的准确性

Andrei Lazarev, Dmitrii Sedov, Alexander Galkin

机构 * Russian Federation(俄罗斯联邦)

专题命中 推理评测 :chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 本文通过对比空间提示与语义提示,发现基于网格的提示方法能显著降低图表数据提取误差,优于传统语义引导策略。

Comments his is the version of the article accepted for publication in SUMMA 2025 after peer review. The final, published version is available at IEEE Xplore: https://doi.org/10.1109/SUMMA68668.2025.11302248

Journal ref 2025 7th International Conference on Control Systems, Mathematical Modeling, Automation and Energy Efficiency (SUMMA), Lipetsk, Russian Federation, 2025, pp. 799-804

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08197 2026-05-12 cs.LG cs.AI 62%

ReplaySCM: A Benchmark for Executable Causal Mechanism Induction from Interventions

ReplaySCM:一个用于从干预证据中可执行因果机制归纳的基准

Serafim Batzoglou

机构 * Frontier LLMs(前沿大语言模型) Original, Extra Worlds, and Counterexample Audit (CEx)(原始、额外世界和反例审计)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 ReplaySCM通过不同设置测试模型从有限干预证据中归纳可执行因果机制的能力,评估可执行回放泛化,提升局部前驱模式覆盖率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07305 2026-05-11 cs.CL cs.AI 62%

MedAction: Towards Active Multi-turn Clinical Diagnostic LLMs

MedAction:迈向主动多轮临床诊断LLMs

Hsin-Ling Hsu, Zizheng Wang, Donghua Zhang, Nai-Chia Chen, Jerry Wang, Jun-En Ding, Chia-Hsuan Hsu, Guoan Wang, Feng Liu, Fang-Ming Hung, Chenwei Wu, Liyue Shen

机构 * National Chengchi University(国立中正大学) Georgetown University(乔治城大学) University of Michigan(密歇根大学) Stevens Institute of Technology(史蒂文斯理工学院) National Taiwan University of Science and Technology(台湾科技大学) Far Eastern Memorial Hospital(东方纪念医院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MedAction,通过LLM与环境交互生成高质量多轮诊断轨迹,解决现有模型在动态证据下推理不足的问题,提升开源医疗LLMs性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15567 2026-05-11 cs.AI cond-mat.mtrl-sci cs.LG physics.chem-ph 62%

Evaluating Large Language Models in Scientific Discovery

评估大型语言模型在科学发现中的表现

Zhangde Song, Jieyu Lu, Yuanqi Du, Botao Yu, Thomas M. Pruyn, Yue Huang, Kehan Guo, Xiuzhe Luo, Yuanhao Qu, Yi Qu, Yinkai Wang, Haorui Wang, Jeff Guo, Jingru Gan, Parshin Shojaee, Di Luo, Andres M Bran, Gen Li, Qiyuan Zhao, Shao-Xiong Lennon Luo, Yuxuan Zhang, Xiang Zou, Wanru Zhao, Yifan F. Zhang, Wucheng Zhang, Shunan Zheng, Saiyang Zhang, Sartaaj Takrim Khan, Mahyar Rajabi-Kochi, Samantha Paradi-Maropakis, Tony Baltoiu, Fengyu Xie, Tianyang Chen, Kexin Huang, Weiliang Luo, Meijing Fang, Xin Yang, Lixue Cheng, Jiajun He, Soha Hassoun, Xiangliang Zhang, Wei Wang, Chandan K. Reddy, Chao Zhang, Zhiling Zheng, Mengdi Wang, Le Cong, Carla P. Gomes, Chang-Yu Hsieh, Aditya Nandy, Philippe Schwaller, Heather J. Kulik, Haojun Jia, Huan Sun, Seyed Mohamad Moosavi, Chenru Duan

机构 * Deep Principle(深原则) Department of Computer Science, Cornell University(计算机科学系,康奈尔大学) Department of Computer Science and Engineering, The Ohio State University(计算机科学与工程系,俄亥俄州立大学) Department of Chemical Engineering & Applied Chemistry, University of Toronto(化学工程与应用化学系,多伦多大学) Department of Computer Science and Engineering, University of Notre Dame(计算机科学与工程系,圣母大学) QuEra Computing Inc.(QuEra计算公司) Department of Pathology, Department of Genetics, Cancer Biology Program, Stanford University School of Medicine(病理学系、遗传学系、癌症生物学项目,斯坦福大学医学院) Harvard Law School(哈佛法学院) Department of Computer Science, Tufts University(计算机科学系,塔夫茨大学) School of Computational Science and Engineering, Georgia Institute of Technology(计算科学与工程学院,佐治亚理工学院) Department of Computer Science, University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校) Department of Computer Science, Virginia Tech(计算机科学系,弗吉尼亚理工大学) Department of Physics, Tsinghua University(物理系,清华大学) Institute for Advanced Study, Tsinghua University(清华大学高级研究所) Laboratory of Artificial Chemical Intelligence, Ecole Polytechnique Federale de Lausanne(人工化学智能实验室,瑞士联邦理工学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一个基于场景的基准测试,评估LLM在生物学、化学、材料科学和物理学中的科学发现能力,揭示了模型在科学发现任务中的性能差距和改进方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21464 2026-05-08 cs.CL cs.AI 62%

Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation

对话式非验证学习:通过元评估自我进化的大型语言模型

Yuan Sui, Bryan Hooi

机构 * National University of Singapore(新加坡国立大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出CoNL框架,通过多智能体自我对弈实现生成、评估与元评估的统一,利用批评质量提升解决方案来改进模型,无需外部评委或真实标签。

Comments Accepted by ICML'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01417 2026-05-05 cs.CL cs.AI 62%

Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks

Medmarks:面向医疗任务的综合性开源大语言模型评估套件

Benjamin Warner, Ratna Sagari Grandhi, Max Kieffer, Aymane Ouraq, Saurav Panigrahi, Geetu Ambwani, Kunal Bagga, Nikhil Khandekar, Arya Hariharan, Nishant Mishra, Manish Ram, Shamus Sim Zi Yang, Ahmed Essouaied, Adepoju Jeremiah Moyondafoluwa, Robert Scholz, Bofeng Huang, Molly Beavers, Srishti Gureja, Anish Mahishi, Sameed Khan, Maxime Griot, Hunar Batra, Jean-Benoit Delbrouck, Siddhant Bharadwaj, Ronald Clark, Ashish Vashist, Anas Zafar, Leema Krishna Murali, Harsh Deshpande, Ameen Patel, William Brown, Johannes Hagemann, Connor Lane, Paul Steven Scotti, Tanishq Mathew Abraham

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出Medmarks,一个包含30个医疗任务基准的开源评估套件,系统评估61个模型,揭示前沿模型在医疗推理中的优势及模型对答案顺序偏见的敏感性。

Comments website: https://medmarks.ai

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00833 2026-05-05 cs.LG cs.AI 62%

Agentopic: A Generative AI Agent Workflow for Explainable Topic Modeling

Agentopic:一种用于可解释主题建模的生成式AI代理工作流

Brice Valentin Kok-Shun, Johnny Chan, Gabrielle Peko, David Sundaram

机构 * The University of Auckland, New Zealand(奥克兰大学,新西兰)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 Agentopic通过多代理协作实现主题识别与解释,提升可解释性与准确性,其在BBC数据集上达到0.95的F1分数,优于LDA并接近BERTopic。

Comments 16 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19298 2026-05-05 cs.CL cs.AI cs.IR 62%

IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text

IndiaFinBench:用于评估大语言模型在印度金融监管文本性能的评估基准

Rajveer Singh Pall

机构 * Gyan Ganga Institute of Technology and Sciences(加延加anga科技与科学学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出IndiaFinBench,首个公开的评估基准,用于评估大语言模型在印度金融监管文本上的性能,包含406个专家标注的问题-答案对,涵盖四个任务类型,评估十二个模型,发现数值推理任务表现最显著。

Comments 24 pages, 4 figures, 11 tables. Dataset and evaluation code at https://github.com/rajveerpall/IndiaFinBench

详情

展开后加载摘要…

URL PDF HTML 收藏