arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 12534 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 12534 篇

2402.07386 2024-07-26 cs.CL 89%

Chain-of-Layer: Iteratively Prompting Large Language Models for Taxonomy Induction from Limited Examples

Qingkai Zeng, Yuyang Bai, Zhaoxuan Tan, Shangbin Feng, Zhenwen Liang, Zhihan Zhang, Meng Jiang

专题命中 领域大模型 :large language model(title);language model(title);prompting(title);分类 cs.CL

Journal ref Published in CIKM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01825 2024-07-09 cs.DC cs.CL cs.HC 89%

Large Language Models to the Rescue: Reducing the Complexity in Scientific Workflow Development Using ChatGPT

Mario Sänger, Ninon De Mecquenem, Katarzyna Ewa Lewińska, Vasilis Bountris, Fabian Lehmann, Ulf Leser, Thomas Kosch

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

Journal ref Sänger et. al: A qualitative assessment of using ChatGPT as large language model for scientific workflow development, GigaScience, Volume 13, 2024, giae030

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15871 2026-08-18 cs.CY cs.CL cs.LG 新提交 88%

Large Language Models as Implicit Sociological Models: Reconstructing Voting Behaviour from Sociodemographic Profiles

大型语言模型作为隐含的社会学模型:从社会人口统计特征重建投票行为

Roman Neruda, Martin Bakoš, Josef Šlerka, Vít Tuček, Petra Vidnerová, Gabriela Kadlecová

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.LG

AI总结 本研究提出将大型语言模型作为隐含社会学模型的方法论框架,以2021年捷克议会选举为案例,证实其可从社会人口统计特征重建投票行为,为计算社会科学提供新探索工具并明确其局限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08636 2026-08-17 cs.CL cs.AI cs.DL cs.IR 版本更新 88%

Enhancing Scientific Named Entity Recognition via Large Language Models: A Type-driven Multi-task Learning Approach

基于大语言模型增强科学命名实体识别:一种类型驱动的多任务学习方法

Tong Bao, Yi Zhao, Heng Zhang, Chengzhi Zhang

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 该研究针对LLMs处理SciNER时因实体类型过多导致准确率低的问题,提出类型驱动的多任务学习方法TdSciNER,通过实体类型筛选、多任务学习和示例选择策略提升性能,相关方法在三个数据集上达到与全监督模型相当的效果。

Journal ref Expert Systems With Applications, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12882 2026-08-11 cs.CL cs.AI 88%

LLMs-Healthcare : Current Applications and Challenges of Large Language Models in various Medical Specialties

LLMs-Healthcare : 大型语言模型在各医疗专科中的当前应用与挑战

Ummara Mumtaz, Awais Ahmed, Summaya Mumtaz

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本文综述了大型语言模型在医疗领域中的最新应用与挑战,涵盖诊断、治疗及各专科的创新贡献,探讨其在医疗诊断和患者护理中的潜力与限制。

Comments 26 pages and one figure

Journal ref Artificial Intelligence in Health ,Published online: 2 April 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29738 2026-08-10 cs.CL cs.AI 版本更新 88%

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions

Multi-Legal-Bench: 跨司法管辖区、语言和法律传统的法律推理评估LLM

Volodymyr Ovcharov

机构 * SecondLayer

专题命中 领域大模型 :LLM(title_cn,summary_cn);pretraining(abstract);prompting(abstract);分类 cs.CL、cs.AI

AI总结 提出Multi-Legal-Bench,首个跨司法管辖区法律基准,在6个国家、4个语系和1.34亿份法院判决上评估LLM,发现少样本效果跨辖区复制、无单一模型主导所有语言、跨语言迁移不遵循语言邻近性、分词器效率不显著预测跨语言准确率。

Comments 17 pages, 5 figures, 9 tables. v2 corrects scorer and taxonomy defects, adds no-model baselines showing label leakage, re-runs the Lithuanian cells on de-leaked text, and withdraws the claim that few-shot helps on judgment-form classification everywhere; all tables and figures regenerated. Dataset: https://huggingface.co/datasets/overthelex/multi-legal-bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05157 2026-08-07 cs.CL cs.AI 新提交 88%

Large Language Models Threaten Double-blind Review

大语言模型对双盲评审构成威胁

Bulambo Mwendelwa Gloire, Prasenjit Mitra

机构 * Carnegie Mellon University Africa(卡内基梅隆大学非洲分校)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 该研究指出双盲评审易受大语言模型(LLMs)破坏,LLMs可通过论文标题摘要高效恢复作者身份,需重新评估AI增强研究生态的匿名性与公平性维护方式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14917 2026-08-03 cs.CL cs.HC cs.LG 88%

Self-reflecting Large Language Models: A Hegelian Dialectical Approach

Sara Abdali, Michael Solodko, Can Goksen, Saeed Amizadeh, Julie E. Maybee, Kazuhito Koishida, Pashmina Cameron

机构 * Microsoft Applied Sciences Group (ASG)(微软应用科学组) Department of Philosophy, Lehman College, City University of New York(哲学系,莱曼学院,新 York 城市大学)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26179 2026-07-30 q-bio.NC cs.AI cs.CL 新提交 88%

Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition

认知趋同:大语言模型与人类认知的深层相似性

Chandra Sripada, Richard Lewis

专题命中 领域大模型 :large language model(title);language model(title);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 该研究指出,尽管大语言模型与人类存在多方面差异,但在推理组织等五个认知维度上与人类认知结构趋同,相关核心原则可用于解释基于大语言模型的系统智能。

Comments 23 pages, 0 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15176 2026-07-22 cs.AI cs.CL cs.HC 版本更新 88%

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

用于科学可视化素养的多模态大语言模型基准测试

Patrick Phuoc Do, Chau M. Ta, Chaoli Wang

机构 * University of Notre Dame(圣母大学)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 研究对六个多模态大语言模型进行科学可视化素养基准测试,涵盖多种技术和任务类型。通过封闭世界协议评估闭源和开源模型,与人类参与者数据对比。发现模型表现不均,Gemini最强,开源模型低于人类基线,明确SciVis素养对评估多模态AI系统的必要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10275 2026-07-14 cs.AI cs.CL 新提交 88%

Information-seeking failures of large language models in agentic clinical reasoning

大语言模型在代理临床推理中的信息获取失败

Krischan Braitsch, Laura K. Schmalbrock, Theresa Weltermann, Andrew F. Berdel, Isabella Miller, Kai Tran, Michael Heider, Sabrina Kraus, Florian Bassermann, Jacqueline Lammert, Sebastian Ziegelmayer, Marcus Makowski, Lisa C. Adams, Keno K. Bressem

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 研究大语言模型在临床推理中信息获取失败问题,开发血液肿瘤学代理评估框架,发现信息利用率是诊断准确性关键指标,推理痕迹与准确性不相关,主要失败模式为搜索满足等,指出模型主要限制是不确定性下信息获取失败。

Comments 66 pages, 9 figures; includes supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20691 2026-07-07 cs.CL cs.AI 新提交 88%

Specific Domain Ontology Construction Using Large Language Models

使用大语言模型构建特定领域本体

Vivian Magri Alcaldi Soares, Renata Wassermann

机构 * University of São Paulo (USP)(圣保罗大学) Center for Artificial Intelligence (C4AI)(人工智能中心)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 利用大语言模型作为领域专家自动构建概念层次结构,以巴西海洋领土为例评估GPT-3.5和GPT-4生成的本体,发现整体连贯但需优化。

Comments Presented at NeLaMKRR@KR, 2025 (arXiv:2511.09575)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23057 2026-06-23 cs.IR cs.CL cs.CY cs.LG 新提交 88%

Who Owns the AI Recommendation? A Multi-Industry Empirical Map of Brand Category Ownership Across Large Language Models

谁拥有AI推荐?跨行业大语言模型品牌类别所有权实证图谱

Dmitrij Żatuchin

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.LG

AI总结 本研究提出三个探索性指标(类别所有权指数、竞争真空指数和置换得分),分析三个大语言模型在50个品牌、五个行业的推荐集中度与竞争结构,发现推荐集中度中等,跨模型一致性仅41.6%,置换程度因行业而异。

Comments 21 pages, 4 figures, 7 tables. Under review at Journal of Marketing Analytics (Palgrave Macmillan). Data and analysis code on Zenodo, https://doi.org/10.5281/zenodo.20788142

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18922 2026-06-18 cs.CL cs.AI 新提交 88%

As Easy as Rocket Science: Assessing the Ability of Large Language Models to Interpret Negation in Figurative Language

像火箭科学一样简单:评估大型语言模型解释比喻语言中否定能力的研究

Jasmine Owers, Edwin Simpson, Martha Lewis

机构 * Intelligent Systems Lab University of Bristol(智能系统实验室 英国布里斯托尔大学) ILLC University of Amsterdam(阿姆斯特丹大学语言学研究所)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本研究通过开发新的注释数据集,测试多种大型语言模型在比喻语言中理解否定的能力,发现否定与比喻的组合对模型构成挑战,且性能高度依赖提示风格。

Comments 16 pages, 16 figures; for associated code and data see https://github.com/jrdowers/Negation-and-Fig-Lang; To be published in Transactions of the Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12953 2026-06-12 cs.AI cs.CV cs.LG eess.IV 新提交 88%

OpenMedQ: Broad Open Pretraining for Medical Vision-Language Models

OpenMedQ:面向医学视觉语言模型的广泛开放预训练

Ibrahim Gulluk, Max Van Puyvelde, Olivier Gevaert

机构 * Stanford University(斯坦福大学) Stanford University School of Medicine(斯坦福大学医学院) Ghent University(根特大学)

专题命中 领域大模型 :language model(title,abstract);pretraining(title,abstract);分类 cs.AI、cs.LG

AI总结 提出OpenMedQ,在14个数据集(约335万样本)上预训练医学视觉语言模型,在PathVQA上BLEU-1达75.9,超越562B参数的Med-PaLM M,并在8个未见医学分类任务上取得最高平均macro-F1(0.757)。

Comments Medical Imaging with Deep Learning (MIDL) 2026, Short Paper Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10587 2026-06-10 cs.LG cs.AI 新提交 88%

Towards Diverse Scientific Hypothesis Search with Large Language Models

面向多样化科学假设搜索的大语言模型

Haorui Wang, Parshin Shojaee, Kazem Meidani, Kunyang Sun, José Miguel Hernández-Lobato, Teresa Head-Gordon, Jiajun He, Chandan K. Reddy, Chao Zhang, Yuanqi Du

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.AI、cs.LG

AI总结 针对科学假设搜索中多样性崩溃问题,提出基于并行回火的多温度进化框架,在固定验证预算下提升假设质量与多样性。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28866 2026-05-29 cs.LG cs.AI 88%

Continuity and Ordinality Matter: Constraining Time Series Tokens for Effective Time Series Analysis with Large Language Models

连续性与序数性至关重要:利用大语言模型进行有效时间序列分析的时间序列令牌约束

Musheng Li, Ziying Zhang, Cheng jin, Yuantao Gu

机构 * Department of Electronic Engineering(电子工程系)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.AI、cs.LG

AI总结 针对令牌化时间序列大语言模型忽略连续性和序数性的问题,提出COM策略,通过几何约束初始化与训练阶段,提升模型在多个时间序列分析基准上的性能与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25920 2026-05-26 cs.CL cs.AI 88%

Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning

LLM 能时间旅行吗?通过强化学习增强法律智能搜索中的时间一致性

Wei Fan, Yining Zhou, Mufan Zhang, Yanbing Weng, Yiran HU, Tianshi Zheng, Baixuan Xu, Chunyang Li, Jianhui Yang, Haoran Li, Yangqiu Song

机构 * Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China(香港科技大学计算机科学与工程系) School of Law, Tsinghua University, Beijing, China(清华大学法学院) Cheriton School of Computer Science, University of Waterloo, Waterloo, Canada(滑铁卢大学丘成桐计算机科学系)

专题命中 领域大模型 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出 LegalSearch-R1 框架,结合本地 statute RAG 和在线搜索,通过强化学习在跨修订期数据上训练,以解决法律 LLM 的时间偏差和搜索代理缺乏时间约束的问题,在13项法律任务上超越现有方法。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03191 2026-05-26 cs.CV cs.AI cs.LG 88%

AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation

AnatomiX:一种解剖学感知的胸部X光解读多模态大语言模型

Anees Ur Rehman Hashmi, Numan Saeed, Christoph Lippert

机构 * Hasso Plattner Institute(霍普夫纳研究所) MBZUAI(穆萨大学人工智能研究所)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出AnatomiX,一种两阶段解剖学感知多模态大语言模型,通过先识别解剖结构再执行下游任务,在解剖定位、短语定位、定位诊断和定位描述任务上相比现有方法提升超过25%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23007 2026-05-25 q-fin.TR cs.AI cs.LG q-fin.PM 88%

MadEvolve: Evolutionary Optimization of Trading Systems with Large Language Models

MadEvolve: 基于大型语言模型的交易系统进化优化

Yurii Kvasiuk, Tianyi Li, Owen Colegrove, Moritz Münchmeyer

机构 * Department of Physics, University of Wisconsin–Madison(威斯康星大学麦迪逊分校物理系) Event Horizon Labs(事件地平线实验室)

专题命中 领域大模型 :large language model(title);language model(title);LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 提出 MadEvolve 框架,利用大型语言模型驱动的进化算法优化量化金融中的交易策略与 alpha 生成,以比特币交易为例,在特征集演化、策略组件优化及联合演化上取得显著改进,并对比了其他智能搜索方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18856 2026-05-20 cs.AI cs.CL 88%

Entry-level guide to the use of large language models for medical research

大型语言模型在医学研究中应用的入门指南

Qiao Jin, Nicholas Wan, Robert Leaman, Shubo Tian, Zhizheng Wang, Yifan Yang, Zifeng Wang, Guangzhi Xiong, Po-Ting Lai, Qingqing Zhu, Benjamin Hou, Maame Sarfo-Gyamfi, Gongbo Zhang, Aidan Gilson, Balu Bhasuran, Zhe He, Aidong Zhang, Jimeng Sun, Chunhua Weng, Ronald M. Summers, Qingyu Chen, Yifan Peng, Zhiyong Lu

机构 * National Library of Medicine (NLM), National Institutes of Health (NIH)(国家医学图书馆(NLM)、国立卫生研究院(NIH)) Department of Computer Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系) Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系) Department of Biomedical Informatics, Columbia University(哥伦比亚大学生物医学信息学系) School of Medicine, Yale University(耶鲁大学医学院) School of Information, Florida State University(佛罗里达州立大学信息学院) Department of Radiology and Imaging Sciences, NIH Clinical Center(国立卫生研究院临床中心放射学与影像科学部) Department of Population Health Sciences, Weill Cornell Medicine(韦尔医学院人口健康科学系)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一套可操作的指南,帮助医疗专业人员更高效地利用大型语言模型(LLMs)进行医学研究,涵盖任务制定、模型选择、提示工程、微调和模型部署等关键步骤,确保安全可靠地将LLMs应用于临床实践。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15000 2026-05-15 cs.CL cs.AI 88%

Quantifying and Mitigating Premature Closure in Frontier LLMs

对前沿大语言模型中过早闭合进行量化与缓解

Rebecca Handler, Suhana Bedi, Nigam Shah

机构 * Department of Medicine, Stanford University(斯坦福大学医学系) Department of Biomedical Data Science, Stanford University(斯坦福大学生物医学数据科学系)

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文研究了大语言模型在医疗任务中过早闭合的问题,通过评估五个前沿模型发现其在不确定情况下仍频繁给出答案,安全提示虽能减少错误,但仍有残留问题,需进一步验证医疗LLM是否能判断何时不应回答。

Comments 14 pages, 3 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13450 2026-05-14 cs.AI cs.CL cs.HC 88%

Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers

评估大语言模型的创造力:测试、限制与新前沿

Samuel Schapiro, Alexi Gladstone, Jonah Black, Heng Ji

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本文通过系统研究人类创造力测试对大语言模型创造力预测的有效性,发现不同构念测试效果差异显著,提出DRAT测试在预测科学构思能力上具有显著优势。

Comments 36 pages. Extended version of work under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13045 2026-05-14 cs.LG cs.CL 88%

Large Language Models Lack Temporal Awareness of Medical Knowledge

大语言模型缺乏医学知识的时间意识

Zihan Guan, Qiao Jin, Guangzhi Xiong, Fangyuan Chen, Mengxuan Hu, Qingyu Chen, Yifan Peng, Zhiyong Lu, Anil Vullikanti

机构 * University of Virginia(弗吉尼亚大学) National Institutes of Health(美国国家卫生研究院) Dana-Farber Cancer Institute(达纳-法伯癌症研究所) Yale University(耶鲁大学) Weill Cornell Medicine(韦氏 Cornell 医学院)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.LG

AI总结 本文提出TempoMed-Bench基准测试,揭示大语言模型在医学知识时间意识方面的不足,包括知识衰减、历史知识回忆困难及时间不一致行为,指出整合代理搜索工具难以解决该问题。

Comments 35 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05410 2026-05-14 cs.AI cs.CL 88%

ChatSR: Multimodal Large Language Models for Scientific Formula Discovery

ChatSR:用于科学公式发现的多模态大语言模型

Yanjie Li, Lina Yu, Weijun Li, Min Wu, Liping Zhang, Jingyi Liu, Yusong Deng, Mingzhu Wan, Xin Ning

机构 * AnnLab, Institute of Semiconductors, Chinese Academy of Sciences, Beijing, China(安 lab,半导体研究所,中国科学院,北京,中国) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing, China(电子、电气与通信工程学院,中国科学院大学,北京,中国) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing 101408, China(先进交叉科学学院,中国科学院大学,北京101408,中国) College of Materials Science and Opto-Electronic Technology, University of Chinese Academy of Sciences, Beijing, 100049, China(材料科学与光电技术学院,中国科学院大学,北京100049,中国) School of Integrated Circuits, University of Chinese Academy of Sciences, Beijing 100049, China(集成电路学院,中国科学院大学,北京100049,中国)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 ChatSR通过设计专用编码器和模态对齐机制,将科学数据映射到可被大语言模型处理的表示空间,从而生成符合领域先验的数学公式,推动科学发现自动化。

Comments 14 pages,

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08257 2026-05-12 cs.CR cs.AI cs.LG 88%

Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks

对抗鲁棒性增强方法研究:用于医疗决策任务的对抗鲁棒大语言模型智能体

Saisai Hu

机构 * Pace University(帕克大学)

专题命中 领域大模型 :large language model(title);language model(title);LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 本文提出ARSM-Agent框架,通过多模块协同提升医疗决策智能体的对抗鲁棒性与安全性,实验表明其在多种攻击下攻击成功率降低至8.7%。

Comments 5 pages, 2 figures, 1 table.Accepted for oral presentation at AINIT 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00803 2026-05-04 cs.SE cs.AI cs.CL 88%

Can Coding Agents Reproduce Findings in Computational Materials Science?

编码代理能否在计算材料科学中复现研究成果?

Ziyang Huang, Yi Cao, Ali K. Shargh, Jing Luo, Ruidong Mei, Mohd Zaki, Zhan Liu, Wyatt Bunstine, William Jurayj, Somdatta Goswami, Tyrel McQueen, Michael Shields, Jaafar El-Awady, Paulette Clancy, Benjamin Van Durme, Nicholas Andrews, William Walden, Daniel Khashabi

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);foundation model(abstract)

AI总结 研究探讨编码代理在计算材料科学领域复现研究结果的能力,提出AutoMat基准测试,发现当前LLM代理在复现复杂科学流程时表现有限,存在流程不完整、方法偏差等问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05190 2026-04-30 cs.CL cs.AI cs.IR 88%

Retrieval-Augmented LLMs for Evidence Localization in Clinical Trial Recruitment from Longitudinal EHR Narratives

基于检索增强的LLM在纵向电子健康记录叙事中证据定位的临床试验招募

Ziyi Chen, Mengxian Lyu, Cheng Peng, Yonghui Wu

机构 * Department of Health Outcomes and Biomedical Informatics(健康结果与生物医学信息学系) University of Florida(佛罗里达大学)

专题命中 领域大模型 :LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了基于编码器和解码器的生成LLM在临床试验招募中的应用,通过三种策略缓解长文档处理问题,MedGemma模型在RAG策略下达到89.05%的微F1分数,提升了长期推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23002 2026-04-28 cs.AI cs.CL 88%

FormalScience: Scalable Human-in-the-Loop Autoformalisation of Science with Agentic Code Generation in Lean

FormalScience: 基于代理代码生成的可扩展人类在环科学自动形式化

Jordan Meadows, Lan Zhang, Andre Freitas

机构 * University of Manchester, UK(英国曼彻斯特大学) Idiap Research Institute, Switzerland(瑞士Idiap研究所) National Biomarker Centre, CRUK-MI, UK(英国国家生物标志物中心,CRUK-MI)

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 FormalScience通过人类在环的代理流程,使非专业领域专家以低成本生成形式化证明,构建了包含200个大学物理问题及解决方案的FormalPhysics数据集,并探讨了现代LLM在自动形式化中的局限性。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01965 2026-04-23 cs.IR cs.AI cs.CL cs.DL 88%

Do We Need Bigger Models for Science? Task-Aware Retrieval with Small Language Models

我们是否需要更大的模型用于科学?基于任务的检索与小型语言模型

Florian Kelber, Matthias Jobst, Yuni Susanti, Michael Färber

机构 * TU Dresden, Germany(德累斯顿理工大学,德国) FIZ Karlsruhe, Germany(卡尔斯鲁厄研究所,德国) ScaDS.AI, TU Dresden, Germany(ScaDS.AI,德累斯顿理工大学,德国)

专题命中 领域大模型 :language model(title,abstract);small language model(title);large language model(abstract);分类 cs.CL、cs.AI

AI总结 本文探讨小型语言模型在科学应用中的可行性,通过设计轻量级检索增强框架,结合全文科学论文和结构化元数据,展示检索与模型规模的互补性。

Comments Accepted at NSLP@LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏