arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-04-01 至 2026-04-01 共收录 23 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 23 篇

2603.29149 2026-04-01 cs.AI cs.DB 89%

Knowledge database development by large language models for countermeasures against viruses and marine toxins

利用大语言模型开发知识库以应对病毒和海洋毒素的对策

Hung N. Do, Jessica Z. Kubicek-Sutherland, S. Gnanakaran

机构 * Physical Chemistry and Applied Spectroscopy Group, Chemistry Division, Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室物理化学与应用光谱学组,化学部)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文利用大语言模型构建病毒和海洋毒素的综合数据库,通过人工输入和模型交叉验证,设计交互式网页以提高决策效率。

Comments Clearance: 26-T-0967 (DOW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03004 2026-04-01 cs.CL cs.AI 88%

SemioLLM: Evaluating Large Language Models for Diagnostic Reasoning from Unstructured Clinical Narratives in Epilepsy

SemioLLM:评估用于癫痫诊断推理的大型语言模型在无结构临床叙述中的表现

Meghal Dani, Muthu Jeyanthi Prakash, Filip Rosa, Zeynep Akata, Stefanie Liebe

机构 * University of Tübingen(蒂宾根大学) Technical University of Munich(慕尼黑工业大学) University Clinic Tübingen(蒂宾根大学医院) Hertie Institute for Clinical Brain Research(赫蒂临床脑研究所) Excellence Cluster Machine Learning, Tübingen University(蒂宾根大学机器学习卓越集群) Hertie Institute for AI in Brain Health (Hertie AI)(赫蒂人工智能脑健康研究所)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI

AI总结 本文评估了八个大型语言模型在癫痫诊断任务中的表现,通过过滤和标准化 seizure 描述短语,将其映射到七个可能的癫痫发作起始区。结果显示,经过提示工程后,多数模型表现接近临床水平,但临床情境下的表现受语言上下文影响显著,需改进模型的可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16790 2026-04-01 cs.SE cs.AI 87%

InCoder-32B: Code Foundation Model for Industrial Scenarios

InCoder-32B:面向工业场景的代码基础模型

Jian Yang, Wei Zhang, Jiajun Wu, Junhang Cheng, Shawn Guo, Haowen Wang, Weicheng Gu, Yaxin Du, Joseph Li, Fanglin Xu, Yizhi Li, Lin Jing, Yuanbo Wang, Yuhan Gao, Ruihao Gong, Chuan Hao, Ran Tao, Aishan Liu, Tuney Zheng, Ganqu Cui, Zhoujun Li, Mingjie Tang, Chenghua Lin, Wayne Xin Zhao, Xianglong Liu, Ming Zhou, Bryan Dai, Weifeng Lv

机构 * Beihang University(北京航空航天大学) IQuest Research Shanghai Jiao Tong University(上海交通大学) ELLIS University of Manchester(曼彻斯特大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Langboat(朗博特)

专题命中 领域大模型 :foundation model(title,abstract);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 本文提出InCoder-32B,首个统一芯片设计、GPU内核优化等工业领域的32B参数代码基础模型,通过高效架构和多阶段训练方法,在通用和工业任务中均取得优异性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18602 2026-04-01 cs.NE cs.AI cs.LG 86%

LLM-Meta-SR: In-Context Learning for Evolving Selection Operators in Symbolic Regression

LLM-Meta-SR: 上下文学习用于符号回归中演化的选择算子

Hengzhe Zhang, Qi Chen, Bing Xue, Wolfgang Banzhaf, Mengjie Zhang

机构 * Centre for Data Science and Artificial Intelligence & School of Engineering and Computer Science, Victoria University of Wellington(惠灵顿维多利亚大学数据科学与人工智能中心及工程与计算机科学学院) Department of Computer Science and Engineering, Michigan State University(密歇根州立大学计算机科学与工程系)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种元学习框架,使LLM自动设计进化符号回归算法的选择算子,解决现有方法中语义指导不足和代码膨胀问题,实验表明LLM设计的选择算子优于九个专家设计基线,达到最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10747 2026-04-01 cs.CL 81%

Codebook LLMs: Evaluating LLMs as Measurement Tools for Political Science Concepts

代码本LLM:评估LLM作为政治科学概念测量工具

Andrew Halterman, Katherine A. Keith

机构 * Michigan State University(密歇根州立大学) Williams College(威廉姆斯学院)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);instruction tuning(abstract)

AI总结 本文评估LLM作为政治科学概念测量工具的有效性,提出五阶段框架,通过代码本和预训练LLM验证其测量准确性,并展示通过监督指令微调提升性能的成果。

Comments Version 2 (v1 Presented at PolMeth 2024)

Journal ref Polit. Anal. 34 (2026) 188-204

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21551 2026-04-01 cs.SE 80%

AI In Cybersecurity Education -- Scalable Agentic CTF Design Principles and Educational Outcomes

AI在网络安全教育中的应用:可扩展的智能体CTF设计原则与教育成效

Haoran Xi, Minghao Shao, Kimberly Milner, Venkata Sai Charan Putrevu, Nanda Rani, Meet Udeshi, Prashanth Krishnamurthy, Brendan Dolan-Gavitt, Siddharth Garg, Sandeep Kumar Shukla, Farshad Khorrami, Alon Hillel-Tuch, Muhammad Shafique, Ramesh Karri

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文研究了基于大语言模型的CTF竞赛设计,探讨了自主性水平和参与者知识背景对问题解决性能和学习行为的影响,提出自主性特定的评分标准和证据要求,以提升网络安全教育的可扩展性和有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29946 2026-04-01 cs.LG 79%

Real-Time Explanations for Tabular Foundation Models

表格基础模型的实时解释

Luan Borges Teodoro Reis Sena, Francisco Galuppo Azevedo

机构 * Kunumi Institute(Kunumi研究所) Universidade Federal de Minas Gerais(米纳斯吉拉斯联邦大学)

专题命中 领域大模型 :foundation model(title,abstract);分类 cs.LG

AI总结 本文提出ShapPFN,通过将Shapley值回归整合到架构中,实现预测与解释的单次前向传递,提升解释效率和质量。

Comments Accepted at the 2nd DATA4Science Workshop at ICLR 2026, Rio de Janeiro, Brazil. OpenReview: https://openreview.net/forum?id=StSMBSZqxx

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29231 2026-04-01 cs.AI 79%

Beyond pass@1: A Reliability Science Framework for Long-Horizon LLM Agents

超越pass@1:面向长时间跨度LLM代理的可靠性科学框架

Aaditya Khanal, Yangyang Tao, Junxiu Zhou

专题命中 领域大模型 :LLM(title,abstract);分类 cs.AI

AI总结 本文提出一个可靠性科学框架,用于评估长时间跨度的LLM代理,通过四个指标揭示任务持续时间增长时可靠性与能力的差异,发现可靠性衰减、方差放大因子、渐进退化分数和熔断点等指标对评估模型的长期表现至关重要。

Comments 23 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28986 2026-04-01 cs.AI cs.LG cs.MA 79%

Mimosa Framework: Toward Evolving Multi-Agent Systems for Scientific Research

Mimosa框架:迈向科学研究的演进多智能体系统

Martin Legrand, Tao Jiang, Matthieu Feraud, Benjamin Navet, Yousouf Taghzouti, Fabien Gandon, Elise Dumont, Louis-Félix Nothias

机构 * Université Côte d’Azur(蔚蓝海岸大学) CNRS(法国国家科学研究中心) ICN(化学研究所) Interdisciplinary Institute for Artificial Intelligence (3iA) Côte d’Azur(蔚蓝海岸跨学科人工智能研究所) Inria(法国国家信息与自动化研究所) I3S(信息系统与软件工程实验室)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 Mimosa框架通过自动合成任务特定的多智能体工作流并迭代优化,提升科学研究的适应性与效率,其在ScienceAgentBench上实现43.1%的成功率,展示了动态工作流演进的优势。

Comments 48 pages, 4 figures, 1 table. Clean arXiv version prepared. Includes main manuscript plus appendix/supplementary-style implementation details and prompt listings. Dated 30 March 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29890 2026-04-01 cs.HC cs.AI 77%

Interview-Informed Generative Agents for Product Discovery: A Validation Study

基于访谈的生成代理在产品发现中的应用:一项验证研究

Zichao Wang, Alexa Siu

机构 * Adobe Research(Adobe研究院)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文探讨了基于访谈的生成代理在概念测试场景中模拟用户反馈的有效性,发现代理在群体层面响应分布上准确但个体层面不够精准,为产品开发流程中合理整合模拟提供了启示。

Comments CHI 2026 Honourable Mention

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29373 2026-04-01 cs.CL 77%

Beyond Idealized Patients: Evaluating LLMs under Challenging Patient Behaviors in Medical Consultations

超越理想化患者:在医疗咨询中评估LLM在挑战性患者行为下的表现

Yahan Li, Xinyi Jie, Wanjia Ruan, Xubei Zhang, Huaijie Zhu, Yicheng Gao, Chaohao Du, Ruishan Liu

机构 * University of Southern California(南加州大学)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文研究了医疗咨询中常见的挑战性患者行为,定义了四种行为类别,并提出了CPB-Bench基准测试集,评估了多种LLM在处理这些行为时的表现,发现模型在处理矛盾或医学上不合理的患者信息时存在困难。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01838 2026-04-01 cs.CL 77%

AXE: Low-Cost Cross-Domain Web Structured Information Extraction

AXE:低成本跨领域网络结构化信息提取

Abdelrahman Mansour, Khaled W. Alshaer, Moataz Elsaban

机构 * Faculty of Computers & Artificial Intelligence - Cairo University(开罗大学计算机与人工智能学院) Ejada Systems Fawry Integrated Systems Microsoft(微软)

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 AXE通过将HTML DOM视为需修剪的树结构,利用专用修剪机制提取高密度上下文,使小型模型生成精确结构化输出,并通过GXR确保提取可追溯性,实现跨领域信息提取的高性能与低成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29039 2026-04-01 astro-ph.IM astro-ph.HE physics.data-an 76%

AI Cosplaying as Astrophysicists: A Controlled Synthetic-Agent Study of AI-Assisted Astrophysical Research Workflows

AI 伪装成天体物理学家:一个受控的合成智能体研究,探讨AI辅助天体物理研究流程

Chun Huang

专题命中 领域大模型 :language model(abstract,comments);LLM(abstract);large language model(abstract)

AI总结 本文通过模拟144名合成研究者完成2592项日常天体物理研究任务,评估AI在不同任务类型和使用风格下的效果,发现AI辅助在特定任务和使用策略下能提升效率,但对推导密集型任务存在风险。

Comments 16 pages, 8 figures. Began as an April Fools' idea, regrettably became a real methods paper, No language models were physically harmed. GitHub repository of this work: https://github.com/ChunHuangPhy/agent_astro

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29141 2026-04-01 cs.CY 75%

Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education

现代化真实数据:四种转变以提升教育中人工智能可靠性和有效性

Danielle R. Thomas, Conrad Borchers, Kirk P. Vanacore, Kenneth R. Koedinger, René F. Kizilcec

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文探讨了教育AI中真实数据质量的挑战,提出通过诊断信号、透明报告、偏见审计和有效性证据四种转变提升可靠性和有效性。

Comments Accepted as full paper to the 27th International Conference on Artificial Intelligence in Education (AIED 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29211 2026-04-01 cs.AI cs.CL cs.CV 73%

Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems

玄武:将通用多模态模型演进为内容生态系统工业级基础模型

Zhiqian Zhang, Xu Zhao, Xiaoqing Xu, Guangdong Liang, Weijia Wang, Xiaolei Lv, Bo Li, Jun Gao

机构 * Computational Intelligence Dept, Hello Group Inc(Hello Group Inc. 计算智能部)

专题命中 领域大模型 :foundation model(abstract);post-training(abstract);分类 cs.CL、cs.AI

AI总结 本文提出Xuanwu VL-2B,通过三阶段训练流程平衡业务专精与通用能力,实现7项多模态指标平均67.90分,7项业务审核任务召回率94.38%,在对抗性OCR场景中超越Gemini-2.5-Pro。

Comments 41 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29366 2026-04-01 cs.AI 70%

AI-Generated Prior Authorization Letters: Strong Clinical Content, Weak Administrative Scaffolding

人工智能生成的预先授权信:强大的临床内容,薄弱的行政框架

Moiz Sadiq Awan, Maryam Raza

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究评估了三种LLM在45个合成场景中生成的预先授权信,发现其临床内容扎实但行政要求方面存在不足,提示需加强行政精度。

Comments 11 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28926 2026-04-01 eess.SY cs.SY 67%

A Computational Framework for Cross-Domain Mission Design and Onboard Cognitive Decision Support

跨领域任务设计与 onboard 认知决策支持的计算框架

J. de Curtò, Adrianne Schneider, Ricardo Yanez, María Begara, Álvaro Rodríguez, Javier López, Martina Fraga, Ignacio Gómez, Arman Akdag, Sumit Kulkarni, Siddhant Nair, Kiyan Govender, Eian Wratchford, Eli Lynskey, Seamus Dunlap, Cooper Nervick, Nicolas Tête, Rocío Fernández, Pablo González, Elena Municio, I. de Zarzà

专题命中 领域大模型 :LLM(abstract);foundation model(abstract)

AI总结 本文提出统一计算方法,评估七种异构任务架构中自主性必要度约束,引入自主性必要度评分,并评估基于 LLM 的自主任务决策支持层,展示基础模型在高 ANS 任务中的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28802 2026-04-01 cs.DL 67%

Interactive Evidence Maps for Visualizing and Understanding Systematic Reviews

交互式证据地图用于可视化和理解系统综述

Aditi Mallavarapu, Rohan Khandare, Mokshagna Kadiyala, Neelesh Yaddanapudi, Noah L. Schroeder, Shan Zhang, Jessica R. Gladstone

专题命中 领域大模型 :large language model(abstract);language model(abstract)

AI总结 本文提出交互式证据地图,通过大语言模型提取主题模型,帮助研究人员动态探索和分析综述数据,提升透明度和发现文献中的模式与空白。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29893 2026-04-01 cs.HC cs.AI cs.CL cs.MA 62%

Perfecting Human-AI Interaction at Clinical Scale. Turning Production Signals into Safer, More Human Conversations

在临床规模上完善人机交互。将生产信号转化为更安全、更人性化的对话

Subhabrata Mukherjee, Markel Sanz Ausin, Kriti Aggarwal, Debajyoti Datta, Shanil Puri, Woojeong Jin, Tanmay Laud, Neha Manjunath, Jiayuan Ding, Bibek Paudel, Jan Schellenberger, Zepeng Frazier Huo, Walter Shen, Nima Shirazian, Nate Potter, Sathvik Perkari, Darya Filippova, Anton Morozov, Austin Mease, Vivek Muppalla, Ghada Shakir, Alex Miller, Juliana Ghukasyan, Mariska Raglow-Defranco, Maggie Taylor, Herprit Mahal, Jonathan Agnew

机构 * Hippocratic AI

专题命中 领域大模型 :LLM(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种基于真实患者与AI交互数据的框架,通过分析实时信号提升医疗AI的安全性和可靠性,减少ASR错误,提高患者体验。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29557 2026-04-01 cs.AI cs.CL 62%

FlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature Exploration

FlowPIE: 测试时的科学思想演变与流引导文献探索

Qiyao Wang, Hongbo Wang, Longze Chen, Zhihao Yang, Guhong Chen, Hamid Alinejad-Rokny, Hui Li, Yuan Lin, Min Yang

机构 * University of Chinese Academy of Sciences(中国科学院大学) Dalian University of Technology(大连理工大学) UNSW Sydney(新南威尔士大学悉尼分校) Shenzhen University of Advanced Technology(深圳理工大学) Xiamen University(厦门大学)

专题命中 领域大模型 :LLM(abstract);分类 cs.CL、cs.AI

AI总结 本文提出FlowPIE框架,通过流引导的蒙特卡洛树搜索扩展文献轨迹,结合LLM生成奖励模型指导适应性检索,生成高质量且多样化的初始种群,进而通过选择、交叉和变异实现测试时的科学思想演变。

Comments 30 pages, 11 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02789 2026-04-01 cs.CV cs.AI cs.LG 62%

Align Your Query: Representation Alignment for Multimodality Medical Object Detection

对齐你的查询:多模态医学目标检测中的表示对齐

Ara Seo, Bryan Sangwoo Kim, Hyungjin Chung, Jong Chul Ye

机构 * KAIST AI(韩国科学技术院人工智能学院) EverEx

专题命中 领域大模型 :pretraining(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过表示对齐提升多模态医学目标检测性能,引入模态令牌和QueryREPA预训练阶段,使查询表示更适应模态特征,提升检测精度。

Comments Project page: https://araseo.github.io/alignyourquery/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29062 2026-04-01 cs.CR cs.AI 57%

CivicShield: A Cross-Domain Defense-in-Depth Framework for Securing Government-Facing AI Chatbots Against Multi-Turn Adversarial Attacks

CivicShield:一种跨领域纵深防御框架,用于保护面向政府的AI聊天机器人免受多轮对抗攻击

KrishnaSaiReddy Patil

专题命中 领域大模型 :LLM(abstract);分类 cs.AI

AI总结 本文提出CivicShield框架,通过多层防御机制提升政府AI聊天机器人安全性,有效检测多轮对抗攻击,理论分析显示多层防御可将攻击概率降低1-2个数量级。

Comments 25 pages, 17 tables, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11872 2026-04-01 cs.CV 50%

PRS-Med: Position Reasoning Segmentation in Medical Imaging

PRS-Med: 医学影像中的位置推理分割

Quoc-Huy Trinh, Minh-Van Nguyen, Jun Zeng, Debesh Jha, Ulas Bagci

机构 * Aalto University(阿尔托大学) Northwestern University(西北大学) Technical University of Denmark(丹麦技术大学) Chongqing University of Posts and Telecommunications(重庆邮电大学) University of South Dakota(南达科他大学)

专题命中 领域大模型 :language model(abstract)

AI总结 PRS-Med提出一种基于位置推理的医学影像分割框架,通过整合视觉-语言模型与分割解码器,提升临床诊断的准确性与可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏