arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-07-15 至 2026-07-15 共收录 18 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 18 篇

2607.00011 2026-07-15 cs.IR cs.AI cs.SE 版本更新 90%

SkillSelect-Serve: QoS-Aware Budgeted Skill Service Recommendation for LLM Agents

SkillSelect-Serve:面向小型LLM代理的预算可控且QoS感知的技能服务推荐与组合

Jingyuan Zheng, Dongjing Wang, Xin Zhang, Hao Chen, Youhuizi Li, Xudong Shen, Haiping Zhang, Butian Huang, Dongjin Yu, Guandong Xu

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出SkillSelect-Serve框架,将技能选择建模为服务推荐与组合问题,通过双粒度效用建模和预算约束优化,在35,353个技能和586个任务查询上优于固定top-k检索基线。

Comments 18 pages (14-page main text + appendices), 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12336 2026-07-15 cs.CL cs.AI cs.CY cs.ET cs.HC 新提交 90%

Evaluating Health Misinformation in Low-Resource Languages: Integrating Small Language Models with a Culturally-Sensitive Responsible NLP Framework (Bangla as a Case Study)

评估低资源语言中的健康错误信息:将小语言模型与文化敏感的负责任自然语言处理框架相结合(以孟加拉语为例)

Farnaz Farid, Raihan Alam, Al Al-Areqi, Farhad Ahamed, Muhammad Hassan Khan, Sadia Hossain, Irena Veljanova, Anika Tabassum Binte Hossain

机构 * Western Sydney University(西悉尼大学) Microsoft(微软公司) Excelsia College(埃克塞尔西亚学院) Faulconbridge Health Centre(福尔康布里奇健康中心)

专题命中 领域大模型 :language model(title,abstract);small language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

AI总结 研究针对低资源语言中健康错误信息难检测问题,提出结合小语言模型与文化敏感的负责任自然语言处理框架,以孟加拉语为例进行实验,证明Phi-4表现优,还设计新框架,为评估低资源语言错误信息提供整体视角。

Comments 39 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06294 2026-07-15 q-bio.QM cs.AI 88%

Enhancing Phenotype Recognition in Clinical Notes Using Large Language Models: PhenoBCBERT and PhenoGPT

利用大型语言模型增强临床笔记中的表型识别:PhenoBCBERT和PhenoGPT

Jingye Yang, Cong Liu, Wendy Deng, Da Wu, Chunhua Weng, Yunyun Zhou, Kai Wang

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本研究提出PhenoBCBERT和PhenoGPT两种模型,利用大型语言模型提升临床笔记中表型术语的自动识别与提取,从而推动疾病相关生物学研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12051 2026-07-15 cs.CL 新提交 86%

Agentic systems for breast cancer treatment recommendations

用于乳腺癌治疗建议的智能体系统

Vinicius Anjos de Almeida, Nícolas Henrique Borges, Leonardo Vicenzi, Helena Kociolek, Sarah Miriã de Castro Rocha, Frederico Nassif Gomes, Júlia Cristina Ferreira Ribeiro, Lucas Emanuel Silva e Oliveira

机构 * Spesia(斯佩西亚) Faculdade de Medicina, Universidade de São Paulo(圣保罗大学医学院) Laboratory of Artificial Intelligence Applied to Bioinformatics, SEPT, Universidade Federal do Paraná (UFPR)(巴拉那联邦大学人工智能应用于生物信息学实验室,SEPT) Pontifícia Universidade Católica do Paraná (PUCPR)(巴拉那天主大学)

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究评估用于乳腺癌治疗建议的智能体LLM系统,用72个真实临床病例和1147个特定病例量表,比较七种流程,最佳配置全局得分为0.594±0.025,工具使用和智能体自主性影响各异,虽能生成相关建议,但用于无监督临床使用仍不足。

Comments Under peer review. Source code available at: https://github.com/GRUPOMED4U/breast_cancer_agents_paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01241 2026-07-15 cs.CY cs.AI 版本更新 86%

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

首先,不伤害:迈向临床安全的大语言模型

David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj, Priyank Jain, Jessica Tran, Matthew Gwiazdon, Arjun Rustagi, Jenelle Jindal, Jacob M. Koshy, Vinay Kadiyala, Anup Agarwal, Bassman Tappuni, Brianna French, Sirus Jesudasen, Christopher V. Cosgriff, Rebanta Chakraborty, Jillian Caldwell, Susan Ziolkowski, David J. Iberri, Robert Diep, Rahul S. Dalal, Kira L. Newman, Kristin Galetta, J. Carl Pallais, Nancy Wei, Kathleen M. Buchheit, David I. Hong, Vartan Pahalyants, Ernest Y. Lee, Allen Shih, Tamara B. Kaplan, Vishnu Ravi, Sarita Khemani, Thomas A. Buckley, April S. Liang, Daniel Shirvani, Advait Patil, Nicholas Marshall, Kanav Chopra, Joel Koh, Adi Badhwar, Anastasia Perez, Austin J. Schoeffler, Mahbuba Tusty, Chase M. Walton, Liam G. McCoy, David J. H. Wu, Yingjie Weng, Sumant Ranji, Kevin Schulman, Nigam H. Shah, Jason Hom, Arnold Milstein, Arjun K. Manrai, Adam Rodman, Jonathan H. Chen, Ethan Goh

机构 * Harvard Combined Dermatology Program(哈佛联合皮肤科项目) Department of Dermatology, Mass General Brigham(麻省总医院皮肤科) Harvard Medical School(哈佛医学院) Stanford Center for Biomedical Informatics Research(斯坦福生物医学信息学研究中心) Stanford University(斯坦福大学) Division of Hospital Medicine, Department of Medicine, Stanford University School of Medicine(斯坦福大学医学院医院医学科) Department of Medicine, Cambridge Health Alliance(剑桥健康联盟医学科) Beth Israel Deaconess Hospital–Plymouth(贝塞斯达德acons医院-普利茅斯) Department of Medicine, University of California, San Francisco(加州大学旧金山分校医学科) Department of Neurology, Stanford University School of Medicine(斯坦福大学医学院神经科) Department of Medicine, Beth Israel Deaconess Medical Center(贝塞斯达德acons医学中心医学科) Division of Cardiology, Department of Medicine, Cambridge Health Alliance(剑桥健康联盟心脏病科) Department of Cardiovascular Medicine, Summa Health System(Summa健康系统心血管医学科) Division of Allergy, Pulmonary, and Critical Care Medicine, Department of Medicine, University of Wisconsin-Madison(威斯康星大学麦迪逊分校医学科过敏、呼吸科和危重医学科) Division of Pulmonary and Critical Care Medicine, Department of Medicine, Massachusetts General Hospital(麻省总医院呼吸科和危重医学科) Center for Immunology and Inflammatory Diseases, Department of Medicine, Massachusetts General Hospital(麻省总医院免疫和炎症疾病中心) Broad Institute of MIT and Harvard(MIT和哈佛Broad研究所) Division of Pulmonary, Critical Care, and Sleep Medicine, Cambridge Health Alliance(剑桥健康联盟呼吸科、危重医学科和睡眠医学科)

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出NOHARM基准,包含1100个初级到专科咨询案例,评估28个LLM的医疗建议安全性,发现高达22.6%的案例存在严重危害风险,其中遗漏错误占80%以上。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09483 2026-07-15 cs.AI 版本更新 83%

CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?

CrochetBench:视觉语言模型能否在钩针领域从描述转向实践?

Peiyu Li, Xiaobao Huang, Ting Hua, Nitesh V. Chawla

机构 * University of Notre Dame(诺特大学)

专题命中 领域大模型 :language model(title,abstract);large language model(abstract);分类 cs.AI

AI总结 探讨视觉语言模型在钩针领域从描述到实践的转变,采用CrochetPARADE DSL进行评估,涵盖多种任务,发现评估转变时性能下降,揭示模型局限性,为评估多模态模型过程能力提供新视角。

Comments ACL2026 main-long

Journal ref Proceedings of ACL 2026, Vol. 1, pp. 13229-13251

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07482 2026-07-15 eess.SP cs.AI 83%

VSLLaVA: a pipeline of large multimodal foundation model for industrial vibration signal analysis

VSLLaVA:一种用于工业振动信号分析的大型多模态基础模型流水线

Qi Li, Xinran Zhang, Jinfeng Huang, Hongliang He, Feibin Zhang, Zhaoye Qin, Fulei Chu

机构 * State Key Laboratory of Tribology, Department of Mechanical Engineering, Tsinghua University(摩擦学国家重点实验室,清华大学机械工程系) Department of Statistics and Data Science, Yale University(耶鲁大学统计与数据科学系)

专题命中 领域大模型 :foundation model(title);LLM(abstract);instruction tuning(abstract);分类 cs.AI

AI总结 VSLLaVA通过专家知识引导的指令微调和双模式评估框架,提升工业振动信号分析的端到端多模态模型性能。

Journal ref Advanced Engineering Informatics 76 (2026): 105023

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12454 2026-07-15 cs.LG 新提交 79%

Exploring Zero-Shot Foundation Models for Multivariate Time Series Anomaly Detection

探索用于多变量时间序列异常检测的零样本基础模型

Martin Uray, Saverio Messineo, Roland Kwitt, Stefan Huber

机构 * Salzburg University of Applied Sciences(萨尔茨堡应用科学大学) Paris Lodron University of Salzburg(萨尔茨堡巴黎洛德龙大学)

专题命中 领域大模型 :foundation model(title,abstract);分类 cs.LG

AI总结 研究多变量时间序列异常检测,探索单变量预测基础模型TimesFM的零样本应用于工业MTSAD,评估两种策略,虽未胜过基线,但发现其在捕获时间动态上过于有效致异常难区分,不过在异常边界误差有峰值,对变化点检测有前景。

Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution will be published in Computer Aided Systems Theory - EUROCAST 2026, Lecture Notes in Computer Science, Springer

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11954 2026-07-15 cs.LG cs.AI cs.IR 新提交 79%

Graph-Constrained Policy Learning for Extreme Clinical Code Prediction

用于极端临床代码预测的图约束策略学习

Amritpal Singh, Sebastian Torres, Khawar Shakeel, Syed Ahmad Chan Bukhari

专题命中 领域大模型 :SFT(abstract,abstract_cn);language model(abstract);分类 cs.AI、cs.LG

AI总结 研究针对临床代码预测任务,提出图约束遍历策略,将其转换为有限时域决策过程。单一语言模型逐层选节点得叶代码,实现极端多标签预测到子集决策的转换。实验表明该策略优于平面基线,增加监督数据可提升性能,简单图约束策略学习效果良好。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12340 2026-07-15 cs.SE cs.CR 新提交 78%

Skills That Don't Exist: A Large-Scale Study of Hallucinated Skill Recommendation in LLM Agents

不存在的技能:对大语言模型代理中幻觉技能推荐的大规模研究

Weifeng Yuan, Wenbo Guo, Feng Dong, Haoyu Wang, Yang Liu

专题命中 领域大模型 :LLM(title,abstract)

AI总结 研究大语言模型代理中技能名称幻觉漏洞,通过大规模测量发现各配置均有此问题,系统生成众多幻觉名称且非随机,测试的模型级防御存在安全与可用性冲突,修复需全生态系统结构改变。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11464 2026-07-15 cs.IR cs.AI cs.CL cs.DB 交叉投稿 73%

FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis

公平的GraphRAG:一种用于语义数据分析的检索增强生成方法

Marlena Flüh, Soo-Yon Kim, Carolin Victoria Schneider, Sandra Geisler

机构 * RWTH Aachen University(亚琛RWTH大学) University Hospital RWTH Aachen(亚琛RWTH大学医院)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究针对现有RAG方法缺乏结构化公平化的问题,引入公平的GraphRAG框架,以公平数字对象为基本单元,利用大语言模型支持相关构建与提取,应用于生物医学数据集,显著提升问答效果,展示了公平数据实践与图检索技术结合的可行性。

Comments Accepted at the IEEE International Conference on Knowledge Graph, 2025. Corrects an error in the published abstract: the evaluation dataset is RNA-sequencing data, not single-cell data

Journal ref 2025 IEEE International Conference on Knowledge Graph (ICKG), Limassol, Cyprus, 2025, pp. 90-97

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02208 2026-07-15 cs.CL cs.AI cs.IR cs.SE 73%

Towards AI Evaluation in Domain-Specific RAG Systems: The AgriHubi Case Study

面向领域特定RAG系统的AI评估:AgriHubi案例研究

Md. Toufique Hasan, Ayman Asad Khan, Mika Saari, Vaishnavi Bankhele, Pekka Abrahamsson

机构 * Faculty of Information Technology and Communication Sciences(信息科技与通讯科学学院)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 AgriHubi通过整合芬兰农业文档与开放模型,结合来源 grounding 和用户反馈,提升了农业决策支持系统的回答完整性、语言准确性和可靠性。

Comments 6 pages, 2 figures, submitted to MIPRO 2026

Journal ref 2026 49th MIPRO ICT and Electronics Convention (MIPRO), Opatija, Croatia, pp. 989-994

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12048 2026-07-15 cs.CV cs.AI cs.CL 新提交 62%

An Empirical Analysis of Continual Learning for Heterogeneous Medical Visual Question Answering

异构医学视觉问答持续学习的实证分析

Mai A. Shaaban, Tausifa Jan Saleem, Alaa Mohamed, Dilnaz Utemissova, Ufaq Khan, Mohammad Yaqub

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Alexandria University(亚历山大大学)

专题命中 领域大模型 :language model(abstract);分类 cs.CL、cs.AI

AI总结 研究异构医学视觉问答持续学习,通过系统评估多种临床目标任务,探究现有CL方法减轻灾难性遗忘能力、对任务排序敏感性及低秩适应参数演变,发现交错不同任务时现有方法难维持稳定性 - 可塑性平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12122 2026-07-15 cs.LG 新提交 57%

An Agentic AI Scientific Community for Automated Neural Operator Discovery

用于自动神经算子发现的智能体人工智能科学社区

Luis Loo, Ulisses Braga-Neto

机构 * Texas A&M University(德克萨斯农工大学)

专题命中 领域大模型 :LLM(abstract);分类 cs.LG

AI总结 该研究基于人工智能科学社区提出智能体方法用于自动神经算子发现,通过虚拟实验室中三个智能体协作,共享通用词汇表,在五个问题上评估,能发现高精度低参数架构,揭示语言模型智能体对保持多样性的作用及神经算子无免费午餐定理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12113 2026-07-15 cs.DC cs.AI 新提交 57%

Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap

迈向可信自主科学:两年社区路线图

Rafael Ferreira da Silva, Milad Abolhasani, Peter Beaucage, Laura Biven, Michael Bussmann, Kyle Chard, Ryan Coffee, Stephen DeWitt, Sagar Dolas, Carrie Eckert, David Elbert, Ian Foster, Tirthankar Ghosal, Anna Giannakou, Tom Gibbs, Leslie Hamilton, Glenn Lockwood, Theresa Mayer, Ben Mintz, Raffi Nazikian, Sal Nimer, Amanda Randles, Woong Shin, Sreenivas Rangan Sukumar, Frédéric Suter, Mitra Taheri, Michela Taufer, Draguna Vrabie

机构 * U.S. Department of Energy, Office of Science, Office of Advanced Scientific Computing Research(美国能源部科学办公室高级科学计算研究办公室)

专题命中 领域大模型 :foundation model(abstract);分类 cs.AI

AI总结 该研究针对自主科学领域发展现状及问题,更新路线图围绕七个维度,评估原有里程碑并新增四个,规划两年发展路径,第一年聚焦接口等搭建验证框架,第二年针对联盟等,强调基层网络的互操作性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15748 2026-07-15 cs.LG cs.CV 版本更新 57%

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors

以大型多模态模型作为事后校正器的视觉物种识别

Tian Liu, Anwesha Basu, James Caverlee, Shu Kong

机构 * Texas A&M University(德克萨斯A&M大学) University of Macau(澳门大学) Institute of Collaborative Innovation(协同创新研究院)

专题命中 领域大模型 :prompting(abstract);分类 cs.LG

AI总结 研究视觉物种识别问题,比较少样本学习专家模型与大型多模态模型后发现后者虽有不足但有互补优势,进而提出事后校正框架,利用多模态提示策略提升少样本学习专家模型准确率,该框架具有通用性和有效性。

Comments website and code: https://tian1327.github.io/POC

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12896 2026-07-15 cs.CV 新提交 50%

UniMedSeg: Unified In-Context Learning for Multi-Paradigm 2D/3D Medical Image Segmentation

UniMedSeg:用于多范式2D/3D医学图像分割的统一上下文学习

Yunzhou Li, Jiesi Hu, Yanwu Yang, Hanyang Peng, Chenfei Ye, Jianfeng Cao, Yixuan Yuan, Ting Ma

机构 * Harbin Institute of Technology at Shenzhen(哈尔滨工业大学(深圳)) Peng Cheng Laboratory(鹏城实验室) University Hospital Tübingen(图宾根大学医院) German Center for Mental Health(德国心理健康中心) Chinese University of Hong Kong(香港中文大学)

专题命中 领域大模型 :foundation model(abstract)

AI总结 研究针对医学图像分割基础模型存在的问题,提出以Transformer为中心的UniMedSeg框架,通过映射多种信息到共享序列空间联合学习异构医学监督,引入解耦分割注意力克服内存瓶颈,在多类分割任务中实现最优性能且无需特定任务微调。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31946 2026-07-15 cs.CV 版本更新 50%

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration

世界叙事模型:从像素采样到物理世界编排的高度可控视频生成范式转变

Ye Chen, Xuanhong Chen, Yupeng Zhu, Liming Tan, Zhewen Wan, Yuxuan Xiong, Tielong Wang, Jinfan Liu, Wuze Zhang, Xiongzhen Zhang, Feifei Li, Xianglin Luo, Zhehan Zhao, Zhifan Zhang, Laisheng Kou, Zhujin Liang, Yugang Chen, Muchun Chen, Xu Miao, Yijing Zhang, Xiaojie Sheng, Qiang Hu, Jialiang Chen, Weimin Zhang, Wenjun Zhang, Bingbing Ni

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 领域大模型 :foundation model(abstract)

AI总结 提出世界叙事模型(WNM),将视频生成解耦为结构化物理叙事与像素渲染,通过协同代理将多模态输入转化为可编辑的4D世界表示,驱动基础模型生成符合创作者意图的视频,大幅提升可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏