arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 840 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 840 篇

2606.22306 2026-06-23 cs.SE cs.AI 新提交 88%

Leveraging Large Language Models to Obscure Code Stylometry: A Comparative Study of GPT-3.5 and GPT-4

利用大型语言模型混淆代码风格计量学:GPT-3.5与GPT-4的对比研究

Saman Pordanesh, Benjamin Tan

机构 * UNIVERSITY OF CALGARY(卡尔加里大学)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 研究GPT-3.5和GPT-4在保持代码功能的同时混淆代码风格计量学的能力,通过随机森林分类器评估不同提示工程策略的效果,发现单次与多次提示方法存在显著差异,并强调了详细结构化提示的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10467 2026-06-10 cs.CL 新提交 88%

Large Language Models as Modal Models in Linguistics

大语言模型作为语言学中的模态模型

Haruto Suzuki, Saku Sugawara

机构 * Keio University(庆应义塾大学) National Institute of Informatics(国立信息学研究所) University of Tokyo(东京大学)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 本文应用科学哲学中的模态建模框架,论证大语言模型作为最小模型具有真正的认知价值,能提供“如何可能解释”,但当前尚不满足“如何实际解释”的条件,其解释力位于两者之间的连续统上。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09632 2026-06-09 cs.CL 新提交 88%

Civil Court Simulation with Large Language Models

基于大型语言模型的民事法庭模拟

Yifan Chen, Haitao Li, Kaiyuan Zhang, Yueyue Wu, Qingyao Ai, Yiqun Liu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 提出多智能体民事法庭模拟框架,通过五阶段审判程序、记忆模块和法规检索实现可靠判决,在责任分配和多项裁决上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07570 2026-06-09 cs.DL cs.LG 新提交 88%

Can LLMs extract scientific consensus? A case study in high-temperature superconductivity

LLMs能否提取科学共识?以高温超导为例

Mouyang Cheng, Wenhao He, Zhuotao Jin, Bowen Yu, Ju Li, Boris Kozinsky, Yao Wang, Pavel Volkov, Liangzi Deng, Ching-Wu Chu, Xiao-Gang Wen, Mingda Li

机构 * Center for Computational Science and Engineering, MIT(MIT计算科学与工程中心) Department of Materials Science and Engineering, MIT(MIT材料科学与工程系) Department of Physics, MIT(MIT物理系) Department of Nuclear Science and Engineering, MIT(MIT核科学与工程系) John A. Paulson School of Engineering and Applied Sciences, Harvard University(哈佛大学约翰·A·保罗森工程与应用科学学院) Department of Chemistry, Emory University(埃默里大学化学系) Department of Physics, University of Connecticut(康涅狄格大学物理系) Department of Physics and Texas Center for Superconductivity, University of Houston(休斯顿大学物理系和德克萨斯超导中心)

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本研究以高温超导领域为测试平台,利用近18,000篇高被引文献构建知识图谱,发现LLM提取的表征能恢复出连贯且物理可解释的结构,表明LLM可作为解码竞争性科学知识的可扩展工具。

Comments 23 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04307 2026-08-06 cs.CL cs.AI cs.MA 新提交 88%

MIDAS: Multi-LLM Iterative Data-Adaptive Summarization

MIDAS:多大语言模型迭代数据自适应摘要生成

Karen Lee, Dhanashree Balaram, Seojun Shon, Umair Rasheed

机构 * Volkswagen Group Innovation(大众集团创新公司)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 针对现有自动提示词优化方法无法适配多样摘要需求的问题,提出多大语言模型迭代数据自适应摘要生成框架MIDAS,在企业客户工单摘要任务中优于CriSPO等方法,还具备跨模型与跨领域泛化能力。

Comments Accepted at the 20th International Conference on Document Analysis and Recognition (ICDAR 2026). 17 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25630 2026-07-29 cs.CL cs.AI cs.HC 新提交 88%

A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries

用于基于大语言模型的科学摘要简化的人工参与语料库

Kyuri Im, Michael Färber

机构 * ScaDS.AI, Technische Universität Dresden(ScaDS.AI,德累斯顿工业大学)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究基于大语言模型的科学文本简化,提出人工参与工作流程,先以GPT-4o-mini生成基线简化,再经不同阶段读者与专家反馈编辑,发布语料库及评估结果,支持跨学科科学交流简化系统相关工作。

Comments Accepted at FGWM@KI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24598 2026-06-24 cs.SE cs.AI cs.LG 新提交 88%

Toward Self-Evolution-Ready Workflow Harnesses: A Reversible Migration Path and Convertibility Taxonomy for Expert LLM Pipelines

面向自进化就绪工作流工具:专家大语言模型管道的可逆迁移路径与可转换性分类法

Yimo Lin, Zhen Zhang, Yibin Li

机构 * Yunyong Century (Beijing) Artificial Intelligence Technology Co., Ltd.(云涌世纪(北京)人工智能科技有限公司) Tsinghua University(清华大学)

专题命中 领域大模型 :LLM(title,summary_cn);分类 cs.AI、cs.LG

AI总结 针对专家验证的“LLM+脚本”工作流静态、无法自适应的问题,提出一种基于绞杀者模式的可逆迁移路径,将遗留工作流重构为可组合、类型化、可审计的阶段,并引入三级可转换性分类法(A/B/C)作为路由阶段诊断工作流就绪状态。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13024 2026-06-12 cs.LG cs.AI 新提交 88%

CausalMoE: A Billion-Scale Multimodal Foundation Model for Granger Causal Discovery with Pattern-Routed Heterogeneous Experts

CausalMoE:基于模式路由异构专家的十亿规模多模态基础模型用于格兰杰因果发现

Bo Liu, Di Dai, Jingwei Liu, Jiarui Jin, Xiaocheng Fang, Guangkun Nie, Hongyan Li, Shenda Hong

机构 * State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院通用人工智能国家重点实验室) National Institute of Health Data Science, and Institute for Artificial Intelligence, Peking University(北京大学健康医疗大数据国家研究院、人工智能研究院)

专题命中 领域大模型 :foundation model(title,abstract);LLM(summary_cn,abstract_cn);分类 cs.AI、cs.LG

AI总结 提出CausalMoE,一种十亿规模多模态格兰杰因果基础模型,通过模式路由混合异构专家解耦动态机制,结合因果自注意力与LLM/VLM先验,实现稀疏因果图恢复,在监督和少样本场景中达到最优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12407 2026-08-14 cs.DB 新提交 88%

From Relational and Property Graph Data to Large Language Models

从关系型和属性图数据到大语言模型

Malcolm Crowe, Fritz Laux

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract)

AI总结 本文提出一种结合关系型与属性图数据的数据管理服务器,通过引用值替代外键协调两种模型,将其知识模型提供给大语言模型生成工具,还给出基于三元组的知识库示例以验证该系统的可行性。

Comments 7 pages, 4 figures, 4 tables

Journal ref IARIA Congress 2026 : The 2026 IARIA Annual Congress on Frontiers in Science, Technology, Services, and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03899 2026-08-05 cs.IR 新提交 88%

ATLAS: Learning to Recommend Across Unseen Domains

ATLAS:学习在未见域间进行推荐

Pervez Shaik, Prosenjit Biswas, Abhinav Thorat, Ravi Kolla, Niranjan Pedanekar

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);pretraining(abstract)

AI总结 ATLAS是一种多源推荐域泛化框架,无需目标域适配或LLM预训练,可从异构源域学习域不变表示,在未见域零样本推荐中优于多种基线,HitRate平均提升24%。

Comments 18 pages, 5 figures, 14 tables. Includes appendix with proofs and additional experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03187 2026-08-05 cs.NE 新提交 88%

NeuroMosaic: Anatomically Grounded Multimodal Large Language Modeling for Molecularly Aware Glioma Reasoning from 3D MRI and Clinical Narratives

NeuroMosaic:基于解剖学的多模态大语言模型,用于从3D MRI和临床叙事中进行分子感知的胶质瘤推理

Yantong Liu, Zheyu Zhang, Runpeng Liu, Mu Xitang, Seong-Yoon Shin, Hyun-Ae Lee

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract)

AI总结 该研究针对多模态医疗大语言模型在神经肿瘤学中的结构缺陷,提出NeuroMosaic模型,通过多模块架构实现胶质瘤的准确推理,在多个数据集上取得优异性能,验证了解剖学索引路由机制的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02617 2026-08-05 cs.CL cs.AI cs.LG 新提交 88%

Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

偏好而非安全:成对偏好并非临床安全的可靠替代指标

Fay Elhassan, David Sasu, Alexandra Kulinkina, Lars Henning Klein, Mary-Anne Hartley

机构 * EPFL(洛桑联邦理工学院) LiGHT Laboratory(LiGHT实验室)

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究表明临床医生成对偏好并非LLM临床安全的可靠替代指标,其排名靠前的模型仍存在大量临床安全失败,提出结合偏好与安全反馈的临床调整排名方法,支持分离偏好与安全的评估实践。

Comments 27 pages,10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24116 2026-07-28 cs.CR 新提交 88%

A Cybersecurity MLPS Large Language Model with Multi-Path Retrieval Fusion

一种具有多路径检索融合的网络安全多级保护方案大语言模型

Qian Li, Zhenyan Qi, Liang Shen, Yuan Zhang, Yifan Wan, Junyuan Ma, Yining Hu

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract)

AI总结 针对网络安全MLPS,提出集成多种检索策略的大语言模型框架,结合分层、基于树及基于词元化的匹配检索,减少无关干扰,采用多维加权评分评估,在十个典型问题实验中该特定领域模型总分更高。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20284 2026-07-23 cs.CV 新提交 88%

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

用于遥感图像理解的多模态大语言模型:领域特定还是通用?

Qiwei Ma, Chunping Qiu, Xinjun Cheng, Xiaoyu Zhang, Puhong Duan, Ke Yang, Xudong Kang, Shutao Li

机构 * School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) Intelligent Game and Decision Lab (IGDL)(智能游戏与决策实验室) Yuelushan Center for Industrial Innovation(岳麓山工业创新中心)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract)

AI总结 本文针对遥感图像理解的多模态大语言模型展开研究,通过系统调查和评估,比较其与通用模型在不同任务上的表现,发现当前模型存在局限,进而给出未来方向,为开发相关模型提供系统参考。

Comments 27 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17397 2026-07-21 physics.soc-ph 新提交 88%

The unintended consequences of large language models as a labor-augmenting technology in science

大语言模型作为科学领域劳动力增强技术的意外后果

Eamon Duede, Kevin Gross, M. J. Crockett, Carl Bergstrom

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract)

AI总结 研究大语言模型作为科学领域劳动力增强技术的意外后果,用简单数学模型说明,指出其改变研究精力分配平衡,使研究人员对发表内容选择性改变,还提高时间机会成本,让论文打磨不那么彻底,降温了相关美好期望。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20002 2026-06-19 cs.LG cs.AI cs.CL 新提交 88%

Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

Connect the Dots:通过强化学习训练具备跨域泛化能力的长期生命周期智能体

Yanxi Chen, Weijie Shi, Yuexiang Xie, Boyi Hu, Yaliang Li, Bolin Ding, Jingren Zhou

机构 * Alibaba Group(阿里巴巴集团)

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出Connect the Dots框架,通过端到端强化学习训练LLM在长期任务中自我更新上下文并泛化到新领域,实验验证了跨域泛化能力。

Comments Work in progress; we will continuously update the codebase and arXiv version

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09818 2026-08-11 cs.CV cs.AI 新提交 87%

MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation

MedPixel:用于医学推理与分割的统一像素-语言模型

Haoyu Yang, Meixing Shi, Zengjie Chen, Haoran Sun, Haitao Leng, Xiaoming Shi, Yuxiang Cai, Yankai Jiang

专题命中 领域大模型 :language model(title,abstract);LLM(abstract,abstract_cn);preference optimization(abstract);分类 cs.AI

AI总结 该研究提出统一医学像素-语言模型MedPixel,引入44万样本的MedPLG-440K数据集,通过联合多任务微调与像素级偏好优化训练,支持多类医学任务,性能优异且具备零样本迁移与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27944 2026-07-31 cs.IR cs.AI 新提交 87%

Interpretable Representation via LLM-Driven Generative Disentanglement for Local-Life Service Recommendation

面向本地生活服务推荐的、基于大语言模型驱动的生成式解缠的可解释表示

Long Zhang, Hao Jiang, Sheng Yu, Fei Pan, Peng Jiang, Kun Gai

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 针对现有语义ID生成框架的语义纠缠与黑箱问题,提出LGRID模型,通过生成式解缠范式提升本地生活服务推荐的性能与可解释性,在公开数据集上取得显著效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24793 2026-07-29 cs.IR cs.CL 新提交 87%

Retrieval, not hallucinations, will be the limiting factor for LLM-based clinical AI tools

检索而非幻觉将成为基于大语言模型的临床人工智能工具的限制因素

Kirk Roberts, Steven Bedrick, Kurt Miller, William R. Hersh, Hongfang Liu

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 探讨临床人工智能中基于大语言模型的错误,将讨论重点从精度错误转向召回错误,特别是患者级数据检索方面,概述错误类型、缓解策略及研究方向,提供检索评估概述。

Comments Perspective piece

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17653 2026-07-21 cs.CV cs.LG cs.MM 新提交 87%

LFM: Leveraging Foundation Models for Source-Free Universal Domain Adaptation

LFM:利用基础模型进行无源通用域适应

Jing Li, Pan Liu, Meng Zhao, Wanli Xue, Yanhong Yang, Xu Cheng, Fan Shi, Jianhua Zhang, Qinghua Hu, Shengyong Chen

机构 * School of Computer Science and Engineering, Tianjin University of Technology(天津理工大学计算机科学与工程学院) Engineering Research Center of Learning-Based Intelligent System, Ministry of Education of the People’s Republic of China, Tianjin University of Technology(中华人民共和国教育部基于学习的智能系统工程研究中心,天津理工大学) School of Artificial Intelligence, Tianjin University(天津大学人工智能学院) Engineering Research Center of City Intelligence and Digital Governance, Ministry of Education of the People’s Republic of China, Tianjin University(中华人民共和国教育部城市智能与数字治理工程研究中心,天津大学)

专题命中 领域大模型 :foundation model(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 研究在无源数据时将预训练源模型适应到目标域的问题,提出LFM框架,利用视觉语言模型计算相似度确定标签转移类型、识别未知样本,通过共识策略精炼伪标签训练目标模型,实验验证了该框架的有效性和优越性。

Comments Accepted by IEEE Transactions on Multimedia (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14659 2026-07-17 cs.SE cs.AI 新提交 87%

LLM-Driven Approach to Modeling Tool Interoperability in Automotive Domain

汽车领域中基于大语言模型驱动的建模工具互操作性方法

Nenad Petrovic, Jiajie Zhang, Vahid Zolfaghari, Alois Knoll

机构 * European Chips Joint Undertaking(欧洲芯片联合体) Federal Ministry of Research, Technology and Space of Germany(德国联邦研究、科技与航天部)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究汽车领域异构建模工具互操作性难题,提出基于大语言模型驱动的方法,涉及模型实例到目标元模型的映射及元模型合并,经案例验证该方法可行,能减少手动转换工作量并生成有效目标模型促进跨工具互操作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10871 2026-07-14 cs.AI cs.HC 新提交 87%

Toward Contemplative LLM: A Modular Framework for Evaluating and Enhancing LLM Alignment in Mental Health

迈向沉思型大语言模型:心理健康领域中评估与增强大语言模型对齐性的模块化框架

Asher Sprigler, Yang-Yang Feng, Iftach Amir, Jonathan E. Bogard, Todd S Braver, Yi Ding, David Kinney, Yixue Zhao

机构 * Purdue University(普渡大学) Washington University in St. Louis(圣路易斯华盛顿大学) Yixue Research Institute(易学研究所)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 针对新模型等涌现使评估沉思原则增强大语言模型对齐性具挑战的问题,提出模块化可扩展评估框架,能集成新元素并重现先进结果,支持交叉评估,其提示模块便于纳入伦理视角,为跨学科研究及有益人机生态系统奠定基础。

Comments Accepted as an oral presentation at HARMONY 2026 (Human-centered AI Research for Mental Health, an Open Networking Symposium), co-located with IEEE/ACM Conference on Connected Health: Applications, Systems, and Engineering Technologies (CHASE 2026) held in Pittsburgh, August 6, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18181 2026-07-08 cs.IR cs.AI cs.CY 新提交 87%

IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information Extraction

IUU+DB:通过LLM驱动的信息提取追踪非法、不报告和不管制捕捞、海鲜欺诈和劳工虐待

Henry Bodwell, Hong Yang, John C. Simeone, Kelvin Gorospe, Bella Sullivan, Lana Huang, Jessica Gephart, Sandy Aylesworth, Molly Masterton, Naren Ramakrishnan

机构 * University Of Washington(华盛顿大学)

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出IUU+概念扩展非法捕捞定义,并构建基于大语言模型的IUU+DB系统,从异构文档中自动提取事件关键信息,支持去重和趋势分析,为渔业监管和研究提供数据支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04579 2026-07-07 cs.SE cs.AI 新提交 87%

LLM-Driven CI-CD Workflow Intelligence for Cyber Systems Engineering

面向网络系统工程的由大语言模型驱动的持续集成与持续交付工作流智能

Bonan Shen, Jiazhou Gao, Tao Ning, Wei-Jung Huang, Xin Liu

机构 * Independent Researcher(独立研究者)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 研究面向网络系统工程的CI/CD工作流智能,通过基于大语言模型的分析管道,结合多种方法处理大量GitHub库数据,发现问题并得出结果,强调CI/CD可观测性应结合诊断、上下文和人工审查。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00003 2026-07-02 cs.IR cs.AI 新提交 87%

From "Strings" to "Things" for Personal Knowledge Graphs: Evaluating LLM Triple Extraction for Recommendation Systems

从“字符串”到“事物”:面向个人知识图谱的LLM三元组提取在推荐系统中的应用评估

Abhirup Dasgupta, Fernando Spadea, Oshani Seneviratne

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出一种可复现的流水线,利用轻量级大语言模型从对话数据中提取结构化用户偏好三元组,构建个人知识图谱,并评估其在推荐任务中的效用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24331 2026-06-24 cs.CL cs.ET 新提交 87%

Transformer-Based Language Models Across Domain Verticals: Architectures, Applications and Critical Assessment

基于Transformer的语言模型在垂直领域中的应用:架构、应用与批判性评估

Guruprakash J, Krithika L. B

机构 * SCOPE, VIT-AP University(VIT-AP大学SCOPE学院) SCORE, VIT(VIT大学SCORE学院)

专题命中 领域大模型 :language model(title,abstract);RLHF(summary_cn);instruction tuning(abstract);分类 cs.CL

AI总结 本文系统梳理Transformer架构变体(编码器、解码器、长上下文等)及后2023年进展(指令微调、RLHF、MoE等),评估其在医疗、金融等领域的部署,并批判性分析模型选择、参数-能耗权衡及“SOTA”含义。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22586 2026-06-23 cs.AI cs.SE 新提交 87%

Text2DSL: LLM-Based Code Generation for Domain-Specific Languages

Text2DSL:基于大语言模型的领域特定语言代码生成

Alexander V. Kozachok, Alexander M. Nazimov, Shamil G. Magomedov

机构 * RTU MIREA --- Russian Technological University(俄罗斯技术大学) Academy of the Federal Guard Service of the Russian Federation(俄罗斯联邦警卫局学院)

专题命中 领域大模型 :LLM(title,summary_cn);分类 cs.AI

AI总结 提出Text2DSL任务,将自然语言描述自动转换为领域特定语言(DSL)代码;构建PolkitBench数据集,通过结构化上下文(BNF语法、API规范等)显著提升LLM生成代码的语法和结构有效性。

Comments 14 pages, 4 figures, 5 tables. Accepted at KES 2026 (Knowledge-Based Intelligent Information and Engineering Systems), Procedia Computer Science, Elsevier

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16905 2026-06-16 cs.CL 新提交 87%

Speaking the Language of Science: Toward a General-Purpose Generative Foundation Model for the Natural Sciences

说科学的语言:面向自然科学的通用生成基础模型

Mingyang Li, Yurou Liu, Jieping Ye, Bing Su, Ji-Rong Wen, Zheng Wang

机构 * Alibaba Group(阿里巴巴集团) Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)

专题命中 领域大模型 :foundation model(title,abstract);LLM(abstract_cn);large language model(abstract);language model(abstract)

AI总结 提出LOGOS模型,通过统一科学语法将异构任务转化为自回归框架中的下一个词预测,在多个科学任务上匹配或超越领域专用基线,验证了“一模型适用于所有”的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12702 2026-06-12 cs.AI 新提交 87%

Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System

以部署为中心的评估:预测临床大语言模型系统中的查询级拒绝风险

Alyssa Unell, Miguel Fuentes, Brenna Li, Bridget Lin, Meena Jagadeesan, Sanmi Koyejo, Nigam Shah

机构 * Department of Computer Science(计算机科学系) Department of Medicine(医学系) Department of Biomedical Data Science(生物医学数据科学系)

专题命中 领域大模型 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 针对临床大语言模型系统,提出基于部署上下文(如提供者类型、科室名称)的预响应分类器,预测用户拒绝风险,AUROC达0.719,并展示其在触发护栏和弃权中的效用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13786 2026-08-17 cs.IR cs.AI cs.CL 新提交 87%

Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

AI聊天机器人能否找到专家会选择的研究?模型、用户角色和样本量对医学问题研究检索的影响

Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt, Zhiyong Lu

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 该研究评估Claude Sonnet 5、Gemini 3.1 Pro、ChatGPT GPT-5.5三款LLM聊天机器人在模拟三种用户角色下的医学研究检索表现,发现其召回率因模型、角色差异显著,且存在偏向大样本临床试验的偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏