arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-08-20 至 2026-08-20 共收录 20 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 20 篇

2608.18279 2026-08-20 physics.optics cs.LG 新提交 91%

A Comprehensive Review of Large Language Models for Nanophotonics: From Surrogate Modeling to Autonomous Design

面向纳米光子学的大语言模型综合综述:从代理建模到自主设计

Huanshu Zhang, Kegeng Tang, Lei Kang, Sawyer D. Campbell, Zihao Wang, Douglas H. Werner

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);foundation model(abstract)

AI总结 该综述探讨大语言模型(LLMs)如何通过语义接口、代码生成及工具编排改进纳米光子学工作流,梳理相关方法的两类模式及跨学科应用,展望具备物理感知的多模态基础模型,推动AI从被动工具向主动科研合作者转变。

Comments Accepted for publication in Advanced Photonics

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19025 2026-08-20 cs.AI cs.DB 新提交 91%

Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models

自提示与跨模型共识:利用大语言模型从科学文献中实现可复现的数据提取

Valentin Romanov, Monique Bax, Steven Niederer

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);prompting(title);分类 cs.AI

AI总结 本研究提出结合自提示与跨模型共识的方法,通过四个工作流程优化LLMs从科学文献提取数据的性能,为科学数据规模化整理提供可行方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18080 2026-08-20 cs.AI 新提交 90%

Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges

心理健康领域的大语言模型:应用、创新与伦理挑战的系统综述

Yisong Chen, Yifan Gao, Sijing Yu, Chuqing Zhao, Yang Lu

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 该系统综述梳理了大语言模型在心理健康领域的多类应用、技术创新,同时探讨了其面临的伦理与监管挑战,并倡导建立保障其安全公平部署的框架。

Comments Systematic review. Published in Journal of Industrial Integration and Management (2025). Applications of large language models in mental health, including social media analysis, clinical conversational agents, therapy support tools, multimodal learning, and ethical considerations

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22664 2026-08-20 cs.AI 版本更新 90%

MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance

WorkstreamBench: 评估LLM代理在金融领域的端到端电子表格任务

Thomson Yen, Julian Poeltl, Harshith Srinivas Gear, Yilin Meng, Joshua Fan, Adam Shen, Yili Liu, Ali Bauyrzhan, Patrick Shea, Siri Du, Haoyang Liu, Daniel Guetta, Hongseok Namkoong

机构 * Decision, Risk, and Operations Division, Columbia Business School(哥伦比亚商学院决策、风险与运营部门) ESB Business School, Reutlingen University(图宾根大学ESB商学院)

专题命中 领域大模型 :LLM(title,title_cn);分类 cs.AI

AI总结 本文提出WorkstreamBench,用于评估LLM代理在金融领域复杂端到端电子表格任务中的能力,重点在于财务建模和情景分析等关键流程,通过三个维度(准确性、公式、格式)的细粒度标准来衡量解决方案质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18260 2026-08-20 cs.AI cs.CL cs.CR cs.IR cs.LG 新提交 88%

Redakto - The Incognito Tab for LLMs

Redakto——面向大语言模型的隐身标签

Saurav Kumar Saha, Tom Röhr, Felix Bießmann

机构 * Berlin University of Applied Sciences(柏林应用科学大学) Einstein Center Digital Future(爱因斯坦数字未来中心)

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 Redakto是一款开源匿名化工具,可通过网页应用、REST API等供用户使用,能移除文本中PII,其匿名化后的文本效用与原文相当,可助力LLM使用时的隐私保护。

Comments Accepted at WIPE-OUT 2026, 2nd Workshop on Machine Unlearning and Privacy Preservation at ECML-PKDD

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16921 2026-08-20 cs.IR cs.CR cs.LG cs.MA 版本更新 86%

MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering Model

MITRE-SAGE:一种多智能体网络安全问答模型

Ali Habibzadeh, Farid Feyzi, Reza Ebrahimi Atani

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 针对网络安全领域LLM的不足,研究提出多智能体框架MITRE-SAGE,结合MITRE-QA基准验证其在多项网络安全问答任务中优于基线方法,轻量级配置表现突出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18116 2026-08-20 cs.CL cs.LG 新提交 81%

You Are What You Prompt: Prompt Quality, Domain Shift, and Uncertainty in Agrifood Vision-Language Models

你即你所提示的:农业食品视觉语言模型中的提示质量、领域偏移与不确定性

Andrea Morales-Garzón, Salvador López-Joya, Miguel López-Pérez, Maria J. Martin-Bautista

机构 * University of Granada(格拉纳达大学)

专题命中 领域大模型 :language model(title,abstract);分类 cs.CL、cs.LG

AI总结 该研究针对农业食品领域,评估了零样本提示集成(ZPE)在分布内与分布外场景的表现,提出PID方法提升严重领域偏移下的故障检测能力,验证了领域特定提示池的优势。

Comments Accepted in the journal Procesamiento del Lenguaje Natural (SEPLN2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18086 2026-08-20 cs.AI cs.LG 新提交 81%

Position: Current Model Cards Are Insufficient for Downstream Governance of Open-Weight Foundation Models

立场:当前模型卡片不足以支持开放权重基础模型的下游治理

Sungwon Chae, Keonwoo Kim, Hoki Kim, Jaeyeon Ju, Sangchul Park

专题命中 领域大模型 :foundation model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文分析Hugging Face的500份模型卡片,指出现有模型卡片无法支持开放权重基础模型下游治理,提出需整合模型卡片、可接受使用政策、许可证的多层治理框架。

Comments Accepted as a position paper at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14870 2026-08-20 hep-ph cs.LG 版本更新 79%

Pre-Training for Simulation-Based Science: A Study on Jet Foundation Model Training Objectives

基于模拟的科学预训练:喷注基础模型训练目标研究

Ibrahim Elsharkawy, Joschka Birk, Vinicius Mikuni, Wahid Bhimji, Gregor Kasieczka, Benjamin Nachman

机构 * Department of Physics, University of Toronto and Vector Institute(物理系,多伦多大学和向量研究所) NERSC, Lawrence Berkeley National Laboratory(NERSC,伯克利国家实验室) Institut für Experimentalphysik, Universität Hamburg(实验物理研究所,汉堡大学) Nagoya University, Kobayashi-Maskawa Institute(名古屋大学,小林昭夫研究所) Department of Particle Physics and Astrophysics, Stanford University(粒子物理与天体物理系,斯坦福大学) Fundamental Physics Directorate, SLAC National Accelerator Laboratory(基础物理局,SLAC国家加速器实验室)

专题命中 领域大模型 :foundation model(title,abstract);分类 cs.LG

AI总结 本文系统比较了高能物理中基础模型的预训练方法,发现纯分类预训练在标签充足时最优,结合自监督掩码粒子建模在低标签场景下表现突出,而流匹配生成预训练对下游分类无益,但必须包含在预训练目标中才能提升生成任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19078 2026-08-20 cs.CV 新提交 78%

Subgroup performance analysis of adaptation strategies for chest X-ray foundation models

胸部X射线基础模型适配策略的子组性能分析

Dhruv Gupta, Emma A.M. Stanley, Fabio De Sousa Ribeiro, Sujal Desai, Ben Glocker

机构 * Imperial College London(帝国理工学院) Royal Brompton Hospital(皇家布朗普顿医院) Causality in Healthcare AI Hub(医疗保健AI因果关系中心) National Heart & Lung Institute(国家心肺研究所)

专题命中 领域大模型 :foundation model(title,abstract)

AI总结 该研究针对胸部X射线基础模型,探究三种参数高效适配技术对病理分类性能与子组公平性的影响,发现整体性能提升未必减少子组差异,公平性影响需直接按任务评估

Comments Accepted at MICCAI Workshop on Fairness of AI in Medical Imaging (FAIMI) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12036 2026-08-20 cs.AI cs.CL cs.HC cs.LG cs.MA 版本更新 75%

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

Mechanist:作为科学仪器的AI,用于发现智能的机制

Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen

机构 * Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学) Heriot-Watt University(赫瑞瓦特大学) Southern University of Science and Technology(南方科技大学) University of California, San Diego(加利福尼亚大学圣迭戈分校) Northeastern University(东北大学)

专题命中 领域大模型 :foundation model(abstract);pretraining(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究推出智能体系统Mechanist,以AI为科学仪器自主发现AI智能机制,其生成假设更有价值、实验更可靠,还能发现安全风险、构建信念机制理论并转化为提升模型性能的干预措施。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18937 2026-08-20 cs.CL cs.AI 新提交 73%

MedUAG: Unified Understanding and Generation for Medical Multimodal Models

MedUAG:面向医学多模态模型的统一理解与生成框架

Zijie Meng, Yuncheng Zhang, Hualiang Wang, Yitian Tang, Xiaotang Gai, Chen Shen, Songtao Jiang, Shaosheng Cao, Jian Wu, Xian Wu, Zuozhu Liu

机构 * Zhejiang University(浙江大学) Hong Kong University of Science and Technology(香港科技大学) Tsinghua University(清华大学) Tencent Jarvis Lab(腾讯Jarvis实验室)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 针对医学多模态领域缺乏统一理解生成框架的问题,本文构建了MedUAGCorpus数据集与MedUAGBench基准,开发了MedUAG模型,其在医学理解与生成任务中表现出色,为下一代医学多模态系统奠定基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18565 2026-08-20 cs.SE 新提交 67%

SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation

SemaPLC:一种基于项目、经验证门控的PLC代码生成智能体框架

Yanlun Tu, Huacan Wang, Ziyue Zhou, Jie Zhou, Ningyan Zhu, Ge Chen, Wangyi Chen, Tengfei Zhou, Yifan Zhou, Dasheng Yang, Xiaofeng Mou, Hui Zhang, Yi Xu

专题命中 领域大模型 :large language model(abstract);language model(abstract)

AI总结 SemaPLC是基于项目、经验证门控的PLC代码生成智能体框架,通过外部检查确认任务完成,在多模型的POU任务及项目上下文任务中均表现最优,动态行为测试得分显著高于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24834 2026-08-20 q-bio.NC cs.AI cs.ET cs.LG 版本更新 62%

Cross-Cohort Spectral-Temporal Dissociation in Frozen EEG Foundation-Model Representations

脑电图基础模型对长程时间相关性视而不见:其跨群体脆弱性背后的频谱-时间解离

Marzieh Zare

专题命中 领域大模型 :foundation model(abstract);分类 cs.AI、cs.LG

AI总结 研究探讨脑电图基础模型对长程时间相关性的表征及跨群体转移情况。通过探测五种模型,发现它们均未按时间顺序表示LRTC,频谱输入模型能恢复1/f但非DFA,跨群体转移方面各模型表现不一,揭示了模型在这方面的局限性。

Comments I found a computational error in one of the tables that needs to be fixed; and it may take 2-3 weeks

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18099 2026-08-20 cs.AI q-fin.PM 新提交 57%

FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management

FinSkillBench:评估用于投资管理的智能体AI与领域技能

Jermyn Zhen Yong Bek, Zhuang Qiang Bok, Zhongtian Sun

机构 * Deep Insight Labs(深洞察实验室) University of Kent(肯特大学) University of Cambridge(剑桥大学)

专题命中 领域大模型 :language model(abstract);分类 cs.AI

AI总结 FinSkillBench是评估投资管理领域AI智能体金融技能的基准,经9个模型的大规模评估发现,精选技能可显著提升性能,而自生成技能益处有限,其成果为该领域研究提供了支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07985 2026-08-20 cs.CL cs.IR 版本更新 57%

Predicting the Benefit of Retrieval Augmentation in Open-Domain Question Answering

基于问答的RAG性能预测

Or Dado, David Carmel, Oren Kurland

机构 * Technion(技术学院) Technology Innovation Institute (TII)(技术创新研究所)

专题命中 领域大模型 :LLM(abstract);分类 cs.CL

AI总结 本文研究了RAG在问答中的性能预测,提出了一种新的监督预测器,通过建模问题、检索段落和生成答案之间的语义关系,提升了预测效果。

Comments 17 pages. 4 figures. 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14473 2026-08-20 cs.CL 版本更新 57%

AI Can Learn Scientific Taste

AI 可以学习科学品味

Jingqi Tong, Mingzhe Li, Hangcheng Li, Yongzhuo Yang, Yurong Mou, Weijie Ma, Hongji Chen, Xiaoran Liu, Qinyuan Cheng, Ming Zhang, Qiguang Chen, Weifeng Ge, Qipeng Guo, Tianlei Ying, Tianxiang Sun, Yining Zheng, Zhiheng Xi, Xinchi Chen, Jun Zhao, Ning Ding, Xuanjing Huang, Yu-Gang Jiang, Xipeng Qiu

专题命中 领域大模型 :LLM(abstract);分类 cs.CL

AI总结 本文提出RLCF框架,通过社区反馈学习科学品味,使AI能提出高潜力研究想法,实验显示其优于现有模型并具备泛化能力。

Comments 47 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18166 2026-08-20 eess.IV cs.CV 新提交 50%

TractoGraphVLM: A Unified Vision-Language Framework for White Matter Tractography

TractoGraphVLM:用于白质纤维束成像的统一视觉-语言框架

Gurucharan Marthi Krishna Kumar, Janine Dale Mendola, Amir Shmuel

专题命中 领域大模型 :language model(abstract)

AI总结 TractoGraphVLM是统一视觉-语言框架,可完成白质纤维束的分类、检索、描述、问答四项任务,在HCP数据集上表现良好,具跨年龄迁移鲁棒性,仅从语言学习神经解剖学知识。

Comments Accepted as a Spotlight at the ECCV 2026 Workshop on Artificial Intelligence for Medical 3D Vision (AI4M3D). Our codebase, including all training and evaluation pipelines, is publicly available at this https URL (https://github.com/AS-Lab/Marthi-et-al-2026-TractoGraphVLM-Unified-Vision-Language-White-Matter-Tractography)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18637 2026-08-20 cs.IR 新提交 50%

PILOT Technical Report

PILOT技术报告

Jiuning Lin, Ruiquan Lan, Xiaodong Zhu, Bin Zhang, Chengyu Lai, Chuxin Chen, Dimin Wang, Hongtao Cheng, Jialin Zhu, Lingqing Zhang, Shuai Zhong, Tao Wang, Weipeng Huang, Yinjiang Cai, Yinnan Song, Yuan Liu, Zhibo Xiao, Zhixin Ma, Zihong Huang

专题命中 领域大模型 :LLM(abstract)

AI总结 该研究针对现有推荐系统优化智能体方法的反应式缺陷,提出PILOT框架,通过三类角色构建控制循环,在淘宝平台对比ROAM验证,实现多项指标提升且搜索效率显著提高,无需人工干预。

Comments Technical Report, 42 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18479 2026-08-20 cs.CV 新提交 50%

COSTA: A Cluster-Centric Paradigm for Annotation-Free Open-Set Semantic Segmentation of Aerial Point Clouds with Domain Shifts

COSTA:面向存在域偏移的航拍点云无标注开放集语义分割的以聚类为中心范式

Yanghong Lin, Li Fang, Tianyu Li, Shudong Zhou, Wei Yao

机构 * Fujian Institute of Research on the Structure of Matter, Chinese Academy of Sciences(中国科学院福建物质结构研究所) Institute of Urban Environment, Chinese Academy of Sciences(中国科学院城市环境研究所) University of Chinese Academy of Sciences(中国科学院大学) School of Resource and Environmental Sciences, Wuhan University(武汉大学资源与环境科学学院)

专题命中 领域大模型 :language model(abstract)

AI总结 COSTA提出以聚类为中心的范式,无需额外训练即可在推理阶段将预训练航拍点云分割模型适配到偏移目标域,在三个基准上实现按需分割,mIoU达70.09%。

详情

展开后加载摘要…

URL PDF HTML 收藏