arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-05-28 至 2026-05-28 共收录 32 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 32 篇

2605.28187 2026-05-28 cs.IR cs.AI cs.CY cs.SI 93%

Whose Name Comes Up? III: Persona Prompting Effects in LLM-Based Scholar Recommendation

谁的名字会出现?III:基于LLM的学者推荐中的人设提示效应

Annabella Sánchez-Guzmán, Lukas Eberhard, Denis Helic, Lisette Espín-Noboa

机构 * Graz University of Technology(格拉茨技术大学) Complexity Science Hub(复杂科学中心)

专题命中 领域大模型 :LLM(title,title_cn);prompting(title);large language model(abstract);language model(abstract)

AI总结 本研究通过构建基准测试,分离模型选择与提示设计对LLM学者推荐的影响,发现提示设计(语言、地点、角色与任务)显著影响推荐质量(事实性、覆盖度)和社会代表性(多样性、均等性)。

Comments 25 pages (10 main, 2 references, 13 appendix), 6 figures in main, 13 figures in appendix (under-review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28201 2026-05-28 cs.AI 92%

Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents

种植、持久化、触发:针对大语言模型智能体的潜伏攻击

Yongxiang Li, Moxin Li, Zhixin Ma, Fengbin Zhu, Dongrui Liu, Wenjie Wang, Fuli Feng

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Singapore Management University(新加坡管理学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 提出潜伏攻击(Sleeper Attack),即攻击者将对抗性内容注入智能体状态并持久化,在后续交互中被良性用户查询触发,导致有害行为;构建包含1896个实例的基准测试,实验表明当前最强LLM智能体仍易受此类攻击。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27856 2026-05-28 cs.IR cs.AI 92%

Fine-Tuned LLM as a Complementary Predictor Improving Ads System

微调LLM作为改进广告系统的互补预测器

Hui Yang, Daiwei He, Kevin Jiang, Taejin Park, Kungang Li, Jiajun Luo, Yuying Chen, Xinyi Zhang, Sihan Wang, Haoyu He, Yu Liu, Lakshmi Manoharan, David Xue, Shubham Barhate, Runze Su, Duna Zhan, Ling Leng, Siping Ji, Jinfeng Zhuang, Alice Wu, Leo Lu, Han Sun, Zhifang Liu

机构 * Pinterest, Inc., USA(Pinterest公司,美国)

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出将微调的开源LLM作为广告特定辅助预测器,从用户画像和历史中预测广告主,增强候选生成并为下游排序提供先验信息,在工业广告系统中取得离线改进和在线业务提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27700 2026-05-28 cs.DL cs.AI 92%

CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text

CiteCheck: 基于检索的科学文本中LLM引用幻觉检测

Khashayar Khajavi, Shaghayegh Sadeghi, Rise Adhikari, Alexander Tessier

机构 * School of Computing Science, Simon Fraser University(西蒙·弗雷泽大学计算科学学院)

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出CiteCheck框架,通过从外部学术来源检索候选出版物、使用结构化LLM验证器比较引用与候选信息,并将验证器得分映射为精确、次要和主要三个标签,以检测LLM生成的引用幻觉,在物理基准上达到88.7 macro-F1和88.9%准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27403 2026-05-28 cs.CY cs.AI 92%

LLM-assisted sentiment analysis for integrated computational and qualitative mixed methods education research: A case study of students' written reflection assignments

LLM辅助情感分析在综合计算与定性混合方法教育研究中的应用:学生书面反思作业案例研究

Xiomara Gonzalez, Gabriella Coloyan Fleming, Andrew Katz, Maya Denton, Jessica Deters

机构 * Chandra Family Department of Electrical and Computer Engineering, University of Texas at Austin(德克萨斯大学奥斯汀分校电子与计算机工程系) Department of Engineering Education, Virginia Polytechnic Institute and State University(弗吉尼亚理工大学工程教育系) Gallogly College of Engineering, University of Oklahoma(俄克拉荷马大学加洛格利工程学院) Department of Mechanical and Materials Engineering, University of Nebraska-Lincoln(内布拉斯加大学林肯分校机械与材料工程系)

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究通过纵向案例,利用LLM辅助情感分析结合统计检验与主题分析,探讨学生身份变量对留学期间语言交流情感的影响,发现海外生活经历是唯一显著变量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01627 2026-05-28 cs.CL cs.AI 91%

JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models

JMedEthicBench:用于评估日语大语言模型医疗安全性的多轮对话基准

Junyu Liu, Zirui Li, Qian Niu, Zequn Zhang, Yue Xun, Wenlong Hou, Shujun Wang, Yusuke Iwasawa, Yutaka Matsuo, Kan Hatakeyama-Sato

机构 * Kyoto University(京都大学) Hohai University(河海大学) The University of Tokyo(东京大学) University of Science and Technology of China(中国科学技术大学) Hong Kong Polytechnic University(香港理工大学)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 提出首个多轮对话基准JMedEthicBench,基于日本医学会67条指南和7种自动越狱策略生成5万+对抗对话,评估27个模型发现医疗专用模型安全性脆弱,且多轮交互中安全性显著下降。

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28524 2026-05-28 cs.AI 90%

Let Relations Speak: An End-to-End LLM-GNN Soft Prompt Framework for Fraud Detection

让关系说话:面向欺诈检测的端到端LLM-GNN软提示框架

Zhixing Zuo, Huilin He, Jiasheng Wu, Dawei Cheng

机构 * School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出LGSPF框架,通过软提示桥接图结构与语义空间,并引入并行GNN编码器将多关系拓扑转化为图令牌,实现端到端优化,在欺诈检测中达到最优性能。

Comments 14 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27853 2026-05-28 cs.AI 90%

MolLingo: Molecule-Native Representations for LLM-Powered Scientific Agents

MolLingo:面向LLM驱动的科学智能体的分子原生表示

Thao Nguyen, Heng Ji

机构 * Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校Siebel计算与数据科学学院)

专题命中 领域大模型 :LLM(title,title_cn);分类 cs.AI

AI总结 提出MolLingo多智能体系统,通过共享内存协调文献、化学家和编排智能体,结合基于BRICS的片段枚举(BFE)表示方法,实现分子块级推理与编辑,在四个基准上优于前沿LLM和专用基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09781 2026-05-28 cs.CR 89%

CLIOPATRA: Extracting Private Information from LLM Insights

CLIOPATRA: 从LLM洞察中提取隐私信息

Meenatchi Sundaram Muthu Selva Annamalai, Emiliano De Cristofaro, Peter Kairouz

专题命中 领域大模型 :LLM(title,title_cn)

AI总结 提出CLIOPATRA攻击,通过注入恶意聊天突破多层启发式隐私保护,从LLM洞察系统中提取目标用户敏感信息,在合成医疗聊天中高达65%成功率且几乎100%精确。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27970 2026-05-28 cs.AI 89%

Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations

人类感知域的几何结构在LLM表征中短暂出现

Simardeep Singh, Paras Chopra

机构 * Indian Institute of Technology Roorkee(印度理工学院罗尔基分校)

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 研究大型语言模型内部表征中是否出现与人类感知组织相似的几何结构,发现多个感知域的几何结构在中间层短暂涌现,且与人类基准对齐。

Comments 19 Pages, 28 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27710 2026-05-28 cs.AI 89%

DeepSciVerify: Verifying Scientific Claim--Citation Alignment via LLM-Driven Evidence Escalation

DeepSciVerify: 通过LLM驱动的证据升级验证科学声明与引文对齐

Shaghayegh Sadeghi, Khashayar Khajavi, Rise Adhikari, Alexander Tessier

机构 * School of Computing Science, Simon Fraser University(西蒙弗雷泽大学计算科学学院)

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出DeepSciVerify两阶段流水线,结合摘要推理与选择性升级到段落证据,在SCitance基准上以86.7 Micro-F1超越纯摘要基线4.5点,同时67%实例无需全文检索。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28338 2026-05-28 cs.AI 88%

SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models

SafeMed-R1: 临床医生审计的安全与伦理对齐用于医疗大语言模型

Chao Ding, Mouxiao Bian, Tianbin Li, Minjia Yuan, Yidong Jiang, Yankai Jiang, Jinru Ding, Jiayuan Chen, Zhuangzhi Gao, Pengcheng Chen, Zhao He, Rongzhao Zhang, Meiling Liu, Luyi Jiang, Jie Xu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Joint Laboratory of Biomedical Artificial Intelligence(生物医学人工智能联合实验室) Shanghai Institute of Infectious Disease and Biosecurity(上海传染病与生物安全研究院) Shanghai Health Development Research Center (Shanghai Medical Information Center)(上海健康发展战略研究中心(上海医疗信息中心)) University of Washington(华盛顿大学) Department of Eye and Vision Sciences, University of Liverpool(利物浦大学眼科与视觉科学系) Liverpool Centre for Cardiovascular Science, University of Liverpool(利物浦大学心血管科学中心) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 提出SafeMed-R1模型,通过可追溯的临床信任信号管道和红队压力测试实现安全与伦理对齐,在临床基准上达到79.6%的宏平均准确率,并将不安全输出减少约3-5%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28328 2026-05-28 cs.LG cs.AI 88%

Learning the Error Patterns of Language Models

学习语言模型的错误模式

Jinwoo Kim, Taylor Berg-KirkPatrick, Loris D'Antoni

机构 * Department of Computer Science and Engineering(计算机科学与工程系) University of California-San Diego(加州大学圣地亚哥分校)

专题命中 领域大模型 :LLM(summary_cn,abstract);language model(title);分类 cs.AI、cs.LG

AI总结 提出前缀过滤器(prefix filters)来捕捉LLM在特定领域中的错误模式,并通过Palla算法高效学习这些过滤器,从而提升输出有效性,例如在TypeScript生成中将编译率提升60%以上。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27958 2026-05-28 cs.CL cs.AI cs.LG 83%

Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations

压力测试LLM中的欺骗探针:扩展性、鲁棒性与欺骗表示的几何结构

Sachin Kumar

机构 * LexisNexis(LexisNexis公司)

专题命中 领域大模型 :LLM(title_cn,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文通过系统压力测试,诊断线性探针在分布偏移下失效的原因,发现风格增强可恢复近完美检测,并证明欺骗编码非单一线性方向或熵代理,而是分布式亚阈值特征。

Comments Accepted at the GEM Workshop @ ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27610 2026-05-28 cs.IR cs.AI cs.HC 81%

Eliot: Interactively $\underline{E}$xploring Fast-Changing Scientific $\underline{Li}$terature Trends with $\underline{O}$nline Da$\underline{t}$a and Learning

Eliot: 通过在线数据和学习交互式探索快速变化的科学文献趋势

Bernardo A. Denkvitts, Nitin Gupta, Biplav Srivastava

机构 * University of South Carolina(南卡罗来纳大学)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出Eliot系统,通过查询时聚类和时间可视化,帮助研究人员可追溯地探索快速变化的科学文献趋势。

Comments Under-review at CIKM Applied Research 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06054 2026-05-28 cs.CL 81%

Are We Truly Innovating? A Qualitative and Quantitative Study of Originality in AI Research Papers

我们真的在创新吗?AI研究论文原创性的定性与定量研究

Abeer Mostafa, Thi Huyen Nguyen, Zahra Ahmadi

机构 * Peter L. Reichertz Institute for Medical Informatics(汉诺威医学院彼得·L·里赫茨医学信息学研究所) L3S Research Center(L3S研究中心) Lower Saxony Center for Artificial Intelligence and Causal Methods in Medicine (CAIMed)(下萨克森人工智能与医学因果方法中心(CAIMed))

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 基于10万+同行评审报告,通过定性与定量方法分析AI研究论文原创性的感知维度,并评估大语言模型在原创性评估中的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28563 2026-05-28 cs.LG cs.AI 81%

A Multi-dimensional Framework for Evaluating Generalization in EEG Foundation Models

评估脑电图基础模型泛化能力的多维框架

Aditya Kommineni, Emily Zhou, Kleanthis Avramidis, Tiantian Feng, Shrikanth Narayanan

机构 * Signal Analysis and Interpretation Laboratory(信号分析与解释实验室)

专题命中 领域大模型 :foundation model(title,abstract);分类 cs.AI、cs.LG

AI总结 提出一个多维评估框架,在低资源条件下系统评估EEG基础模型(如LaBraM、CSBrain、CBraMod)的泛化能力,发现其在长上下文任务中表现优异,但在短窗口BCI任务中与监督模型相当,且对通道限制鲁棒性不足。

Comments 24 pages, 5 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12887 2026-05-28 cs.CV 78%

Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification

重新审视用于可扩展3D医学图像分类的2D基础模型

Han Liu, Bogdan Georgescu, Yanbo Zhang, Youngjin Yoo, Michael Baumgartner, Riqiang Gao, Jianing Wang, Gengyan Zhao, Eli Gibson, Dorin Comaniciu, Sasa Grbic

机构 * Digital Technology and Innovation, Siemens Healthineers, Princeton NJ, USA(西门子医疗数字技术与创新,普林斯顿新泽西州,美国) Digital Technology and Innovation, Siemens Healthineers, Erlangen, Germany(西门子医疗数字技术与创新,埃尔兰根,德国)

专题命中 领域大模型 :foundation model(title,abstract)

AI总结 本文针对当前3D医学图像分类基础模型的数据偏差、适应不足和任务覆盖不全问题,提出AnyMC3D框架,通过冻结2D基础模型并添加轻量插件实现高效多任务扩展,并在12项任务上达到领先性能。

Comments 1st Place in VLM3D Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27571 2026-05-28 cs.AI cs.CL cs.DB 73%

Discovery Agents for Real-Time Analytics: Toward Proactive Insight Systems

实时分析发现代理:迈向主动洞察系统

Gaetano Rossiello, Dharmashankar Subramanian

机构 * IBM

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出一种多智能体架构,通过持续发现循环(假设生成、编译、验证、可视化)实现实时数据流的自主洞察发现,支持从查询驱动向主动发现的范式转变。

Comments Accepted at Supporting Our AI Overlords (SAO) at the ACM Conference on AI and Agentic Systems (CAIS), May 26 2026, San Jose, CS, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10185 2026-05-28 cs.CL cs.AI cs.MA 73%

Auditing medical multi-agent AI reveals risks of false consensus

审计医疗多智能体AI揭示虚假共识风险

Yinghao Zhu, Lei Gu, Zixiang Wang, Haoran Sang, Dehao Sui, Wen Tang, Lan Mi, Yasha Wang, Junyi Gao, Liang Yao, Tianfan Fu, Ewen Harrison, Lequan Yu, Liantao Ma

机构 * National Engineering Research Center for Software Engineering, Peking University(北京大学软件工程国家工程研究中心) School of Computing and Data Science, The University of Hong Kong(香港大学计算机与数据科学学院) Department of Nephrology, Peking University Third Hospital(北京大学第三医院肾内科) Key Laboratory of Carcinogenesis and Translational Research (Ministry of Education), Department of Lymphoma, Peking University Cancer Hospital & Institute(教育部癌症发生与转化研究重点实验室、北京大学肿瘤医院淋巴瘤科) Department of Automation, Tsinghua University(清华大学自动化系) Centre for Medical Informatics, The University of Edinburgh(爱丁堡大学医学信息学中心) Health Data Research UK(英国健康数据研究机构) Lee Kong Chian School of Medicine, Nanyang Technological University(南洋理工大学李科贤医学院) State Key Laboratory for Novel Software Technology, School of Computer Science, Nanjing University(南京大学新型软件技术国家重点实验室、计算机科学学院)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出MedAgentAudit框架,通过专家验证的审计流程诊断医疗多智能体系统中的协作失败模式,发现虚假共识、权威偏差等系统性风险。

Comments Code and Data: https://github.com/MedX-PKU/MedAgentAudit

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27873 2026-05-28 cs.AI 70%

AIBuildAI-2: A Knowledge-Enhanced Agent for Automatically Building AI Models

AIBuildAI-2:一种用于自动构建AI模型的知识增强智能体

Ruiyi Zhang, Peijia Qin, Qi Cao, Li Zhang, Pengtao Xie

机构 * Department of Electrical and Computer Engineering, University of California San Diego(加州大学圣地亚哥分校电气与计算机工程系) Department of Medicine, University of California San Diego(加州大学圣地亚哥分校医学系)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 针对现有自动构建AI模型的智能体因依赖大语言模型静态参数知识而性能受限的问题,提出AIBuildAI-2,通过引入分层、可进化的外部知识系统,动态加载相关上下文,实现设计决策的专家知识支撑,在MLE-Bench上取得70.7%奖牌率并在心脏病预测竞赛中排名前6.6%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27404 2026-05-28 cs.CY cs.AI 70%

Smaller, Younger, and More Impactful: How AI-Assisted Writing Transforms Research Teams

更小、更年轻、更具影响力:AI辅助写作如何改变研究团队

Haoyang Wang, Mingze Zhang, Yi Bu, Star Xing Zhao, Meijun Liu

机构 * School of Information, University of Texas at Austin(德克萨斯大学奥斯汀分校信息学院) National Science Library, Chinese Academy of Sciences(中国科学院国家科学图书馆) Department of Information Resources Management, School of Economics and Management, University of Chinese Academy of Sciences(中国科学院大学经济管理学院信息资源管理系) Department of Information Management, Peking University(北京大学信息管理系) Institute of Big Data, Fudan University(复旦大学大数据研究院) Institute for Global Public Policy, Fudan University(复旦大学全球公共政策研究院) Faculty of Finance, City University of Macau(澳门城市大学金融学院)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本研究利用2020年以来的PLoS和Nature系列期刊全文,通过多种回归方法发现AI辅助写作使研究团队更年轻、规模更小,且不影响甚至提升科学影响力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06915 2026-05-28 cs.LG 70%

LLMs are not (consistently) Bayesian: Quantifying internal (in)consistencies of LLMs' probabilistic beliefs

LLMs 并非(一致地)贝叶斯:量化 LLMs 概率信念的内部(不)一致性

Chacha Chen, Matthew Jörke, Adam Goliński, Masha Fedzechkina, Guillermo Sapiro, Sinead Williamson, Nicholas Foti

机构 * Apple(苹果公司) Stanford University(斯坦福大学) Princeton University(普林斯顿大学)

专题命中 领域大模型 :LLM(abstract,abstract_cn);分类 cs.LG

AI总结 本文引入信息处理差距来研究 LLMs 在更新概率信念时的内部不一致性,发现非贝叶斯启发式更新在下游任务中常优于精确贝叶斯计算,表明 LLMs 的世界概率模型存在错误设定。

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.12986 2026-05-28 cs.CL cs.AI 62%

Measuring Massive Multitask Chinese Understanding

测量大规模多任务中文理解

Hui Zeng

机构 * Besteasy (Beijing) Language Technology Co., Ltd.(北京最佳语言科技有限公司)

专题命中 领域大模型 :language model(abstract);分类 cs.CL、cs.AI

AI总结 针对中文大语言模型缺乏能力评估的问题,提出一个涵盖医学、法律、心理学和教育四大领域共23个子任务的多任务测试,通过零样本准确率评估模型性能,发现最佳模型平均领先最差模型18.6个百分点,且所有模型在法律领域表现最差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28521 2026-05-28 cs.CL 57%

ClinicalEncoder26AM: A Multlilingual Diagnosable ColBERT Model; Evidences from the MultiClinNER Shared Task

ClinicalEncoder26AM:一个多语言可诊断的ColBERT模型——来自MultiClinNER共享任务的证据

François Remy

机构 * Parallia AI

专题命中 领域大模型 :post-training(abstract);分类 cs.CL

AI总结 本文提出ClinicalEncoder26AM,一个基于BGE-M3的多语言可诊断ColBERT模型,通过多适配器蒸馏和ColBERT式检索目标进行临床后训练,在MultiClinNER任务中微调为BIO标注器,实现了最先进的多语言实体召回率和字符加权F1分数前五。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28211 2026-05-28 cs.CL 57%

When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR

当有用上下文泄露:领域自适应ASR中的隐私风险

Maike Züfle, Jan Niehues

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

专题命中 领域大模型 :prompting(abstract);分类 cs.CL

AI总结 本文识别并系统研究了领域自适应ASR中因上下文提示或微调导致模型泄露隐私的风险,通过构建控制数据集测量泄露率,并评估了提示级缓解策略及精度-泄露权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26277 2026-05-28 cs.CV cs.AI 57%

VesselSim: learning 3D blood vessel segmentation without expert annotations

VesselSim: 无需专家标注的3D血管分割学习

Erin Rainville, Melissa Ananian, Tristan Mirolla, Hassan Rivaz, Yiming Xiao

机构 * Department of Computer Science and Software Engineering, Concordia University, Montreal, Canada(计算机科学与软件工程系,康科迪亚大学,蒙特利尔,加拿大) Department of Electrical and Computer Engineering, Concordia University, Montreal, Canada(电气与计算机工程系,康科迪亚大学,蒙特利尔,加拿大)

专题命中 领域大模型 :foundation model(abstract);分类 cs.AI

AI总结 提出VesselSim两阶段框架,通过几何驱动的合成血管生成和自监督测试时适应,实现无需真实标注的3D血管分割,在多个临床数据集上达到与有监督方法竞争的性能。

Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution will be published as part of the MICCAI 2026 proceedings in October

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28129 2026-05-28 cs.AI 57%

Do Clinical Models Change Treatment Decisions?

临床模型是否会改变治疗决策?

Dongkyu Cho, Miao Zhang, Rumi Chunara

机构 * New York University(纽约大学)

专题命中 领域大模型 :foundation model(abstract);分类 cs.AI

AI总结 本研究提出ClinPivot基准,通过生物医学关系和变化的患者情境评估临床基础模型在治疗决策中的适应性,发现强医学QA能力不能可靠预测决策表现。

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10073 2026-05-28 cs.CL 57%

Heterogeneous Dependency Graph-Guided Attentionfor Patent Representation Learning

异构依赖图引导的专利表示学习注意力机制

Yongmin Yoo, Qiongkai Xu, Zhangkai Wu, Longbing Cao

机构 * Frontier AI Research Centre, Macquarie University School of Computing, FSE, Macquarie University(前沿人工智能研究中心,麦考瑞大学计算机学院,FSE,麦考瑞大学)

专题命中 领域大模型 :language model(abstract);分类 cs.CL

AI总结 针对专利权利要求间的依赖层次被忽略的问题,提出专利异构注意力图编码器(PHAGE),通过构建类型图区分法律引用与技术关系,并引入可学习偏置的连通性掩码将权利要求级拓扑投射到令牌级注意力,结合双粒度对比学习,在分类、检索和聚类任务上超越领域自适应和引用感知基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02097 2026-05-28 cs.CL 57%

ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs

ClinConsensus:一个用于评估中文医疗大模型临床评分标准覆盖率的医师校准基准

Xiang Zheng, Han Li, Wenjie Luo, Weiqi Zhai, Yiyuan Li, Chuanmiao Yan, Xue Yang, Kailuan Wu, Ruyi Xu, Tianyun Lu, Tianyi Tang, Yubo Ma, Kexin Yang, Dayiheng Liu, Sen Yang, Lin Qu, Bing Zhao, Hu Wei

机构 * Alibaba Group(阿里巴巴集团)

专题命中 领域大模型 :LLM(abstract);分类 cs.CL

AI总结 为解决开放域医疗大模型评估缺乏医师校准的临床响应标准覆盖率问题,提出包含2500个专家病例的ClinConsensus基准,并引入医师锚定覆盖率评分(CACS)及双裁判框架,发现前沿模型存在19.2-21.9分的覆盖率差距。

详情

展开后加载摘要…

URL PDF HTML 收藏