arXivDaily arXiv每日学术速递 周一至周五更新

科学与医疗

医学 AI

医学智能、临床 AI、医学影像、病理、诊断和医疗健康大模型。

共收录 2124 信号源:cs.CV, cs.LG, q-bio, eess.IV, eess.SP

1. 医学数据与评测 2124 篇

2603.22820 2026-03-25 cs.CL 50%

RadTimeline: Timeline Summarization for Longitudinal Radiological Lung Findings

RadTimeline:纵向放射学肺部发现的时间线摘要

Sitong Zhou, Meliha Yetisgen, Mari Ostendorf

专题命中 医学数据与评测 :radiology(abstract)

AI总结 本文提出RadTimeline时间线摘要任务,通过生成时间线来总结纵向放射学报告,提升疾病进展识别效率。实验表明,分组名称生成对发现分组至关重要,最佳配置在召回率和分组性能上接近人工标注。

Comments Accepted at Language Resources and Evaluation Conference (LREC) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23743 2026-03-24 cs.SE cs.AI 50%

Hybrid-Code v2: Zero-Hallucination Clinical ICD-10 Coding via Neuro-Symbolic Verification and Automated Knowledge Base Expansion

Hybrid-Code v2:通过神经符号验证和自动知识库扩展实现零幻觉的临床ICD-10编码

Yunguo Yu

机构 * AI Innovation & Prototyping, Zyter(人工智能创新与原型开发,Zyter)

专题命中 医学数据与评测 :medical AI(abstract)

AI总结 Hybrid-Code v2结合神经网络与符号验证,实现零类型I幻觉,同时保持高覆盖率和精度,通过自动知识库扩展解决规则系统扩展性问题。

Comments Version 2: Substantially extended version with (1) multi-layer verification framework (format, evidence, negation, temporal, exclusion), (2) automated knowledge base expansion from unlabeled clinical text, (3) formal zero Type-I hallucination guarantees, and (4) expanded experimental evaluation on 5,000 cases with detailed error analysis. 28 pages, 3 figure, original research paper;

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02868 2026-03-24 cs.AI 50%

PrecLLM: A Privacy-Preserving Framework for Efficient Clinical Annotation Extraction from Unstructured EHRs using Small-Scale LLMs

PrecLLM: 一种用于从非结构化电子健康记录中高效提取临床注释的隐私保护框架,使用小型语言模型

Yixiang Qu, Yifan Dai, Shilin Yu, Pradham Tanikella, Malvika Pillai, Walter Chen, Jialiu Xie, Yishan Ren, Duan Wang, Yikai Wang, Sid Sheth, Guanting Chen, Yufeng Liu, Travis Schrank, Trevor Hackman, Didong Li, Di Wu

机构 * Department of Biostatistics, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校生物统计学系) Department of Genetics, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校遗传学系) Curriculum for Bioinformatics and Computational Biology, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校生物信息学与计算生物学课程) Carolina Health Informatics Program, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校健康信息学计划) Department of Statistics and Operations Research, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校统计学与运筹学系) Department of Otolaryngology/Head and Neck Surgery, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校耳鼻喉科及头颈外科系) Department of Statistics, University of Michigan(密歇根大学统计学系) Department of Biomedical Sciences, Adams School of Dentistry, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校阿德姆牙科学院生物医学科学系) Computational Medicine Program, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校计算医学计划) Lineberger Comprehensive Cancer Center, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校林伯格综合癌症中心)

专题命中 医学数据与评测 :clinical LLM(abstract)

AI总结 本文提出PrecLLM框架,利用小型语言模型高效处理非结构化电子健康记录,通过正则表达式和RAG技术提升隐私保护下的临床注释提取性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19313 2026-03-23 cs.CL cs.AI 50%

Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs

基于记忆的角色扮演:评估和增强LLM中人设知识的利用

Kai Wang, Haoyang You, Yang Zhang, Zhongjie Wang

机构 * Harbin Institute of Technology(哈尔滨工业大学) Macquarie University(麦考瑞大学)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 本文提出Memory-Driven Role-Playing框架,通过MREval、MRPrompt和MRBench评估LLM在角色扮演中的记忆能力,验证了小模型能效媲美大模型,并证明上游记忆提升下游响应质量。

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19274 2026-03-23 cs.CL cs.AI 50%

CURE: A Multimodal Benchmark for Clinical Understanding and Retrieval Evaluation

CURE:临床理解和检索评估的多模态基准

Yannian Gu, Zhongzhen Huang, Linjie Mu, Xizhuo Zhang, Shaoting Zhang, Xiaofan Zhang

机构 * Shanghai Jiao Tong University, Shanghai, China(上海交通大学) SenseTime Research, China(商汤研究院) Shanghai Innovation Institute, Shanghai, China(上海创新研究院)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 CURE基准通过控制证据环境评估多模态模型的推理与检索能力,揭示模型在依赖权威文献时性能显著下降,强调整合临床证据与精确检索的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18637 2026-03-20 cs.CR cs.CL 50%

MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment

MOSAIC:多目标切片感知迭代校准用于对齐

Yipu Dou, Wang Yang

机构 * School of Cyber Science and Engineering(网络科学与工程学院)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 本文提出MOSAIC框架,通过统一的L1-L3评估接口实现闭环数据混合搜索,在有限预算下平衡多目标,提升对齐性能并优于随机基线。

Comments 9 pages, 5 figures. Code available at https://github.com/douyipu/mosaic

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00917 2026-03-18 cs.CL cs.AI 50%

Prompt Sensitivity and Answer Consistency of Small Open-Source Language Models for Clinical Question Answering in Low-Resource Healthcare

小开源语言模型在低资源医疗场景中的提示敏感性与答案一致性

Shravani Hariprasad

机构 * Independent Researcher(独立研究者)

专题命中 医学数据与评测 :clinical AI(abstract)

AI总结 研究评估了五种开源模型在三个医疗问答数据集上的表现,发现一致性与准确性独立,Llama 3.2在准确性与可靠性上表现最佳,但领域预训练不足以保证结构化医疗问答的正确性。

Comments 30 pages, 7 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15483 2026-03-17 cs.AI 50%

Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error Analysis

谈话、评估、诊断:基于用户意识的代理评估与自动错误分析

Penny Chong, Harshavardhan Abichandani, Jiyuan Shen, Atin Ghosh, Min Pyae Moe, Yifan Mai, Daniel Dahlmeier

机构 * SAP Stanford University(斯坦福大学)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 本文提出TED框架,通过用户角色模板、自动评分和错误分析,提升代理评估的全面性与效率,实验显示在模型和用户水平上取得8-10%的性能提升。

Comments Accepted as a conference paper at ICLR 2026. Code and dataset are available in the repository https://github.com/SAP-samples/agent-quality-inspect

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14265 2026-03-17 cs.CL cs.MA 50%

MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-End Question Answering

MedPriv-Bench:医疗开放问答中大语言模型隐私-效用权衡的基准测试

Shaowei Guan, Yu Zhai, Hin Chi Kwok, Jiawei Du, Xinyu Feng, Jing Li, Harry Qin, Vivian Hui

专题命中 医学数据与评测 :medical AI(abstract)

AI总结 MedPriv-Bench首次提出医疗开放问答中评估隐私保护与临床效用的基准,通过多代理人机协同流程生成敏感医疗情境,利用预训练RoBERTa-NLI模型量化数据泄露,揭示大语言模型在隐私与效用间的普遍权衡。

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12710 2026-03-16 cs.AI cs.CL 50%

AI Planning Framework for LLM-Based Web Agents

基于大语言模型的网络代理的AI规划框架

Orit Shahnovsky, Rotem Dror

机构 * Faculty of Computer and Information Science, University of Haifa(计算机与信息科学学院,海法大学)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 本文提出一种将网络任务视为序列决策过程的AI规划框架,通过将现代代理架构映射到传统规划范式,提出五种新的评估指标,验证了全计划提前代理在技术指标上的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12458 2026-03-16 cs.CL cs.AI 50%

Shattering the Shortcut: A Topology-Regularized Benchmark for Multi-hop Medical Reasoning in LLMs

打破捷径:一种用于LLMs多跳医学推理的拓扑正则化基准

Xing Zi, Xinying Zhou, Jinghao Xiao, Catarina Moreira, Mukesh Prasad

机构 * University of Technology Sydney(悉尼技术大学)

专题命中 医学数据与评测 :medical AI(abstract)

AI总结 本文提出ShatterMed-QA基准,通过拓扑正则化知识图谱评估深度诊断推理能力,揭示LLMs在多跳医学推理中的缺陷,并验证RAG方法的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08363 2026-03-13 cs.IR cs.CL 50%

PosIR: Position-Aware Heterogeneous Information Retrieval Benchmark

PosIR:位置感知的异构信息检索基准

Ziyang Zeng, Dun Zhang, Yu Yan, Xu Sun, Cuiqiaoshu Pan, Yudong Zhou, Yuqing Yang

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 PosIR是一个标准化的异构信息检索基准,通过长度控制分桶策略系统诊断位置偏差,揭示了检索模型在长文档中存在优先偏差及近期偏差,暴露了短文本评估的局限性。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14916 2026-03-11 cs.CL cs.AI 50%

MKE-Coder: Multi-Axial Knowledge with Evidence Verification in ICD Coding for Chinese EMRs

MKE-Coder:基于中国电子病历的ICD编码中的多轴知识与证据验证

Xinxin You, Xien Liu, Xue Yang, Ziyi Wang, Ji Wu

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 MKE-Coder通过多轴知识与证据验证方法提升中文电子病历的ICD自动编码准确性与效率。

Comments We identified an error in the data preprocessing script that led to inconsistent results in the tables. As the current version contains inaccurate data, we are withdrawing it for further correction and verification

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22758 2026-03-10 cs.AI stat.AP 50%

Decomposing Physician Disagreement in HealthBench

分解健康基准中医生的分歧

Satya Borgohain, Roy Mariathas

专题命中 医学数据与评测 :medical AI(abstract)

AI总结 研究通过分解健康基准数据集中的医生分歧,发现可减少的不确定性显著影响分歧程度,而不可减少的不确定性则无影响,提示评估设计改进的方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00657 2026-03-10 cs.SD 50%

XPPG-PCA: Reference-free automatic speech severity evaluation with principal components

XPPG-PCA:无参考自动语音严重程度评估的主成分分析

Bence Mark Halpern, Thomas B. Tienkamp, Teja Rebernik, Rob J. J. H. van Son, Sebastiaan A. H. J. de Visscher, Max J. H. Witjes, Defne Abur, Tomoki Toda

机构 * Nagoya University, Japan(日本名古屋大学) Netherlands Cancer Institute(荷兰癌症研究所) University of Groningen(格罗宁根大学) University Medical Center Groningen(格罗宁根大学医学中心) CNRS, Sorbonne Nouvelle(法国国家科学研究中心、索邦-努瓦塞勒大学) University of Amsterdam(阿姆斯特丹大学) University Medical Hospital Groningen(格罗宁根大学医学中心)

专题命中 医学数据与评测 :pathology(abstract)

AI总结 XPPG-PCA提出了一种无参考、无监督的语音严重程度评估方法,通过主成分分析在三个荷兰口腔癌数据集上验证了其性能,展示了在临床应用中的潜力。

Comments 14 pages, 4 figures. Author Accepted Manuscript version of the IEEE Selected Topics in Signal Processing with the same title

Journal ref IEEE Journal of Selected Topics in Signal Processing 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07078 2026-03-10 cs.AI cs.CL 50%

CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs

CoTJudger: 一种基于图的框架,用于自动评估链式推理效率和冗余性在LRMs中

Siyi Li, Jiajun Shi, Shiwen Ni, Ge Zhang, Shuaimin Li, Shijian Wang, Zhoufutu Wen, Yizhi Li, Hamid Alinejad-Rokny, Jiaheng Liu, Min Yang, Wenhao Huang

机构 * University of Science and Technology of China(科学技术大学) Shenzhen University of Advanced Technology(深圳先进技术大学) Shenzhen Institutes of Advanced Technology, CAS(深圳先进技术研究所,中国科学院) Southeast University(东南大学) Nanjing University(南京大学) Beihang University(北航) University of Manchester(曼彻斯特大学)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 CoTJudger通过构建依赖图提取最短有效路径,评估链式推理的效率与冗余,揭示模型中的冗余问题及失败模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15796 2026-03-05 cs.AI 50%

From Privacy to Trust in the Agentic Era: A Taxonomy of Challenges in Trustworthy Federated Learning Through the Lens of Trust Report 2.0

从隐私到信任在代理时代:通过信任报告2.0的视角,对可信联邦学习中挑战的分类

Nuria Rodríguez-Barroso, Mario García-Márquez, M. Victoria Luzón, Francisco Herrera

机构 * Department of Computer Science and Artificial Intelligence, Andalusian Research Institute in Data Science and Computational Intelligence (DaSCI) University of Granada(计算机科学与人工智能系,数据科学与计算智能安达卢西亚研究 institute,格拉纳达大学) Department of Software Engineering, Andalusian Research Institute in Data Science and Computational Intelligence (DaSCI) University of Granada(软件工程系,数据科学与计算智能安达卢西亚研究 institute,格拉纳达大学)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 本文通过信任报告2.0提出可信联邦学习的挑战分类,强调信任作为持续维持的操作条件,并引入协调蓝图以处理跨要求的权衡和治理对齐。

Comments Already published in Information Fusion

Journal ref Rodríguez-Barroso, et. al. (2026). From Privacy to Trust in the Agentic Era: A Taxonomy of Challenges in Trustworthy Federated Learning Through the Lens of Trust Report 2.0. Information Fusion, 104236

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02799 2026-03-04 cs.CR 50%

Understanding the Resource Cost of Fully Homomorphic Encryption in Quantum Federated Learning

理解量子联邦学习中全同态加密的资源成本

Lukas Böhm, Arjhun Swaminathan, Anika Hannemann, Erik Buchmann

专题命中 医学数据与评测 :MRI(abstract)

AI总结 本研究评估了全同态加密在量子联邦学习中带来的资源开销,发现其在现实应用中面临内存和通信开销大的挑战,需在隐私与模型复杂性间做出权衡。

Comments Experiments with Quantum Federated Learning using Homomorphic Encryption to encrypt the gradients

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02528 2026-03-04 cs.AI cs.RO 50%

LLM-MLFFN: Multi-Level Autonomous Driving Behavior Feature Fusion via Large Language Model

LLM-MLFFN: 通过大语言模型的多级自动驾驶行为特征融合

Xiangyu Li, Tianyi Wang, Xi Cheng, Rakesh Chowdary Machineni, Zhaomiao Guo, Sikai Chen, Junfeng Jiao, Christian Claudel

机构 * Department of Civil, Architectural, and Environmental Engineering, The University of Texas at Austin(德克萨斯大学奥斯汀分校土木、建筑与环境工程系) Systems Engineering Program, Cornell University(康奈尔大学系统工程项目) Department of Electrical and Computer Engineering, University of Michigan(密歇根大学电气与计算机工程系) Department of Civil and Environmental Engineering, University of Wisconsin-Madison(威斯康星大学麦迪逊分校土木与环境工程系) School of Architecture, The University of Texas at Austin(德克萨斯大学奥斯汀分校建筑学院)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 LLM-MLFFN通过大语言模型的多级特征融合提升自动驾驶行为分类的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01780 2026-02-27 cs.AI cs.CL 50%

LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?

LiveMCPBench: 代理能否在MCP工具的海洋中导航?

Guozhao Mo, Wenliang Zhong, Jiawei Chen, Qianhao Yuan, Xuanang Chen, Yaojie Lu, Hongyu Lin, Ben He, Xianpei Han, Le Sun

机构 * University of Chinese Academy of Sciences(中国科学院大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 LiveMCPBench通过大规模基准测试揭示MCP代理在检索和工具组合上的性能差距,强调检索错误是主要瓶颈,并提供可重复的评估框架和工具套件。

Comments Our code and data will be publicly available at https://icip-cas.github.io/LiveMCPBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19139 2026-02-26 cs.AI cs.CL 50%

A Multi-faceted Analysis of Cognitive Abilities: Evaluating Prompt Methods with Large Language Models on the CONSORT Checklist

对认知能力的多维度分析:利用大型语言模型在CONSORT清单上评估提示方法

Sohyeon Jeon, Hyung-Chul Lee

专题命中 医学数据与评测 :medical AI(abstract)

AI总结 本研究通过比较通用和领域专用LLM在三种提示策略下的表现,揭示了其在评估临床试验报告时存在显著的校准错误和过度自信问题,强调了改进校准和提示工程的重要性。

Comments We have decided to withdraw this manuscript because we believe it requires further revision and substantial improvement before it is suitable for dissemination to the academic community

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20823 2026-02-25 cs.SD eess.AS 50%

Geometric Analysis of Speech Representation Spaces: Topological Disentanglement and Confound Detection

语音表示空间的几何分析:拓扑解缠与混淆检测

Bipasha Kashyap, Pubudu N. Pathirana

机构 * Networked Sensing \& Biomedical Engineering (NSBE) Research lab, School of Engineering

专题命中 医学数据与评测 :pathology(abstract)

AI总结 本文提出四指标聚类框架,评估语音特征在多语言环境中的几何解缠与混淆检测,为构建公平可靠的语音健康系统提供指导。

Comments Submitted to INTERSPEECH 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19843 2026-02-24 cs.SE cs.AI 50%

MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems

MAS-FIRE:基于大语言模型的多智能体系统故障注入与可靠性评估

Jin Jia, Zhiling Deng, Zhuangbin Chen, Yingqi Wang, Zibin Zheng

机构 * Sun Yat-sen University(中山大学)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 MAS-FIRE通过故障注入和分层分析,揭示多智能体系统在容错和鲁棒性方面的关键因素,为提升系统可靠性提供系统性方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16777 2026-02-20 cs.CL 50%

DistillNote: Toward a Functional Evaluation Framework of LLM-Generated Clinical Note Summaries

DistillNote:一种评估大语言模型生成临床笔记摘要功能性的框架

Heloisa Oss Boll, Antonio Oss Boll, Leticia Puttlitz Boll, Ameen Abu Hanna, Iacer Calixto

机构 * Department of Medical Informatics, Amsterdam UMC, University of Amsterdam(医学信息学系,阿姆斯特丹大学) Amsterdam Public Health, Methodology(阿姆斯特丹公共健康与方法学) Institute of Mathematics and Statistics, University of São Paulo(数学与统计学研究所,圣保罗大学)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 DistillNote提出了一种评估大语言模型生成临床摘要功能性的新方法,通过下游任务验证摘要的诊断信号保留,揭示压缩与性能的权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16051 2026-02-18 cs.CL 50%

Moving Beyond Medical Exams: A Clinician-Annotated Fairness Dataset of Real-World Tasks and Ambiguity in Mental Healthcare

超越医学考试:一个由临床专家标注的现实任务及精神卫生领域模糊性公平性数据集

Max Lamparth, Declan Grabb, Amy Franks, Scott Gershan, Kaitlyn N. Kunstman, Aaron Lulla, Monika Drummond Roots, Manu Sharma, Aryan Shrivastava, Nina Vasan, Colleen Waickman

机构 * Stanford University(斯坦福大学) University of Colorado(科罗拉多大学) Northwestern University(西北大学) University of Wisconsin(威斯康星大学) Yale School of Medicine(耶鲁医学院) University of Chicago(芝加哥大学) Ohio State University(俄亥俄州立大学)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 本文提出一个由临床专家标注的现实任务公平性数据集,用于评估模型在精神卫生领域决策中的性能和偏见问题。

Comments Camera-ready version for ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14655 2026-02-17 cs.CL cs.AI 50%

Breaking Data Efficiency Dilemma: A Federated and Augmented Learning Framework For Alzheimer's Disease Detection via Speech

突破数据效率困境:一种联邦学习与增强学习框架用于通过语音检测阿尔茨海默病

Xiao Wei, Bin Wen, Yuqin Lin, Kai Li, Mingyang gu, Xiaobao Wang, Longbiao Wang, Jianwu Dang

机构 * Tianjin Key Laboratory of Cognitive Computing and Application(认知计算与应用天津重点实验室) College of Intelligence and Computing, Tianjin University(智能计算学院,天津大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) College of Computer and Data Science, Fuzhou University(计算机与数据科学学院,福州大学) Huiyan Technology (Tianjin) Co., Ltd(慧研科技(天津)有限公司)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 FAL-AD通过联邦学习与数据增强框架,实现阿尔茨海默病语音检测中的数据效率提升,达到91.52%的多模态准确率。

Comments 5 pages, 1 figures, accepted by ICASSP 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06968 2026-02-12 cs.AI cs.CL 50%

Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report

向指令集的信息边界扩展:Infinity Instruct主体技术报告

Li Du, Hanyu Zhao, Yiming Ju, Tengfei Pan

机构 * BAAI(北京人工智能研究院)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 本文提出系统化的指令数据构建框架,构建高质量的Infinity Instruct Subject数据集,提升模型在复杂指令任务上的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07740 2026-02-10 stat.ME 50%

Hyperbolic statistical inference for Treatment Effects with Circular biomarker of astigmatism

双曲统计推断用于治疗效应的圆环生物标志物:散光

Buddhananda Banerjee, Surojit Biswas, Daitari Prusty

专题命中 医学数据与评测 :biomedical(abstract)

AI总结 本文提出基于双曲几何的圆环数据双样本检验方法,用于分析白内障手术引起的散光治疗效应,通过几何嵌入提升参数空间的连续表示和检验统计量的可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17118 2026-02-10 cs.CR 50%

Constant-Size Cryptographic Evidence Structures for Regulated AI Workflows

固定大小的加密证据结构用于受监管的AI工作流

Leo Kao

专题命中 医学数据与评测 :medical AI(abstract)

AI总结 本文提出固定大小的加密证据结构,用于受监管AI工作流,以实现强绑定、固定存储和统一验证,适用于临床试验、医疗决策支持等场景。

Comments 16 pages, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04813 2026-02-05 cs.AI cs.CY 50%

Agentic AI in Healthcare & Medicine: A Seven-Dimensional Taxonomy for Empirical Evaluation of LLM-based Agents

医疗与医学中的代理AI:一个七维分类法用于评估基于大语言模型的代理

Shubham Vatsal, Harsh Dubey, Aditi Singh

机构 * Department of Computer Science, New York University, CIMS, New York, USA(纽约大学计算机科学系,CIMS,纽约,美国) Department of Computer Science, Cleveland State University, Cleveland, USA(克利夫兰州立大学计算机科学系,克利夫兰,美国)

专题命中 医学数据与评测 :diagnosis(abstract)

AI总结 本文提出一个七维分类法,用于评估基于大语言模型的医疗代理在认知能力、知识管理、交互模式等方面的能力分布与实现情况。

Journal ref IEEE Access, vol. 14, pp. 4840-4863, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏