arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-07-15 至 2026-07-15 共收录 215 信号源:cs.CL, cs.AI, cs.LG

1. 效率与部署 34 篇

2511.11132 2026-07-15 cs.CV 版本更新 80%

From Hindsight to Foresight: Self-Encouraged Hindsight Distillation for Knowledge-based Visual Question Answering

从回顾到前瞻:面向知识驱动视觉问答的自我鼓励回顾蒸馏

Yu Zhao, Ying Zhang, Xuhui Sui, Baohang Zhou, Xinying Qian, Li Shen, Dacheng Tao

机构 * College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 效率与部署 :large language model(abstract);language model(abstract);preference optimization(abstract);prompting(abstract)

AI总结 本文提出HinD框架,通过知识鼓励偏好优化提升多模态大语言模型的知识推理能力,实验表明其在OK-VQA和A-OKVQA上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12188 2026-07-15 cs.AI cs.DB cs.IR 新提交 79%

Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems

成本治理的检索增强生成模型:多租户大语言模型系统中跨检索与生成的统一租户成本归因

Navnit Shukla

专题命中 效率与部署 :LLM(title,abstract);分类 cs.AI

AI总结 研究多租户大语言模型系统中成本治理问题,提出成本治理的RAG架构,集成TurboVec与治理网关实现统一可观测堆栈,能按租户联合归因成本,准确率高且降低成本,还形式化三层成本模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12747 2026-07-15 cs.AI cs.CL 新提交 79%

Tracing Agentic Failure from the Flow of Success

从成功流中追溯智能体失败

Samuel Yeh, Yiwen Zhu, Shaleen Deep, Sharon Li

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Microsoft Research(微软研究院)

专题命中 效率与部署 :LLM(abstract);post-training(abstract);prompting(abstract);分类 cs.CL、cs.AI

AI总结 研究基于大语言模型的智能体系统失败归因问题,提出OAT方法将其转化为单类学习,用神经控制微分方程建模成功轨迹动态模式,实验表明该方法比基线快且F1分数更高,是诊断智能体系统失败的有效方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12875 2026-07-15 cs.MA cs.SE 新提交 78%

MetaInfer: A Knowledge Only LLM Inference Engine Generator SKILL Toolbox

MetaInfer:一个仅基于知识的大语言模型推理引擎生成器SKILL工具包

Zhenwen Miao, Honglin Wang, Mingheng Mi

专题命中 效率与部署 :LLM(title,abstract)

AI总结 研究针对大语言模型推理框架代码复杂、维护成本高的问题,提出MetaInfer方法,通过用户指定运行时约束,利用大语言模型驱动多智能体协作系统结合契约知识库自动生成定制推理框架,并从多视角评估,实现从显式知识生成可运行方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12555 2026-07-15 cs.IT math.IT 新提交 78%

FM-Receiver: A Foundation Model Enabled Unified Inner and Outer Neural Receiver Towards AI-Native Wireless Communications

FM-接收机:一种基于基础模型的统一内外部神经接收机,迈向人工智能原生无线通信

Tianyue Zheng, Chao Jiang, Linglong Dai

专题命中 效率与部署 :foundation model(title,abstract)

AI总结 针对多数神经接收机内外接收机未联合优化的问题,提出基于基础模型的FM-Receiver,利用分组纠错码Transformer实现符号级信道解码,集成内外接收机,并设计预训练策略,提升泛化能力,在不同配置下性能优于基线且有零样本泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12696 2026-07-15 cs.CL cs.AI cs.DC 新提交 73%

Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts

更少专家,更快解码:用于专家混合模型的成本感知推测解码

Jincheng Xie, Runheng Liu, Heyan Huang, Yawen Ling, Hanbin Dai, Yu Zheng, Wen Hu

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究针对大规模专家混合模型推理效率受专家激活模式影响的问题,提出成本感知推测解码框架EcoSpec,通过纳入预测的边际专家激活成本进行草稿选择,在多个模型和基准测试中减少活跃专家足迹、提高解码速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12634 2026-07-15 cs.AI 新提交 70%

Atomic Units of X: The Compression Layer of Intelligence

X的原子单位:智能的压缩层

Sachin Dev Duggal, Pradyumna Swarnalatha Ramanna, Alexandros Vassiliades

机构 * SeKondBrain AI Labs(第二大脑人工智能实验室)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该论文提出将智能视为原子压缩与组合复用过程的理论框架,借助多学科证据阐述原子单位概念,给出压缩演算等核心方法,为设计自进化知识系统奠基,为智能相关多方面提供统一视角。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12765 2026-07-15 cs.LG 版本更新 70%

Inference-Time Machine Unlearning via Gated Activation Redirection

推理时的机器去学习 via 门控激活重定向

Vinícius Conte Turani, Otávio Parraga, João Vitor Boer Abitante, Kristen K. Arguello, Joana Pasquali, Ramiro N. Barros, Flavio du Pin Calmon, Christian Mattjie, Rodrigo C. Barros, Lucas S. Kupssinskü

机构 * MALTA, Machine Learning Theory and Applications Lab, PUCRS, Porto Alegre, Brazil(MALTA机器学习理论与应用实验室,PUCRS,波士顿-阿尔格雷,巴西) Harvard University(哈佛大学) Kunumi Institute, Brazil(库努米研究所,巴西)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出了一种无需训练和梯度的机器去学习方法GUARD-IT,通过在推理时依赖输入的激活引导来消除特定数据集的影响,同时保持模型性能,且在量化部署下仍有效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16918 2026-07-15 cs.CV cs.AI 版本更新 70%

Xray-Visual Models: Scaling Vision models on Industry Scale Data

Xray-Visual模型:在产业级数据上扩展视觉模型

Shlok Mishra, Tsung-Yu Lin, Linda Wang, Hongli Xu, Yimin Liu, Michael Hsu, Chaitanya Ahuja, Hao Yuan, Jianpeng Cheng, Hong-You Chen, Haoyuan Xu, Chao Li, Sreya Dutta Roy, Abhijeet Awasthi, Jihye Moon, Don Husa, Michael Ge, Sumedha Singla, Arkabandhu Chowdhury, Phong Dingh, Satya Narayan Shukla, Yonghuan Yang, David Jacobs, Qi Guo, Jun Xiao, Xiangjun Fan, Aashu Singh

机构 * Meta-AI MIT(麻省理工学院) University of Maryland(马里兰大学)

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 Xray-Visual通过三阶段训练流程和LLM2CLIP技术,在产业级数据上实现高效多模态视觉模型,取得最佳性能并提升鲁棒性与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11443 2026-07-15 cs.CL 版本更新 70%

Predict the Retrieval! Test time adaptation for Retrieval Augmented Generation

预测检索!面向检索增强生成的测试时自适应方法

Xin Sun, Zhongqi Chen, Qiang Liu, Shu Wu, Bowen Song, Weiqiang Wang, Zilei Wang, Liang Wang

专题命中 效率与部署 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 提出TTARAG方法,在测试时动态更新语言模型参数以预测检索内容,从而提升RAG系统在专业领域的泛化性能。

Comments ICASSP 2026

Journal ref ICASSP 2026 - 2026 IEEE International Conference on Acoustics, ICASSP 2026 - 2026 IEEE International Conference on Acoustics, ICASSP 2026 - 2026 IEEE International Conference on Acoustics,

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12894 2026-07-15 cs.CV 新提交 67%

Hy-Embodied-VLM-1.0: Efficient Physical-World Agents

Hy-Embodied-VLM-1.0:高效的物理世界智能体

Ziyi Wang, Xumin Yu, Yongming Rao, Yonggen Ling, Yunheng Li, Oran Wang, Mingqi Gao, Yuchen Zhou, Yves Liang, Zuyan Liu, Yani Zhang, Rui Huang, Xiaoran Xu, Bowen Yuan, Yifu Yuan, Xu Tan, He Zhang, Yufei Huang, Shenghao Zhang, Hongsheng Wu, Han Hu, Zhengyou Zhang

机构 * Tencent Robotics X(腾讯Robotics X团队) Hy Vision Team(腾讯混元视觉团队) Futian Laboratory(福田实验室)

专题命中 效率与部署 :foundation model(abstract);post-training(abstract)

AI总结 研究旨在构建物理世界具身智能体,介绍Hy-Embodied-VLM-1.0模型。定义以行动为中心的能力分类法,开发数据管道。基于特定主干和编码器构建模型,用专家混合架构提升效率。在多基准测试中性能出色,较上一代有显著提升,在具身智能任务中也表现强大。

Comments Tech Report. Code and models are open-sourced at https://github.com/Tencent-Hunyuan/HY-Embodied

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06678 2026-07-15 quant-ph 67%

Learning Variational Quantum Circuit Parameters with Classical Artificial Intelligence for Quantum Phase Transition Detection

利用经典人工智能学习变分量子电路参数以检测量子相变

Xin Li, Zhang-Qi Yin

专题命中 效率与部署 :large language model(abstract);language model(abstract)

AI总结 本文提出利用经典人工智能学习变分量子电路参数,以无监督方式检测量子相变,尤其在拓扑相变识别中表现优异。

Journal ref Phys. Rev. B 113, 235157 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27321 2026-07-15 cs.LG cs.AI 版本更新 62%

Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders

超越硬预算:用于更可解释的Top-k稀疏自编码器的稀疏正则化器

Nathanaël Jacquier, Maria Vakalopoulou, Mahdi S. Hosseini

机构 * Université Paris-Saclay, CentraleSupélec, France(巴黎-萨克雷大学,中央圣艾克苏佩里大学,法国) Department of Computer Science and Software Engineering (CSSE), Concordia University, Montreal, QC, Canada(计算机科学与软件工程系,康科迪亚大学,加拿大) Mila–Quebec AI Institute, Montreal, QC, Canada(魁北克人工智能研究所,加拿大) Université Paris-Saclay, CentraleSupélec, Gustave Roussy, INSERM, IHU PRISM, Cancer Data Science Unit, France(巴黎-萨克雷大学,中央圣艾克苏佩里大学,路易·德·鲁斯医院,国家医学研究院,PRISM机构,癌症数据科学单位,法国) Université Paris-Saclay, CentraleSupélec, MICS Laboratory, France(巴黎-萨克雷大学,中央圣艾克苏佩里大学,MICS实验室,法国)

专题命中 效率与部署 :foundation model(abstract);分类 cs.AI、cs.LG

AI总结 针对Top-k稀疏自编码器固定预算k和过拟合问题,提出两种稀疏正则化器(ℓ1惩罚和ℓ1/ℓ2比例惩罚),在不损失重构质量的前提下提高单语义性,并增强对推理时k选择的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12863 2026-07-15 cs.SE cs.LG 新提交 57%

Toward Localizing and Repairing Bias in Transformer Attention Heads

迈向Transformer注意力头中偏差的定位与修复

Sigma Jahan

机构 * Sigma Jahan

专题命中 效率与部署 :language model(abstract);分类 cs.LG

AI总结 研究Transformer语言模型中偏差输出难以定位和修复的问题,提出白盒头级公平性调试方法ROBIN,通过对注意力头排序并去除偏差子空间来修复偏差,在试点研究中减少WinoBias差距且更好保留语言建模质量。

Comments Accepted in ICSME NIER track, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12341 2026-07-15 cs.CL cs.CR cs.DB 新提交 57%

Policy-Conditioned Constrained Decoding for Column-Level Access Control in Text-to-SQL

文本到SQL中列级访问控制的策略条件约束解码

Ryoto Miyamoto, Xin Fan, Hayato Yamana

机构 * Waseda University(早稻田大学)

专题命中 效率与部署 :prompting(abstract);分类 cs.CL

AI总结 研究文本到SQL中列级访问控制问题,提出PCC-SQL系统,通过整合策略并应用逐令牌对数掩码,在单次解码中消除违规,在三个基准和开源模型上实现0%泄漏率与88.7%覆盖率,还评估了语义对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12169 2026-07-15 cs.CL 新提交 57%

We Hebben Een Serieus Translatie: Modeling Intercomprehension as Probabilistic Inference

我们有一个严肃的翻译:将互理解建模为概率推理

Thomas Hikaru Clark, Edward Gibson, Roger Levy

专题命中 效率与部署 :prompting(abstract);分类 cs.CL

AI总结 研究如何实现零样本跨语言理解,通过扩展噪声信道推理算法模型,在贝叶斯框架下建模互理解,用L1语言模型和通用噪声模型推断L2与L1单词映射,实验表明完整模型性能优于消融模型及大模型零样本提示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26373 2026-07-15 cs.CR cs.AI cs.IR 版本更新 57%

Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model

混合隐私感知语义搜索:受限威胁模型下基于SVD截断文档几何与CKKS加密查询重排序

Sergey Kurilenko

机构 * Moscow Institute of Physics and Technology(莫斯科物理技术学院)

专题命中 效率与部署 :LLM(abstract);分类 cs.AI

AI总结 提出一种混合隐私保护方案,利用SVD截断和秘密旋转保护文档集合,CKKS同态加密保护查询,在百万文档规模下实现亚秒级延迟并保持排序质量,证明对受限攻击者的重构误差下界。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12371 2026-07-15 econ.TH 新提交 50%

First They Came for the Others: A Theory of Divide-and-Conquer

首先他们针对其他人:分而治之理论

Yeon-Koo Che, Jinyuqi Huang, Wooyoung Lim

专题命中 效率与部署 :prompting(abstract)

AI总结 研究分而治之策略成功原因,源于对攻击者意图的认知摩擦,通过分析攻击成本、受害者命运相关性等因素,揭示其分化机制,并探讨行为反应等如何影响对攻击意图的推断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12557 2026-07-15 cs.CV 新提交 50%

Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding

用于长视频理解中事件感知视觉分配的高斯混合模型

Yifan Lu, Ziqi Zhang, Chunfeng Yuan, Jun Gao, Bing Li, Weiming Hu

机构 * Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information, CASIA(中国科学院自动化所多模态信息超智能安全北京市重点实验室) State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(中国科学院自动化所多模态人工智能系统国家重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Hello Group(未知(保留英文)) School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院)

专题命中 效率与部署 :language model(abstract)

AI总结 针对长视频理解中视觉分配问题,提出GMM-EVA方法,利用高斯混合模型建模事件级结构,采用差异化分配策略,在多个长视频基准实验中显著优于均匀采样,以约一半视觉令牌预算达可比性能,凸显高效性。

Comments accepted at PRCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12297 2026-07-15 cs.CV 新提交 50%

MobileSAM2: Lightweight Segment Anything for Spatial Intelligence

MobileSAM2:用于空间智能的轻量级图像分割模型

Kai Jiang, Jiaxing Huang, Jingyi Zhang, Weiying Xie, Yunsong Li, Yufei Wang, Aoran Xiao, Dacheng Tao

机构 * Hong Kong Polytechnic University(香港理工大学) Nanyang Technological University(南洋理工大学) Xidian University(西安电子科技大学) SparcAI Inc.(SparcAI公司)

专题命中 效率与部署 :foundation model(abstract)

AI总结 研究旨在使SAM2更适用于移动设备,提出超图知识蒸馏方法HyperKD,由时间和粒度超图知识蒸馏构成,能有效建模转移知识。还推出MobileSAM2家族,经实验验证其在多基准测试及具身AI任务上有良好泛化性能。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12220 2026-07-15 cs.RO 新提交 50%

Contract-Grounded Behavior Tree Synthesis via Coding Agents

通过编码代理实现基于契约的行为树合成

Jonathan Salfity, Robert Blake Anderson, Mitch Pryor

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 效率与部署 :LLM(abstract)

AI总结 研究从自然语言合成机器人行为树时基础设定易出现的问题,提出基于契约的合成架构,编码代理查询服务器获取契约后合成行为树,经实验评估两个大语言模型,结果显示该架构能实现高验证率和成功率,且可转移到物理硬件。

Comments IEEE RA-L Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12181 2026-07-15 cs.RO cs.HC 新提交 50%

Analysis of Mutual and Referential Human and Robot Gazes in a Collaborative Word Association Game

协作式词语联想游戏中人与机器人相互注视和参照注视的分析

Jens V. Rüppel, Tim Schreiter, Andrey Rudenko, Achim J. Lilienthal

机构 * Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM)(慕尼黑工业大学慕尼黑机器人与机器智能研究所 (MIRMI)) Centre for Applied Autonomous Sensor Systems (AASS), Örebro University(厄勒布鲁大学应用自主传感器系统中心 (AASS)) Robotics Insitute Germany (RIG)(德国机器人研究所 (RIG))

专题命中 效率与部署 :LLM(abstract)

AI总结 研究在协作式词语联想游戏中机器人注视对人类视觉注意力的影响及人类是否向机器人寻求确认注视,通过与NAO机器人实验,用多种方式分析互动,发现任务语言方面掩盖注视影响,为相关设计提供见解。

Comments This paper has been accepted as a Late-Breaking Report to the 35th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), which will be held in Kitakyushu, Netherlands on August 24-28, 2026. Copyright 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 领域大模型 18 篇

2607.00011 2026-07-15 cs.IR cs.AI cs.SE 版本更新 90%

SkillSelect-Serve: QoS-Aware Budgeted Skill Service Recommendation for LLM Agents

SkillSelect-Serve:面向小型LLM代理的预算可控且QoS感知的技能服务推荐与组合

Jingyuan Zheng, Dongjing Wang, Xin Zhang, Hao Chen, Youhuizi Li, Xudong Shen, Haiping Zhang, Butian Huang, Dongjin Yu, Guandong Xu

专题命中 领域大模型 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出SkillSelect-Serve框架,将技能选择建模为服务推荐与组合问题,通过双粒度效用建模和预算约束优化,在35,353个技能和586个任务查询上优于固定top-k检索基线。

Comments 18 pages (14-page main text + appendices), 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12336 2026-07-15 cs.CL cs.AI cs.CY cs.ET cs.HC 新提交 90%

Evaluating Health Misinformation in Low-Resource Languages: Integrating Small Language Models with a Culturally-Sensitive Responsible NLP Framework (Bangla as a Case Study)

评估低资源语言中的健康错误信息:将小语言模型与文化敏感的负责任自然语言处理框架相结合(以孟加拉语为例)

Farnaz Farid, Raihan Alam, Al Al-Areqi, Farhad Ahamed, Muhammad Hassan Khan, Sadia Hossain, Irena Veljanova, Anika Tabassum Binte Hossain

机构 * Western Sydney University(西悉尼大学) Microsoft(微软公司) Excelsia College(埃克塞尔西亚学院) Faulconbridge Health Centre(福尔康布里奇健康中心)

专题命中 领域大模型 :language model(title,abstract);small language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

AI总结 研究针对低资源语言中健康错误信息难检测问题,提出结合小语言模型与文化敏感的负责任自然语言处理框架,以孟加拉语为例进行实验,证明Phi-4表现优,还设计新框架,为评估低资源语言错误信息提供整体视角。

Comments 39 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06294 2026-07-15 q-bio.QM cs.AI 88%

Enhancing Phenotype Recognition in Clinical Notes Using Large Language Models: PhenoBCBERT and PhenoGPT

利用大型语言模型增强临床笔记中的表型识别:PhenoBCBERT和PhenoGPT

Jingye Yang, Cong Liu, Wendy Deng, Da Wu, Chunhua Weng, Yunyun Zhou, Kai Wang

专题命中 领域大模型 :large language model(title,abstract);language model(title,abstract);分类 cs.AI

AI总结 本研究提出PhenoBCBERT和PhenoGPT两种模型,利用大型语言模型提升临床笔记中表型术语的自动识别与提取,从而推动疾病相关生物学研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12051 2026-07-15 cs.CL 新提交 86%

Agentic systems for breast cancer treatment recommendations

用于乳腺癌治疗建议的智能体系统

Vinicius Anjos de Almeida, Nícolas Henrique Borges, Leonardo Vicenzi, Helena Kociolek, Sarah Miriã de Castro Rocha, Frederico Nassif Gomes, Júlia Cristina Ferreira Ribeiro, Lucas Emanuel Silva e Oliveira

机构 * Spesia(斯佩西亚) Faculdade de Medicina, Universidade de São Paulo(圣保罗大学医学院) Laboratory of Artificial Intelligence Applied to Bioinformatics, SEPT, Universidade Federal do Paraná (UFPR)(巴拉那联邦大学人工智能应用于生物信息学实验室,SEPT) Pontifícia Universidade Católica do Paraná (PUCPR)(巴拉那天主大学)

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 研究评估用于乳腺癌治疗建议的智能体LLM系统,用72个真实临床病例和1147个特定病例量表,比较七种流程,最佳配置全局得分为0.594±0.025,工具使用和智能体自主性影响各异,虽能生成相关建议,但用于无监督临床使用仍不足。

Comments Under peer review. Source code available at: https://github.com/GRUPOMED4U/breast_cancer_agents_paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01241 2026-07-15 cs.CY cs.AI 版本更新 86%

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

首先,不伤害:迈向临床安全的大语言模型

David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj, Priyank Jain, Jessica Tran, Matthew Gwiazdon, Arjun Rustagi, Jenelle Jindal, Jacob M. Koshy, Vinay Kadiyala, Anup Agarwal, Bassman Tappuni, Brianna French, Sirus Jesudasen, Christopher V. Cosgriff, Rebanta Chakraborty, Jillian Caldwell, Susan Ziolkowski, David J. Iberri, Robert Diep, Rahul S. Dalal, Kira L. Newman, Kristin Galetta, J. Carl Pallais, Nancy Wei, Kathleen M. Buchheit, David I. Hong, Vartan Pahalyants, Ernest Y. Lee, Allen Shih, Tamara B. Kaplan, Vishnu Ravi, Sarita Khemani, Thomas A. Buckley, April S. Liang, Daniel Shirvani, Advait Patil, Nicholas Marshall, Kanav Chopra, Joel Koh, Adi Badhwar, Anastasia Perez, Austin J. Schoeffler, Mahbuba Tusty, Chase M. Walton, Liam G. McCoy, David J. H. Wu, Yingjie Weng, Sumant Ranji, Kevin Schulman, Nigam H. Shah, Jason Hom, Arnold Milstein, Arjun K. Manrai, Adam Rodman, Jonathan H. Chen, Ethan Goh

机构 * Harvard Combined Dermatology Program(哈佛联合皮肤科项目) Department of Dermatology, Mass General Brigham(麻省总医院皮肤科) Harvard Medical School(哈佛医学院) Stanford Center for Biomedical Informatics Research(斯坦福生物医学信息学研究中心) Stanford University(斯坦福大学) Division of Hospital Medicine, Department of Medicine, Stanford University School of Medicine(斯坦福大学医学院医院医学科) Department of Medicine, Cambridge Health Alliance(剑桥健康联盟医学科) Beth Israel Deaconess Hospital–Plymouth(贝塞斯达德acons医院-普利茅斯) Department of Medicine, University of California, San Francisco(加州大学旧金山分校医学科) Department of Neurology, Stanford University School of Medicine(斯坦福大学医学院神经科) Department of Medicine, Beth Israel Deaconess Medical Center(贝塞斯达德acons医学中心医学科) Division of Cardiology, Department of Medicine, Cambridge Health Alliance(剑桥健康联盟心脏病科) Department of Cardiovascular Medicine, Summa Health System(Summa健康系统心血管医学科) Division of Allergy, Pulmonary, and Critical Care Medicine, Department of Medicine, University of Wisconsin-Madison(威斯康星大学麦迪逊分校医学科过敏、呼吸科和危重医学科) Division of Pulmonary and Critical Care Medicine, Department of Medicine, Massachusetts General Hospital(麻省总医院呼吸科和危重医学科) Center for Immunology and Inflammatory Diseases, Department of Medicine, Massachusetts General Hospital(麻省总医院免疫和炎症疾病中心) Broad Institute of MIT and Harvard(MIT和哈佛Broad研究所) Division of Pulmonary, Critical Care, and Sleep Medicine, Cambridge Health Alliance(剑桥健康联盟呼吸科、危重医学科和睡眠医学科)

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出NOHARM基准,包含1100个初级到专科咨询案例,评估28个LLM的医疗建议安全性,发现高达22.6%的案例存在严重危害风险,其中遗漏错误占80%以上。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09483 2026-07-15 cs.AI 版本更新 83%

CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?

CrochetBench:视觉语言模型能否在钩针领域从描述转向实践?

Peiyu Li, Xiaobao Huang, Ting Hua, Nitesh V. Chawla

机构 * University of Notre Dame(诺特大学)

专题命中 领域大模型 :language model(title,abstract);large language model(abstract);分类 cs.AI

AI总结 探讨视觉语言模型在钩针领域从描述到实践的转变,采用CrochetPARADE DSL进行评估,涵盖多种任务,发现评估转变时性能下降,揭示模型局限性,为评估多模态模型过程能力提供新视角。

Comments ACL2026 main-long

Journal ref Proceedings of ACL 2026, Vol. 1, pp. 13229-13251

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07482 2026-07-15 eess.SP cs.AI 83%

VSLLaVA: a pipeline of large multimodal foundation model for industrial vibration signal analysis

VSLLaVA:一种用于工业振动信号分析的大型多模态基础模型流水线

Qi Li, Xinran Zhang, Jinfeng Huang, Hongliang He, Feibin Zhang, Zhaoye Qin, Fulei Chu

机构 * State Key Laboratory of Tribology, Department of Mechanical Engineering, Tsinghua University(摩擦学国家重点实验室,清华大学机械工程系) Department of Statistics and Data Science, Yale University(耶鲁大学统计与数据科学系)

专题命中 领域大模型 :foundation model(title);LLM(abstract);instruction tuning(abstract);分类 cs.AI

AI总结 VSLLaVA通过专家知识引导的指令微调和双模式评估框架,提升工业振动信号分析的端到端多模态模型性能。

Journal ref Advanced Engineering Informatics 76 (2026): 105023

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12454 2026-07-15 cs.LG 新提交 79%

Exploring Zero-Shot Foundation Models for Multivariate Time Series Anomaly Detection

探索用于多变量时间序列异常检测的零样本基础模型

Martin Uray, Saverio Messineo, Roland Kwitt, Stefan Huber

机构 * Salzburg University of Applied Sciences(萨尔茨堡应用科学大学) Paris Lodron University of Salzburg(萨尔茨堡巴黎洛德龙大学)

专题命中 领域大模型 :foundation model(title,abstract);分类 cs.LG

AI总结 研究多变量时间序列异常检测,探索单变量预测基础模型TimesFM的零样本应用于工业MTSAD,评估两种策略,虽未胜过基线,但发现其在捕获时间动态上过于有效致异常难区分,不过在异常边界误差有峰值,对变化点检测有前景。

Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution will be published in Computer Aided Systems Theory - EUROCAST 2026, Lecture Notes in Computer Science, Springer

详情

展开后加载摘要…

URL PDF HTML 收藏