arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Carnegie Mellon University(卡内基梅隆大学)

2026-07-10 至 2026-07-10 共收录 8
2607.08756 2026-07-10 cs.SD cs.LG 新提交

MulTTiPop: A Multitrack Transcription Dataset for Pop Music

MulTTiPop:一个用于流行音乐的多轨转录数据集

Nathan Pruyne, Benjamin Stoler, William Chen, Chien-yu Huang, Shinji Watanabe, Chris Donahue

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 介绍用于评估自动音乐转录模型的MulTTiPop数据集,通过对Lakh MIDI和TheoryTab数据集歌曲片段基于元数据匹配等方式收集,评估模型发现有改进空间,最佳模型起始F1得分为38%。

Comments 8 pages, 4 figures. Associated web preview available at https://gclef-cmu.org/multtipop

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08319 2026-07-10 cs.DB cs.AI 新提交

GitLake: Git-for-data for the agentic lakehouse

GitLake:面向智能数据湖的数据版Git

Weiming Sheng, Jinlang Wang, Manuel Barros, Aldrin Montana, Jacopo Tagliabue, Luca Bigon

机构 * Columbia University(哥伦比亚大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Carnegie Mellon University(卡内基梅隆大学) Bauplan Labs(Bauplan实验室)

AI总结 研究面向智能体的数据湖的Git设计,核心方法是将单表快照提升为全湖操作,贡献是实现智能体与人类协作,管道原子性输出,还分享了生产经验和正确性见解。

Comments Pre-print of the paper accepted at DASHSys, VLDB 2026, Boston, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07980 2026-07-10 cs.SE cs.AI 新提交

3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse

人工智能时代关于代码审查的3100种观点:从从业者话语中构建因果理论

Shyam Agarwal, Courtney Miller, Christian Kästner, Bogdan Vasilescu

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 研究人工智能对代码审查的影响,通过收集从业者话语构建因果模型,明确审查是控制点,团队决定编码代理对软件影响,转化相关命题,还提供LLM辅助灰色文献理论构建方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07957 2026-07-10 cs.AI cs.CV cs.LG 新提交

Evaluating the Effect of Frame Rate in Sequence-Based Classification of Autism-Related Self-Stimulatory Hand Idiosyncrasies

评估帧率在基于序列的自闭症相关自我刺激手部特质分类中的作用

Raunak Mondal, Peter Washington

机构 * School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院) Department of Information and Computer Sciences, University of Hawai‘i at Mānoa(夏威夷大学马诺亚分校信息与计算机科学系)

AI总结 研究针对自闭症相关自我刺激行为检测,通过在不同帧率下训练LSTM和GRU模型确定最佳架构与采样率,应用十种数据增强策略并做消融研究,采用个性化机器学习方法,为视频行为分类提供架构、采样率及增强策略指导。

Comments 15 pages, 5 figures, 3 tables. Preliminary version presented as a poster at the AMIA 2024 Informatics Summit

Journal ref 2024 AMIA Informatics Summit. https://knowledge.amia.org/Info2024/content?act=Info2024a249

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07840 2026-07-10 cs.PL cs.LG 新提交

GradInf: Gradient Estimation as Probabilistic Inference

GradInf:作为概率推理的梯度估计

Gaurav Arya, Mathieu Huot, Moritz Schauer, Alexander K. Lew, Feras A. Saad

机构 * Carnegie Mellon University Pittsburgh USA Massachusetts Institute of Technology Cambridge USA Chalmers University of Technology \& University of Gothenburg Gothenburg Sweden Yale University New Haven USA Carnegie Mellon University Massachusetts Institute of Technology Chalmers University of Technology \& University of Gothenburg Yale University

AI总结 研究针对概率程序梯度估计难题,提出梯度推理新方法,通过形式规约将其转化为概率推理问题,借助耦合和因式分解操作设计梯度估计器,介绍GradInf系统,经案例研究验证其能表达并构建高性能梯度估计器。

Journal ref Proc. ACM Program. Lang. 10, PLDI, Article 243 (June 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07717 2026-07-10 cs.LG cs.CV 新提交

Who Gets Missed in the Tail? Thresholded Subgroup Underdiagnosis in Long-Tailed Chest X-ray Classification

长尾胸部X光分类中谁在尾部被遗漏?阈值化子组诊断不足

Ha-Hieu Pham, Hai-Dang Nguyen, Dang P. M. Cao, Thanh-Huy Nguyen, Min Xu, Trung-Nghia Le, Ulas Bagci, Huy-Hieu Pham

机构 * University of Science, Ho Chi Minh City(胡志明市科学大学) Vietnam National University, Ho Chi Minh City(胡志明市越南国立大学) VinUni-Illinois Smart Health Center, VinUniversity(VinUni-伊利诺伊智能健康中心,Vin大学) Carnegie Mellon University(卡内基梅隆大学) Northwestern University(西北大学) College of Engineering & Computer Science, VinUniversity(Vin大学工程与计算机科学学院)

AI总结 研究胸部X光分类中长尾数据下子组诊断不足问题,通过诊断阶梯分离相关因素,在VinDr-CXR和MIMIC-CXR/CXR-LT数据集上实验,表明CXR罕见标签公平性取决于发现、子组和阈值,非仅标签频率或排序指标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12145 2026-07-10 cs.LG 版本更新

Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling

阈值差分注意力:用于无汇、超稀疏和非分散语言建模

Xingyue Huang, Xueying Ding, Mingxuan Ju, Yozen Liu, Neil Shah, Tong Zhao

机构 * University of Oxford(牛津大学) Carnegie Mellon University(卡内基梅隆大学) Snap Inc

AI总结 本文提出阈值差分注意力机制,解决长上下文下的注意力汇问题,实现超稀疏和鲁棒性提升,无需投影方法或标准注意力的噪声积累。

Journal ref ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15543 2026-07-10 cs.CL cs.AI 版本更新

ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation

ParamMute:抑制知识关键的前馈神经网络以实现忠实的检索增强生成

Pengcheng Huang, Zhenghao Liu, Yukun Yan, Haiyan Zhao, Xiaoyuan Yi, Hao Chen, Zhiyuan Liu, Maosong Sun, Tong Xiao, Ge Yu, Chenyan Xiong

机构 * School of Computer Science and Engineering, Northeastern University, China(东北大学计算机科学与工程学院) Department of Computer Science and Technology, Institute for AI, Tsinghua University, China(清华大学人工智能研究院计算机科学与技术系) Microsoft Research Asia, Beijing, China(微软亚洲研究院) Language Technologies Institute, Carnegie Mellon University, United States(卡内基梅隆大学语言技术研究所)

AI总结 研究RAG中LLMs不忠实生成问题,提出ParamMute框架抑制相关FFNs激活并校准模型,通过CoFaithfulQA基准评估,显著提升忠实性,减少对参数记忆依赖,为提高LLM在RAG中的可信度提供新方向。

Comments 26 pages, 7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏