arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Johns Hopkins University(约翰斯·霍普金斯大学)

2026-06-09 至 2026-06-09 共收录 4
2603.08977 2026-06-09 eess.AS cs.SD 版本更新

Universal Speech Content Factorization

通用语音内容分解

Henry Li Xinyuan, Zexin Cai, Lin Zhang, Leibny Paola García-Perera, Berrak Sisman, Sanjeev Khudanpur, Nicholas Andrews, Matthew Wiesner

机构 * Center for Language and Speech Processing, Johns Hopkins University, USA(约翰霍普金斯大学语言与语音处理中心) Human Language Technology Center of Excellence (COE), Johns Hopkins University, USA(约翰霍普金斯大学人类语言技术卓越中心(COE))

AI总结 本文提出USCF方法,通过线性可逆方法提取低秩语音表示,抑制说话人音色同时保留语音内容。该方法扩展了语音内容分解,通过最小二乘优化学习通用语音到内容映射,并从少量目标语音中推导出说话人特定转换。

Comments Accepted to Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22919 2026-06-09 cs.CV 版本更新

Chain of Flow: ECG-Conditioned 4D Cardiac Cine Generation from Patient-Specific Anatomical Anchor

流动链:基于患者特定解剖锚点的ECG条件4D心脏电影生成

Haofan Wu, Nay Aung, Theodoros N. Arvanitis, Joao A. C. Lima, Steffen E. Petersen, Le Zhang

机构 * School of Engineering, College of Engineering and Physical Sciences, University of Birmingham(英国伯明翰大学工程学院) William Harvey Research Institute, NIHR Barts Biomedical Research Centre, Queen Mary University London(伦敦Queen Mary大学威廉·哈里维研究所) Barts Heart Centre, St Bartholomew’s Hospital, Barts Health NHS Trust(巴特勒医院心脏中心,圣巴塞洛缪医院,巴特勒健康 NHS信托) Division of Cardiology, Johns Hopkins University School of Medicine(约翰霍普金斯大学医学院心脏病科)

AI总结 提出Chain of Flow (COF)框架,利用患者特定MRI和当前ECG生成4D心脏电影,在UK Biobank上实现高图像保真度和下游功能性能。

Comments 10 pages, 8 figures. Submitted to IEEE Transactions on Medical Imaging (TMI). Code will be released after review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22736 2026-06-09 cs.LG cs.AI 版本更新

UA-DCM: Uncertainty-aware Causal Decision Making via Effect Bound Decomposition

UA-DCM: 基于效应界分解的不确定性感知因果决策

Md Musfiqur Rahman, Ziwei Jiang, Hilaf Hasson, Murat Kocaoglu

机构 * Electrical and Computer Engineering, Purdue University(帕克大学电气与计算机工程系) Computer Science, Johns Hopkins University(约翰霍普金斯大学计算机科学系) Cohesity

AI总结 提出一种新框架,通过分解因果效应值的可消除与不可消除部分,区分收集更多样本能否帮助识别最优行动,并利用神经因果模型近似实现该分解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08238 2026-06-09 cs.SD eess.AS 版本更新

CodecFake+: Codec-Based Resynthesized Data as a Proxy for Detecting CodecFake Speech

CodecFake+: 基于编解码器的重合成数据作为检测CodecFake语音的代理

Xuanjun Chen, Jiawei Du, Haibin Wu, Lin Zhang, I-Ming Lin, I-Hsiang Chiu, Wenze Ren, Yuan Tseng, Yu Tsao, Jyh-Shing Roger Jang, Hung-yi Lee

机构 * Graduate Institute of Communication Engineering, National Taiwan University(国家交通大学通信工程研究院) Department of Computer Science and Information Engineering, National Taiwan University(国家交通大学计算机科学与信息工程系) Center for Language and Speech Processing at Johns Hopkins University(约翰霍普金斯大学语言与语音处理中心) Department of Electrical Engineering, National Taiwan University(国家交通大学电子工程系) Research Center for Information Technology Innovation, Academia Sinica(学术院信息技术创新研究中心) NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)(国家交通大学人工智能研究中心)

AI总结 针对新兴的CodecFake深度伪造语音检测挑战,提出大规模数据集CodecFake+,包含31种开源编解码器重合成训练数据和17种先进CoSG模型网络数据,并建立编解码器分类体系,验证了重合成语音作为训练数据的有效性。

Comments Accepted by TASLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏