ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment
ImmersiveTTS:基于多模态扩散Transformer和领域特定表示对齐的环境感知文本转语音
机构 * Department of Artificial Intelligence, Korea University(韩国大学人工智能系)
AI总结 提出ImmersiveTTS模型,通过多模态扩散Transformer和领域特定表示对齐,实现与环境音频自然融合的文本到语音生成。
Comments Accepted to ACL 2026 main conference. Code is available at https://github.com/jjunak-yun/ImmersiveTTS