Conversational Human Audio-visual Talking Dialogue Generation
对话式人类视听对话生成
机构 * Department of Computing, Imperial College London(伦敦帝国理工学院计算系) ; Department of Earth Science & Engineering, Imperial College London(伦敦帝国理工学院地球科学与工程系) ; Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; Department of Computer Science, Heriot-Watt University(赫瑞瓦特大学计算机科学系) ; School of Computer Science, Northumbria University(诺森比亚大学计算机科学学院) ; College of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件学院) ; School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) ; Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(深圳大学广东省智能信息处理重点实验室) ; School of Engineering and Design, Hunan Normal University(湖南师范大学工程与设计学院) ; Department of Computer Science, University of Oxford(牛津大学计算机科学系) ; Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系)
AI总结 提出CHAT框架,统一大语言模型和说话人脸模型,通过交互音频和面部行为细化模块,从单一文本提示生成多样、配对且相互响应的语音-面部对话片段,优于现有方法,合成数据集可作预训练数据。
Comments Accepted to ECCV 2026 as a main paper