Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning
观看合成视频:针对零样本视频字幕生成的视觉合成跨模态表征对齐
机构 * School of Software, Northwestern Polytechnical University(西北工业大学软件学院) ; School of Information Science and Technology, Beijing University of Technology(北京工业大学信息科学与技术学院) ; School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) ; School of Computer Science, The University of Sydney(悉尼大学计算机科学学院)
AI总结 本文提出WSV零样本视频字幕生成框架,通过文本到视频生成模型、抛光器、提示器弥合跨模态差距,在三个公开数据集上取得了B@4 52、CIDEr 95.7的成绩。