JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation
JAVEDIT: 联合音频-视觉指令引导视频编辑与智能体数据策展
机构 * Zhejiang University(浙江大学) ; Tencent Youtu Lab(腾讯优图实验室) ; Nanjing University(南京大学) ; University of Auckland(奥克兰大学) ; Fudan University(复旦大学) ; National University of Singapore(新加坡国立大学)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV
AI总结 针对联合音频-视觉编辑缺乏数据集和基准的问题,提出首个大规模高质量数据集JAVEdit-100k、基准JAVEditBench以及基线模型JAVEdit,在六项指标中五项超越所有基线。
Comments Equal contributions from first two authors. Project page: https://ryanchenyn.github.io/projects/JAVEdit Code: https://github.com/RyanChenYN/JAVEdit Dataset: https://huggingface.co/datasets/Coraxor/JAVEdit-100k