CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
CaptionFormer:时空对象的统一分割、跟踪与描述
机构 * Inria, École Normale Supérieure, CNRS, PSL Research University(法国国家科学研究中心、巴黎高等师范学院、国家科学研究中心、巴黎综合理工研究所) ; Google DeepMind(谷歌DeepMind)
AI总结 提出 CaptionFormer 模型,通过利用 VLM 生成合成描述并扩展数据集,实现视频中对象轨迹的联合检测、分割、跟踪与描述,在三个基准上达到最优。
Comments 17 pages, 10 figures