OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
OmniVTG:一种大规模数据集和开放世界视频时间定位的训练范式
机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) ; State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) ; Central Media Technology Institute, Huawei Technologies Ltd.(华为技术有限公司中央媒体技术研究所) ; PKU-WUHAN Institute for Artificial Intelligence, Peking University(北京大学武汉人工智能研究所)
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
AI总结 本文提出OmniVTG数据集和Self-Correction Chain-of-Thought训练范式,通过语义覆盖迭代扩展管道构建大规模数据集,并利用多模态大语言模型的密集描述能力提升视频时间定位性能。
Comments CVPR 2026