Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
Molmo2:具有视频理解和 grounding 的开放权重和数据的视觉语言模型
机构 * Allen Institute for AI(艾伦人工智能研究所) ; University of Washington(华盛顿大学)
AI总结 Molmo2 是一种开放源代码的视频语言模型,通过7个新视频数据集和2个多图像数据集,实现了单图像、多图像和视频任务中的点驱动 grounding 能力,其8B模型在短视频计数和描述任务中表现优异,在视频 grounding 任务中超越了现有开源和专有模型。
Comments Updated first authors