Factorized Learning for Temporally Grounded Video-Language Models
分解学习用于时间感知的视频-语言模型
机构 * National University of Singapore(新加坡国立大学)
专题命中 预训练与数据 :language model(title,abstract);preference optimization(abstract);分类 cs.CL、cs.AI
AI总结 本文提出D$^2$VLM框架,通过分解学习方法提升视频-语言模型在时间定位和文本响应任务中的性能,引入证据标记和FPO算法以优化学习过程。
Comments ICCV 2025 paper. This arXiv version updates Figure 1 to include the concurrent work Qwen2.5-VL to ensure consistency with Table 1