Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
通过一次剪辑检索增强多模态大语言模型的长视频理解
机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China(多媒体可信感知与高效计算教育部重点实验室) ; Xiamen University(厦门大学)
专题命中 长视频与时序推理 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV
AI总结 本文提出OneClip-RAG方法,通过一次剪辑检索提升多模态大语言模型对长视频的理解能力,改进知识完整性和语义连贯性,并设计新数据集和训练方案。