Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding
用于长视频理解中事件感知视觉分配的高斯混合模型
机构 * Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information, CASIA(中国科学院自动化所多模态信息超智能安全北京市重点实验室) ; State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(中国科学院自动化所多模态人工智能系统国家重点实验室) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; Hello Group(未知(保留英文)) ; School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院)
专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV
AI总结 针对长视频理解中视觉分配问题,提出GMM-EVA方法,利用高斯混合模型建模事件级结构,采用差异化分配策略,在多个长视频基准实验中显著优于均匀采样,以约一半视觉令牌预算达可比性能,凸显高效性。
Comments accepted at PRCV 2026