Dynamic Token Compression for Efficient Video Understanding through Reinforcement Learning
通过强化学习实现高效的视频理解动态令牌压缩
机构 * University of Science and Technology of China(中国科学技术大学) ; State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)
AI总结 本文提出SCORE框架,通过强化学习学习自适应令牌压缩策略,有效解决视频理解中的计算成本和上下文旋转问题,实现16倍的预填充加速且性能损失仅0.5%。