Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
MLLMs中分解注意力融合用于无训练视频推理分割
机构 * Yonsei University(延世大学) ; Inha University(inha大学) ; NAVER Cloud(NAVER云)
专题命中 视频理解 :video reasoning(title,abstract);video understanding(abstract);分类 cs.CV
AI总结 本文提出DecAF方法,通过对比物体背景融合和互补视频帧融合精炼注意力图,实现无训练视频推理分割,取得优于无训练方法和与有训练方法相当的性能。
Comments Accepted to ICLR 2026. Code is available at https://github.com/HYUNJS/DecAF