From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
从逐字到概要:基于语义信息瓶颈的金字塔多模态记忆压缩用于长视界视频智能体
机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) ; Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳校区) ; Peng Cheng Laboratory(鹏城实验室)
专题命中 逻辑推理 :reasoning(abstract);分类 cs.CL、cs.AI
AI总结 本文提出MM-Mem架构,通过金字塔多模态记忆压缩,结合语义信息瓶颈目标,实现长视界视频理解中的高效记忆组织与任务相关信息保留。
Comments Accepted by ACL 2026 Main. 17 pages, 7 figures, 8 tables. TL;DR: We propose MM-Mem, a cognition-inspired, dual-trace hierarchical memory framework for long-horizon video understanding grounded in Fuzzy-Trace Theory. It features adaptive memory compression via the Information Bottleneck and employs an entropy-driven top-down retrieval to access fine-grained details only when necessary