MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models
MemoryVLA++:通过记忆与想象在视觉-语言-动作模型中进行时间建模
机构 * Tsinghua University(清华大学) ; The University of Hong Kong(香港大学) ; Dexmal ; StepFun
AI总结 提出MemoryVLA++框架,通过工作记忆、感知-认知记忆库和想象未来状态的世界模型,实现完整时间建模,在模拟和真实机器人任务上显著提升长时域和依赖记忆与想象的任务性能。
Comments The project is available at https://shihao1895.github.io/MemoryVLA-PP-Web