Agent Memory Below the Prompt: Persistent Q4 KV Cache for Multi-Agent LLM Inference on Edge Devices
代理记忆低于提示:持久化Q4 KV缓存用于边缘设备上的多代理LLM推理
专题命中 记忆与上下文管理 :agent(title,abstract);multi-agent(title,abstract);workflow(abstract);分类 cs.AI、cs.LG
AI总结 本文提出了一种持久化Q4 KV缓存的方法,用于在边缘设备上提高多代理LLM推理效率,通过减少缓存恢复时间,提升模型性能。
Comments 24 pages, 6 figures, 16 tables. Open-source implementation at https://github.com/yshk-mxim/agent-memory