MKA: Memory-Keyed Attention for Efficient Long-Context Reasoning
MKA:用于高效长上下文推理的内存键注意
机构 * University of California, Los Angeles(加州大学洛杉矶分校) ; Columbia University(哥伦比亚大学) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
AI总结 本文提出MKA机制,通过多级KV缓存和动态路由提升长上下文推理效率,FastMKA在保持准确率的同时显著提升训练速度和评估效率。
Comments Accepted to the ACM Computing Frontiers 2026 Conference (Oral Presentation) and the ICML 2025 Long Context Modeling Workshop