Blog

Stay updated with our latest news and announcements.

Insights

[Summary] MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference

Myunghyun RheeMyunghyun Rhee (in IEEE Computer Architecture Letters)

MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference

 

Breaking the Memory Barrier for Efficient Long-Sequence AI Models

The era of massive context windows in Large Language Models (LLMs) has arrived. While this unlocks incredible applications, it also exposes a severe bottleneck for AI service infrastructure: the Key-Value (KV) cache. Even with modern optimization techniques, memory bandwidth requirements scale linearly with batch size, leading to significant GPU under-utilization. Simply adding more memory capacity isn't enough to solve the fundamental problem.

 

Figure 1: Hardware Requirement Challenges.

 

To tackle this challenge, we are excited to introduce MoSKA (Mixture of Shared KV Attention), an architecture that leans into the principles of memory-centric AI. MoSKA strategically exploits the difference between per-request unique data and massively reused shared data. 

 

How MoSKA Works: Rethinking Attention

The secret sauce is the Shared KV Attention mechanism. Standard attention on unique, per-request data remains a memory-bound operation. However, when multiple requests access the same shared data, MoSKA batches these concurrent queries into a single, massive compute-bound GEMM operation.

 

To manage massive shared contexts efficiently, MoSKA incorporates a lightweight routing layer inspired by Mixture-of-Experts (MoE). Similar to concepts in embedding augmented generation, a training-free router dynamically calculates relevance scores to select only the most pertinent shared KV chunks. This sparse attention strategy drastically prunes the search space, ensuring we only compute what is absolutely necessary.

 

Figure 2: The MoSKA Architecture.

 

Disaggregated AI Infrastructure

To fully realize these benefits, MoSKA proposes an attention-centric Disaggregated Infrastructure. The system separates hardware based on distinct computational profiles:

  • Unique KV Nodes: Optimized for latency-sensitive, memory-bound unique sequences.
  • Shared KV Nodes: Designed for throughput-oriented, compute-bound Shared KV Attention tasks.

This separation prevents resource contention and allows the system to independently scale its shared knowledge processing capacity without over-provisioning expensive hardware.

 

Figure 3: Proposed disaggregated LLM serving infrastructure for MoSKA.

 

The Results: A Massive Leap Forward

This architectural shift delivers remarkable results. In workloads with high context sharing, MoSKA demonstrates a staggering throughput increase of up to 538.7x compared to state-of-the-art baselines. By transforming the workload into a compute-bound task, the Shared Node achieves an impressive Model FLOPS Utilization (MFU) of over 80% even with a massive 16M token shared context.

 

Looking Ahead: Universal MoSKA

MoSKA offers a clear, architectural path toward truly scalable LLM inference. The long-term vision points toward "Universal MoSKA," where shared KV chunks become modular, composable blocks of knowledge. This will pave the way for a future where complex queries can dynamically assemble context from a distributed network on demand. 

 

Conclusion

MoSKA represents a pivotal step forward in overcoming the memory bottlenecks of long-sequence AI by turning a previously memory-bound challenge into a highly efficient compute-bound solution. As we continue to develop this architecture, we are excited to move close to a future where AI models can dynamically compose and access vast amounts of shared knowledge on the fly.


The Publication : Link



Popular Insights

Previous Next List