[Summary] PNM Meets Sparse Attention: Enabling Multi-Million Tokens Inference at Scale
PNM Meets Sparse Attention: Enabling Multi-Million Tokens Inference at Scale
Large‑language‑model (LLM) inference is dominated by the quadratic cost of full‑attention, which becomes a severe bottleneck for sequences of millions of tokens. Prior approaches mitigate this by off‑loading weights, compressing activations, or using static sparsity, but they still suffer from high latency and memory pressure.
Key Concepts and System Design
- NELSSA Architecture: A dynamic sparse‑attention accelerator that selects a small subset of tokens per decoding step via a high‑recall vector‑search (token‑selection) stage, then computes attention only on the selected tokens.
- Token‑selection Ratio: The design operates with a 2 % selection ratio (shown to retain high accuracy), but can adapt up to 10 % with modest latency impact.
- Hardware Configuration: Implemented on a CXL‑based near‑memory processing (PNM) platform: 4 × H100 GPUs plus four PNM modules (either stacked‑LPDDR5 or 4‑channel DDR5), interconnected via PCIe 5.0 × 16.
- Baseline Comparisons: Systems evaluated include Hermes (hot/cold neuron partitioning), FlexGen, InfiniGen, RetrievalAttention (CPU‑based sparse attention), and a full‑attention CXL‑PNM configuration.

Figure 1: NELSSA Architecture and Workflow.
Performance Evaluation
- Throughput: NELSSA delivers 11 × to 40 × higher decode throughput than the primary baseline Hermes (Fig. 2). The advantage grows with sequence length, enabling inference on multi‑million‑token prompts that are infeasible for baseline systems (OOM errors shown as ‘X’ markers).
- Latency Breakdown: The token‑selection stage contributes negligible overhead; the majority of latency savings stem from eliminating the full‑attention computation (Fig. 6, left axis). Even when the selection ratio rises to 10 %, per‑token decode latency increases by only 2–3 × relative to the 2 % baseline, still well below baseline systems (Fig. 3, right axis).
- Robustness: NELSSA remains stable across a range of selection ratios, confirming its suitability as a versatile platform for various dynamic sparse‑attention algorithms.

Figure 2: Relative decode throughput comparison for NELSSA and baseline systems.

Figure 3: Latency breakdown and sensitivity analysis of a NELSSA decode step.
Conclusion
The NELSSA architecture demonstrates that dynamic sparse attention combined with near‑memory processing can dramatically accelerate LLM inference, achieving order‑of‑magnitude throughput gains while maintaining low latency and modest accuracy loss. Its low token‑selection overhead and resilience to different sparsity‑accuracy trade‑offs make it a promising foundation for future long‑context LLM deployments.
The Publication : Link

![[Summary] A Memristor-based In-Memory Computing SoC with Efficient Depthwise Convolution](https://mis-prod-koce-research-user-cdn-01-blob-ep.azureedge.net/web/blog/20260806/thumb_sEVmfv3K.20260806085519833.jpg)
![[Insights] Evolving Role of Emerging Memories in Next-Generation Computing](https://mis-prod-koce-research-user-cdn-01-blob-ep.azureedge.net/web/blog/20260806/thumb_NgJoNLYn.20260806080458980.jpg)
![[Summary] Electrical Characteristics of the 4F2 Vertical Gate (VG) DRAM integrated with Bit-Line Shielding (BLS) and Back Gate (BG) Transistor](https://mis-prod-koce-research-user-cdn-01-blob-ep.azureedge.net/web/blog/20260806/thumb_DOUurl5m.20260806080624234.png)