Blog

Stay updated with our latest news and announcements.

Insights

[Summary] H3: Hybrid Architecture Using High Bandwidth Memory and High Bandwidth Flash for Cost-Efficient LLM Inference

Minho HaMinho Ha (in IEEE Computer Architecture Letters)

H3: Hybrid Architecture Using High Bandwidth Memory and High Bandwidth Flash for Cost-Efficient LLM Inference

 

Key Contributions

  1. Proposes H3: A hybrid memory architecture integrating HBM and HBF for cost-efficient LLM inference.
  2. Identifies Ideal Use Case: Gigantic read-only workloads (e.g., cache-augmented generation with precomputed KV caches).
  3. Demonstrates Cost-Effectiveness: Simulations show significant throughput and energy efficiency gains, even with HBF’s limitations.

 

Architecture Overview

  1. Memory Mapping: HBF stores read-only data (weights, precomputed KV cache); HBM stores dynamic, frequently updated data.
  2. Interconnect: HBM and HBF are daisy-chained via the HBM base die, allowing unified address space and direct GPU access.
  3. Latency Mitigation: data is prefetched to latency hiding buffer based on predictable LLM layer-by-layer computation patterns.

 

Evaluation Results

  1. Batch Size: H3 supports up to 2.6x larger batches (1M sequence) and 18.8x larger batches (10M sequence) vs. HBM-only.
  2. Throughput: Up to 6.14x higher TPS for 10M sequences.
  3. Throughput per Power: Up to 2.69x improvement, despite HBF’s higher power draw.
  4. Robustness: Even with 50% reduced HBF bandwidth, H3 still outperforms HBM-only systems.

 

Conclusion

H3 provides a practical, cost-efficient solution for scaling LLM inference with long sequences by intelligently combining HBM and HBF. It exploits the read-only nature of key LLM data to mitigate HBF’s weaknesses, making it viable for next-generation AI infrastructure. Future work will explore alternative hybrid architectures and broader HBF applications.


The Publication : Link



Popular Insights

Previous Next List