Publications

Explore our latest research papers, articles, and whitepapers.

Nex Gen. Memory & Storage Solution

H3: Hybrid Architecture Using High Bandwidth Memory and High Bandwidth Flash for Cost-Efficient LLM Inference

in IEEE Computer Architecture Letters
Large language model (LLM) inference requires massive memory capacity to process long sequences, posing a challenge due to the capacity limitations of high bandwidth memory (HBM). High bandwidth flash (HBF) is an emerging memory device based on NAND flash that offers HBM-comparable bandwidth with much larger capacity, but suffers from disadvantages such as longer access latency, lower write endurance, and higher power consumption. This paper proposes H3, a hybrid architecture designed to effectively utilize both HBM and HBF by leveraging their respective strengths. By storing read-only data in HBF and other data in HBM, H3-equipped systems can process more requests at once with the same number of GPUs than HBM-only systems, making H3 suitable for gigantic read-only use cases inLLMinference, particularly those employing a shared pre-computed key-value cache. Simulation results show that a GPU system with H3 achieves up to 2.69x higher throughput per power compared to a system with HBM-only. This result validates the cost-effectiveness of H3 for handling LLM inference with gigantic read-only data.

DOI

https://doi.org/10.1109/LCA.2026.3660969
PreviousList