H3: Hybrid Architecture Using High Bandwidth Memory and High Bandwidth Flash for Cost-Efficient LLM Inference
Large language model (LLM) inference requires massive memory capacity to process long sequences, posing a challenge due to the capacity limitations of high bandwidth memory (HBM). High bandwidth flash (HBF) is an emerging memory device based on NAND flash that offers HBM-comparable bandwidth with much larger capacity, but suffers from disadvantages such as longer access latency, lower write endurance, and higher power consumption. This paper proposes H3, a hybrid architecture designed to effectively utilize both HBM and HBF by leveraging their respective strengths. By storing read-only data in HBF and other data in HBM, H3-equipped systems can process more requests at once with the same number of GPUs than HBM-only systems, making H3 suitable for gigantic read-only use cases inLLMinference, particularly those employing a shared pre-computed key-value cache. Simulation results show that a GPU system with H3 achieves up to 2.69x higher throughput per power compared to a system with HBM-only. This result validates the cost-effectiveness of H3 for handling LLM inference with gigantic read-only data.
Co-Optimizing Cell, Non-Cell, and Page Schemes for Energy Efficient Analog Computing in 3D FeNAND
This work presents a holistic co-optimization of 3D ferroelectric NAND (FeNAND) for energy-efficient and high-throughput analog computation-in-memory. The co-optimization covers cell properties, non-cell peripherals, and computational schemes. Our approach enhances multi-level capability through write algorithm tuning, reduces word-line transition power, and improves analog multiply-and-accumulate efficiency via in-page pipelining and splitting. As a result, computational throughput and energy efficiency of 3D FeNAND are improved by up to 16× and 4950×, respectively, compared to conventional 2D arrays.
A Highly Scalable Isolated Charge Trap Nitride Layer Implemented in a 176-layer 3D NAND Flash with Superior Threshold Voltage Distribution and Charge Retention
In this work, we present the successful implementation of a novel charge trap nitride isolation (CTI) technology in a 176-layer production-scale 3D NAND device. In this approach, the charge trap nitride (CTN) layer is confined by a pocket with sidewalls formed through a deposition process, a first-time implementation that improves CTI scalability. We achieved sufficient process maturity to produce fully functional chips across a substantial wafer area, enabling chip-level evaluation of threshold voltage (Vth) distribution and retention. This technology achieves an 11% reduction in read time (tR), a 6.9% narrower Vth distribution width, and a 45% improvement in high-temperature retention after cycling compared to conventional continuous CTN cells. These results suggest the current tier pitch limit could be reduced by 4 nm in terms of Vth distribution and by more than 10 nm from a retention perspective, supporting our CTI technology as a compelling solution for continued 3D NAND scaling.
cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
Message Passing Interface (MPI) is a foundational programming model for high-performance computing. MPI libraries traditionally employ network interconnects (e.g., Ethernet and InfiniBand) and network protocols (e.g., TCP and RoCE) with complex software stacks for cross-node communication. This paper presents cMPI, the first work to optimize MPI point-to-point communication (both one-sided and two-sided) using CXL memory sharing on a real CXL platform, transforming cross-node communication into memory transactions and data copies within CXL memory, bypassing traditional network protocols. We analyze performance across various interconnects and find that CXL memory sharing achieves 7.2×- 8.1× lower latency than TCP-based interconnects deployed in smalland medium-scale clusters. We address challenges of CXL memory sharing for MPI communication, including data object management over the dax representation [50], cache coherence, and atomic operations. Overall, cMPI outperforms TCP over standard Ethernet NIC and high-end SmartNIC by up to 49× and 72× in latency and bandwidth, respectively, for small messages.
Integrating Distributed SQL Query Engines with Object-Based Computational Storage
Existing object storage systems like AWS S3 and MinIO offer only limited in-storage compute capabilities, typically restricted to simple SQL WHERE-clause filtering. Consequently, high-impact operators such as aggregation and top-N are still executed entirely at the compute layer. Recent advances in Object-based Computational Storage (OCS) enable these complex operators to run natively within storage, creating opportunities for substantial reductions in data movement and query time. To demonstrate these benefits in distributed SQL engines, we used Presto as a case study and developed the Presto-OCS connector, which analyzes execution plans to identify pushdown-eligible operators and offloads them to OCS for efficient in-storage execution. Evaluations with real-world HPC analytics queries and the TPC-H benchmark show that our approach achieves up to 4.07× speedup and 99% data movement reduction compared to filter-only pushdown. When combined with compression techniques, our approach delivers 1.39× speedup over compressed filter-only pushdown, demonstrating that advanced query pushdown complements existing optimizations.
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
The escalating context length in Large Language Models (LLMs) creates a severe performance bottleneck around the Key-Value (KV) cache, whose memory-bound nature leads to significant GPU under-utilization. This paper introduces Mixture of Shared KV Attention (MoSKA), an architecture that addresses this challenge by exploiting the heterogeneity of context data. It differentiates between per-request unique and massively reused shared sequences. The core of MoSKA is a novel Shared KV Attention mechanism that transforms the attention on shared data from a series of memory-bound GEMV operations into a single, compute-bound GEMM by batching concurrent requests. This is supported by an MoE-inspired sparse attention strategy that prunes the search space and a tailored Disaggregated Infrastructure that specializes hardware for unique and shared data. This comprehensive approach demonstrates a throughput increase of up to 538.7 × over baselines in workloads with high context sharing, offering a clear architectural path toward scalable LLM inference.
PNM Meets Sparse Attention: Enabling Multi-Million Tokens Inference at Scale
Processing multi-million tokens for advanced Large Language Models (LLMs) poses a significant memory bottleneck for existing AI systems. This bottleneck stems from a fundamental resource imbalance, where enormous memory capacity and bandwidth are required, yet the computational load is minimal. We propose NELSSA (Processing Near Memory for Extremely Long Sequences with Sparse Attention), an architectural platform that synergistically combines the high-capacity Processing Near Memory (PNM) with the principles of dynamic sparse attention to address this issue. This approach enables capacity scaling without performance degradation, and our evaluation shows that NELSSA can process up to 20M-token sequences on a single node (Llama-2-70B), achieving an 11× to 40× speedup over a representative DIMM-based PNM system. The proposed architecture radically resolves existing inefficiencies, enabling previously impractical multi-million-token processing and thus laying the foundation for next-generation AI applications.
224 Tops/W-Level Analog Computation in Memory Cell Using Hybrid Ferroelectric Tunnel Junction Having Enhanced on-State Conductance
Here, we demonstrate hybrid ferroelectric tunnel junctions (H-FTJs) that combine ferroelectric and resistive switching for analog computation in memory. The key of H-FTJs lies in the modulation of the effective tunneling thickness by the control of oxygen vacancy-based unconnected filaments in HfZrO2. The H-FTJs exhibited high on-state conductance (1.6×103S/cm2) and on/off ratios (32,000) compared to recent studies on HfZrO2-based FTJs. We also fully integrated one-transistor-one-(H-FTJ) cross-bar arrays, and confirmed their analog multiply-accumulate operations (accuracy 91.9%) with high energy efficiency (224.4 TOPS/W) for inference tasks.
4F2DRAM Integration with Vertical Gate (VG) Cell Transistor and Peri-Under-Cell (PUC) Architecture
Process integration of 4F2 DRAM array with peri-under-cell (PUC) architecture has been successfully demonstrated for the first time, employing vertical gate (VG) cell transistor and wafer bonding process. Fusion wafer bonding and inter-wafer contact techniques enabled wafer-to-wafer integration to provide robust electrical connections between cell and peripheral circuit wafers. Compared to conventional 2D DRAM, superior control over threshold voltage is achieved from back gate bias in VG cell transistor. Precise junction engineering by thermal annealing optimizes cell transistor performance.