[Summary] Analog Computation in Ultra-High Density 3D FeNAND for TB-level Hyperscale AI Models
Large language models have become prominent in generating and comprehending human-like text. Their model sizes have been dramatically increasing, even reaching trillions of parameters. The exponential growth in model size needs to have substantial memory resources to handle the vast amounts of data involved in AI inference process. In this regard, 3D vertical NAND string structures (Fig. 1 (a)), known for its high areal density and non-volatile, emerge as a potential candidate for energy efficient analog compute-in-memory (ACiM). These vertical NAND structures especially provide a density advantage by utilizing the same peripheral circuits with digital, ADC, and MUX even as the number of cell layers increases, avoiding additional area requirement (Fig. 1 (b)).

Fig. 1. (a) Schematic illustration of 3D FeNAND array for ACiM, (c) Comparison of cell density of 2D array vs 3D array
In this work, we demonstrated an analog computation in high-density 3D vertical ferroelectric NAND (FeNAND) for the first time. In the vertical string structure of 3D NAND arrays, as the number of word line (WL) stacks increases, the available conductance ranges for analog states decrease (Fig. 2 (a)). Higher subthreshold slope (SS) of FeNAND cells can induce more conductance states with lower input variation then low SS, improving the analog properties of FeNAND cells. In order to make more conductance states, the interface state density within the FeNAND gate stacks was engineered (Fig. 2 (b)). As a results, the SS of 3D FeNAND S2 was up to twice as large as S1 at maximum (Fig. 2 (c)).

Fig. 2. (a) Schematic illustration of 3D FeNAND string and available current range for multi-leveling. (b) Gate stack of 3D FeNAND with interface engineering. (c) Subthreshold slope as a function of ∆Vth on Split 1 and 2.
The ID-VG curves were measured as a function of program bias amplitude with fixed pulse width for both S1 and S2 of 3D FeNAND (Fig. 3 (a)). The conductance states of 3D FeNAND S1 and 3D FeNAND S2 were confirmed to have ≥128 and ≥256 levels/cell, respectively (Fig. 3 (b)). Furthermore, we verified MAC operations of 3D FeNAND arrays by using Ohm’s law and Kirchhoff's law (MAC currents [IMAC] = Σ[G×V]) (Fig. 3 (c)). The 3D FeNAND arrays with R1 exhibited higher MAC accuracy (87.8 ± 8.2%) than those with R2 (78.4 ± 6.1%) because the R1 conditions are more resilient to IR drop effects than the R2 conditions.

Fig. 3. (a) Measured ID-VD curves with various PGM bias and (b) Conductance as a function of pulse numbers in 3D FeNAND S1 and S2. (c) Experimental data of MAC operations and MAC accuracy in 3D FeNAND.
The simulation revealed that the inference accuracy of all samples (85.7% for 3D FeNAND S1, 85.6% for 3D FeNAND S2, and 85.4% for 2D ReRAM) was similar to the software level inference accuracy (90.3% for FP32 and 89.4% for 4-bit integer post-training-quantization) (Fig. 4 (a)). With the assumption that the weight parameters were all written in cell arrays to reduce data movements, 2D arrays require over 4,000 times more area than 3D arrays for inference (Fig. 4 (b)). The 3D FeNAND arrays exhibited 1000-fold higher compute efficiency (TOPS/mm²) compared to the 2D ReRAM arrays (Fig. 4 (c)). We also calculated the power consumption of the samples for inference (Fig. 4 (d)). The power consumption of 3D FeNAND S2 was observed to be lower than that of 3D FeNAND S1 and, notably even lower than that of the 2D array ReRAM array. The reduction is attributed to improved multi-leveling properties.

Fig. 4. Performance simulation; (a) Inference accuracy, (b) total cell array area, (c) Compute efficiency, (d) Power consumption.
In summary, we achieved the multi-level weight conductance states (≥ 256 levels/cell) of 3D FeNAND cells by engineering the interface trap density of gate stacks. The developed 3D FeNAND arrays demonstrated stable MAC operations with high accuracy (87.8%) and exhibited ultra-high compute efficiency, achieving a factor of 1,000x compared to 2D arrays. Our study provides an efficient method for developing ultra-high-density ACiM chips for edge computing with hyperscale AI models.
The Publication : Link

![[Summary] A Memristor-based In-Memory Computing SoC with Efficient Depthwise Convolution](https://mis-prod-koce-research-user-cdn-01-blob-ep.azureedge.net/web/blog/20260806/thumb_sEVmfv3K.20260806085519833.jpg)
![[Insights] Evolving Role of Emerging Memories in Next-Generation Computing](https://mis-prod-koce-research-user-cdn-01-blob-ep.azureedge.net/web/blog/20260806/thumb_NgJoNLYn.20260806080458980.jpg)
![[Summary] Electrical Characteristics of the 4F2 Vertical Gate (VG) DRAM integrated with Bit-Line Shielding (BLS) and Back Gate (BG) Transistor](https://mis-prod-koce-research-user-cdn-01-blob-ep.azureedge.net/web/blog/20260806/thumb_DOUurl5m.20260806080624234.png)