Blog

Stay updated with our latest news and announcements.

Insights

[Summary] Co-Optimizing Cell, Non-Cell, and Page Schemes for Energy Efficient Analog Computing in 3D FeNAND

Wontae Koo, Jihun KimWontae Koo, Jihun Kim (in IEDM 2025)



The rapid growth of large‑scale AI models is driving the development of AI‑centric computing systems, but today’s AI accelerators consume a lot of energy because data must be moved repeatedly between the processor and memory. This creates a strong demand for next‑generation platforms that combine high energy efficiency with high throughput. Analog computation‑in‑memory (A‑CiM) is emerging as a promising solution because of its low power consumption, yet many implementations cannot provide the ultra‑high density required for large AI models. In this regard, we recently proposed a 3D ferroelectric NAND (FeNAND)‑based A‑CiM that achieves both lower power consumption and ultra‑high density. Nevertheless, there are some challenges: (1) large string resistance, (2) high power consumption during word‑line (WL) transitions, and (3) inefficient parallelism due to page‑level sequential MAC operations.


In this work, we co‑optimize 3D FeNAND to improve both energy efficiency (TOPS/W) and throughput (TOPS) for A‑CiM applications (Fig. 1). Our approach includes the optimization of cell properties, non-cell components, and computational schemes, thereby boosting multi‑level capability, lowering word‑line transition power, and enhancing MAC efficiency.

 

Fig. 1. (Left) Schematic illustration of 3D FeNAND arrays for A‑CiM applications. (Right) Challenges of 3D FeNAND arrays for A‑CiM applications and their corresponding co‑optimizations.
 

First, we improved cell properties by optimizing the write algorithm. As the number of WL stacks increases, the current range available for analog states shrinks because higher string resistance induces a larger IR drop (Fig. 2a). In order to mitigate this effect, we introduced pulse‑number modulation in the 3D FeNAND. The 3D FeNAND operating with the V2 scheme showed better multi‑level characteristics under IR‑drop‑aware low‑current conditions (Fig. 2b). Consequently, the number of usable multi‑level states per WL stack is increased through pulse‑number modulation (Fig. 2c). 

 

Fig. 2. (a) Schematic illustration of a 3D FeNAND string. (b) Currents of 3D FeNAND cells as a function of programming pulses. (c) Available multi‑level capability of the 3D FeNAND.
 

Second, we analyzed the impact of the WL RC on the performance of 3D FeNAND. In particular, the transition between selected and unselected WL during MAC operations markedly increases both power consumption and latency (Fig. 3a). Therefore, reducing the WL RC can markedly improve performance, achieving up to 1.4 × higher TOPS and 2.2 × higher TOPS/W (Fig. 3b). 

 

Fig. 3. (a) Schematic illustration of parasitic capacitance and resistance in 3D FeNAND arrays. (b) TOPS and TOPS/W of 3D FeNAND arrays as a function of WL RC. 

 

Third, we optimized the computational schemes in 3D FeNAND. To simultaneously reduce the number of forward operations and WL transitions, we optimized the in‑page pipelining (IPP), which combines neural‑network duplication with the mapping of multiple neural‑network layers (Fig. 4 [Left]). In addition, we investigated the effect of page‑splitting on the A‑CiM performance of 3D FeNAND. Proper page‑splitting effectively improves A‑CiM performance by reducing the number of pages allocated for a given workload (Fig. 4 [Right]). 

  

Fig. 4. (Left) Schematic illustration of the IPP method in 3D FeNAND arrays: (S1) reference and (S2) IPP.
     (Right) Schematic illustration of page‑splitting in 3D FeNAND arrays, with page sizes of (P1) 16 KB and (P2) 4 KB.

 

Lastly, we evaluated the overall improvement of 3D FeNAND achieved by our co‑optimization method (Fig. 5a). The co‑optimization of 3D FeNAND arrays showed 20.4‑fold higher TOPS and 7.17‑fold higher TOPS/W, compared to the previous 3D FeNAND arrays. We also benchmarked the performance of various emerging memory‑based A‑CiM for large‑scale AI models (Fig. 5b). With the limited chip size, 2D memory arrays must frequently load weight parameters because of their low capacity, even for inference‑only workloads. In contrast, the ultra‑high‑density 3D FeNAND arrays eliminate the need for weight updates, achieving up to 16× higher TOPS and 4,950× higher TOPS/W compared with 2D arrays.

 

Fig. 5. A‑CiM performance of (a) co‑optimized 3D FeNAND and (b) emerging memory‑based A‑CiM.


The Publication : Link





Popular Insights

Previous Next List