Blog Spotlight
Current hot topic
See All- Insights[Summary] A Memristor-based In-Memory Computing SoC with Efficient Depthwise Convolution
The rapid growth of AI foundation models has intensified the "memory wall" problem, where constant data movement between processors and memory leads to high power consumption, latency, and heat. To address this, the industry is shifting from "compute-centric" to "memory-centric" architectures. Leading this change, SK hynix (G-RTC) and TetraMem have partnered to implement a next-generation Analog Compute-in-Memory (A-CiM) architecture. This collaboration marks a significant milestone in AI infrastructure by directly tackling data movement inefficiencies. The key achievements of this joint research are summarized in three points:
The first key achievement is the world’s first end-to-end hardware demonstration of Depthwise-Separable Convolution (DSC) on a Resistive Switching Cell (RSC)/CMOS-based System-on-Chip (SoC). By performing matrix operations directly within the memory arrays where AI model weights are stored ("Compute-in-Memory"), this work fundamentally removes the need for data movement and overcomes traditional bottlenecks. While DSC has proven difficult for conventional A-CiM architectures, this design successfully accelerates it on silicon. To enable this, the team proposed a novel Zig-Zag Selection-line (SEL) topology. As detailed in Fig. 1, unlike the straight lines in normal crossbars, the SELs in depthwise arrays are routed in a zig-zag pattern to connect specific transistor gates within the 1T1R cells. This unique wiring allows the architecture to activate cells diagonally, boosting weight utilization in depthwise layers to nearly 100% and facilitating efficient sharing of peripheral circuits. Consequently, signal paths for both depthwise and standard operations are maximized for efficiency across the entire network.

Fig. 1. (a) Optical image of a system-on-chip (SoC). Each chip has 10 NPUs, each containing one 256 × 256 crossbar array. (b) Architecture of one depthwise NPU with eight blocks. Each block includes one 252 × 28 depthwise crossbar. Cells enabled by SEL0 are shown, while other cells are masked with blue. (c) Schematic of one depthwise crossbar. SEL0 is highlighted to show the zig–zag shape of SELs. (d) Schematic of normal 1T1R crossbars. (e) 1T1R cell structure showing the direction of WL, BL, and SEL. (f) Cross-sectional TEM image of the RSC cell fabricated. TEC, TE contact; HM, hard mask; TE, top electrode; RSV, reservoir; SWL, switching layer; BE, bottom electrode; BEC, BE contact.
The second major achievement is the demonstration of exceptional energy efficiency alongside validation on real-world AI models. Fabricated using a standard 65nm CMOS process, the chip delivers 0.254 TOPS per NPU and achieves an energy efficiency of 21.3 TOPS/W at 100MHz (Fig. 2a). This performance delivers more than 10 times the energy efficiency of the NVIDIA A100 GPU, reaching levels comparable to state-of-the-art 28nm SRAM-based CiM accelerators despite the older process node. Furthermore, the chip’s practical viability was confirmed through real-world benchmark testing. When executing the lightweight MobileNetV1 model on the Visual Wake Words (VWW) task, the hardware achieved an inference accuracy of 80.36%. As shown in Fig. 2b, this result is slightly higher than the 4-bit software quantization accuracy (79.34%) and remains stable across repeated runs (80.30%~80.39%), confirming negligible run-to-run variation. This proves that the proposed architecture has advanced beyond a theoretical concept to a fully practical, chip-level platform ready for deployment.

Fig.2. (a) VMM-state power and efficiency summary. (b) Software accuracy as the weight precision is swept from 8 → 2 bits, with the measured hardware point overlaid. The hardware accuracy of 80.36% is slightly higher than the 4-bit quantization accuracy of 79.34%.
The third key achievement is the establishment of a comprehensive "mapping and deployment flow" for complex artificial neural networks. Moving beyond a simple chip demonstration, the team developed an end-to-end software stack that orchestrates layer partitioning, heterogeneous mapping, zig-zag kernel placement, calibration-based programming, and RISC-V-based layer scheduling. As illustrated in Fig. 3, this workflow enables the smooth execution of complex CNN models on the hybrid NPU architecture. The figure details how convolutional layers (e.g., DSConv1~5) are optimally partitioned and distributed across multiple NPU units (labeled NPU0–NPU5), maximizing hardware utilization through efficient kernel placement. This capability marks a significant transition from a prototype to a fully functional platform, proving that intricate AI workloads can be deployed reliably. This achievement stems from close engineering collaboration, merging SK hynix’s deep expertise in memory technology with TetraMem’s advanced A-CiM platform capabilities to deliver a complete AI solution.

Fig. 3. (a) The MobileNetV1Small architecture and placement of each layer onto the standard or Depthwise NPU. (b) Number of MAC operations per inference for the DSConv model versus an equivalent full-convolution baseline. (c) End-to-end VWW inference accuracy of the fabricated SoC versus the same software model quantized to 4-bit precision.
In conclusion, this research successfully overcame critical hardware design challenges and demonstrated the seamless execution of real-world AI workloads. These achievements were recognized through publication as a Cover Feature in Advanced Intelligent Systems. The developed SoC establishes a solid foundation for sustainable, next-generation memory-centric AI infrastructure, enabling high-efficiency inference from edge devices to future data centers.
The Publication : Link

- Insights[Insights] Evolving Role of Emerging Memories in Next-Generation Computing
- Insights[Summary] Electrical Characteristics of the 4F2 Vertical Gate (VG) DRAM integrated with Bit-Line Shielding (BLS) and Back Gate (BG) Transistor
- Insights[White Paper] Enabling inference at massive scale with hybrid storage for KV cache offloading
NEWS
Open ExternalResearch & Publications
Recent Publications
See All