Blog

Stay updated with our latest news and announcements.

Insights

[Invited Insights] Challenges of ACiM as a key player in sustainable progress for machine learning

Doo Seok JeongDoo Seok Jeong, Professor , Hanyang University

In the past decade, the level of artificial intelligence (AI) technology has increased rapidly, and its applications have expanded significantly. Currently, the mainstream AI technology is based on neural networks, and with the emergence of large language models, the computational requirements and memory usage for running neural networks have increased exponentially. Graphics Processing Units (GPUs), which use high-bandwidth and high-capacity memory, are the leading AI processor due to their high parallelism, which greatly reduces computation time. However, given the current GPU shortages and rapid price increases, the sustainability of this cycle seems uncertain. Additionally, the on-device AI market has recently been emerging, which is expected to expand significantly in the near future. Obviously, power-hungry GPUs are an unfavorable solution to on-device AI.


The high thermal design power of GPUs mainly arises from a vast amount of data transfer between their main memory and processors. As an alternative, Compute-in-Memory (CiM) aims to co-integrate the memory and processing units on a single chip, so that CiM units work as a memory with computing units around. CiM avoids consuming a large amount of power and latency on data transfer. Particularly, the significant reduction in latency is effectively equivalent to a large increase in memory band-width. Among diverse types of CiM, nonvolatile memory-based analog CiM (ACiM) is deemed a forward-looking CiM that makes use of highly efficient analog computation for GEMV (GEneral Matrix Vector multiplication) and/or GEMM (GEneral Matrix Matrix multiplication). Note that GEMV and GEMM are the major type of operation in machine learning. Resistance-based nonvolatile random-access memory (RAM) like resistive RAM (RRAM) allows the current through each bit-cell (sharing the same bit-line) to be summed at the bit-line according to Kirchhoff’s current law, and the bit-cell current is gated by its word-line signal. This is the substrate of analog multiply-accumulate (MAC) operations that are elementary operations for GEMV and GEMM. This analog computation is advantageous in terms of power and area overheads over digital computation. Further, its nonvalatility is capable of always-on operations at zero standby power, suitable for on-device AI applications.


Albeit promissing, there exist several technical obstacles that should be overcome for ACiM to come into practical play, which include computational unreliability and limited memory capacity. The inherent unreliability of analog computation is likely improved by optimally adopting digital processing the extent to which the advantages of anlog computation are valid. The limited memory capacity can be improved by applying advanced tech nodes and higher dimensional memory architecture. These technical obstacels aside, open ecosystems for ACiM need to be built to accelerate practical applications of this promising technology. Of course, the ecosystems are based on ACiM hardware, but importantly include (i) its software stack like software development kits (SDKs) and compliers with front-ends compatiable with existing machine learning libraries, and (ii) user-friendly software libraries that machine learning software engineers can easily use. Most importantly, open platforms that glue the users would be the key to the success of this emerging hardware technology. Surely, GPUs were not designed for large language models back in 1990s. The open ecosystems based on GPUs, which include machine learning libraries like PyTorch and TensorFlow, have been creating new application domains. We keep our eyes on this lesson for the success of ACiM.



Popular Insights

Previous Next List