Blog

Stay updated with our latest news and announcements.

Insights

[Invited Insights] Emerging Memory in Servers: Opportunities and Challenges

Yun ho OhYun ho Oh, Professor, Korea University

Modern datacenters rely heavily on DRAM to meet the performance demands of online services. However, DRAM is expensive and difficult to scale, making it an unsustainable choice for large-scale memory expansion. The increasing gap between memory demand and available DRAM capacity has prompted the industry to explore alternative memory technologies. Several techniques, including Selector Only Memory (SOM) [1], Magnetic Random Access Memory (MRAM), and Ferroelectric RAM (FeRAM), have emerged as promising solutions. SOM, a high-density, non-volatile memory, leverages selector-based addressing to improve scalability while minimizing leakage power. Compared to DRAM, SOM provides higher density and lower standby power, making it well-suited for high-capacity server environments.


These emerging memory technologies should be strategically placed within server architectures to maximize efficiency and performance. While AI and machine learning workloads are often associated with high memory demands, servers supporting online services, such as cloud applications, search engines, and databases, require scalable and low-latency memory solutions. Online services power a wide range of applications, including caching systems, search engines, web servers, and media streaming platforms, all of which require efficient memory management to handle large datasets and meet strict latency requirements. Memcached, a widely used in-memory key-value store, accelerates web applications by caching frequently accessed data. Web search engines process massive datasets, relying on fast access to indexed information to deliver quick query responses. High-throughput web servers maintain session states and frequently accessed content with minimal delay. Media streaming platforms require large memory pools to buffer high-definition content, reducing dependence on slow storage and ensuring smooth playback. These applications demand scalable, high-performance memory architectures to support millions of concurrent users.


Online services exhibit highly skewed data access patterns, where only a small fraction of data is frequently accessed, while the majority remains infrequently used. This imbalance makes it impractical to store all data in expensive DRAM. Instead, DRAM can serve as a cache for hot data, while SOM acts as primary host memory, expanding capacity at a lower cost while preserving low-latency access for performance-critical operations. By integrating SOM into server architectures, online services can scale efficiently, reduce memory costs, and maintain the responsiveness required for modern cloud workloads.


We analyze the DRAM miss ratio while varying the DRAM-to-emerging memory capacity ratio for CloudSuite [2] workloads, assuming a memory hierarchy with DRAM and emerging memory (replacing flash with SOM in this exmaple). Also, we examine the trade-off between DRAM capacity and the bandwidth required to refill DRAM. We calculate the emerging memory bandwidth required per core using the following equation:


BWEM = (BWDRAM / Cache Block Size) × Miss Rate × Page Size


Using 0.5 GBps as the average DRAM bandwidth, a 4KB page size, and a 64B cache block size, we derive the bandwidth requirements for different DRAM configurations.


Figure 1 Miss rate depending on DRAM capacity.


Figure 1 shows the average cache miss ratio across workloads and the required emerging memory bandwidth for different DRAM capacities. The miss rates flatten at 3% DRAM capacity, which corresponds to 60 GBps of emerging memory bandwidth for a 64-core system. With PCIe Gen5 specifications supporting up to 128 GBps bandwidth, it is feasible to meet these bandwidth requirements using multiple emerging memory devices. We assume a system with 1TB of data hosted in emerging memory and a DRAM cache of 3% capacity (32GB), requiring 60 GBps aggregate bandwidth for 64 cores. Since SOM exhibits shorter latencies and better endurance than flash, they would be an ideal fit for modern datacenters where long memory latencies should be minimized.


Challenges

While emerging memory technologies provide significant advantages, their integration into existing server architectures presents several challenges that should be addressed for optimal performance.

One key challenge is DRAM cache design, as managing SOM alongside DRAM requires efficient caching policies to ensure frequently accessed data is kept in DRAM while minimizing unnecessary data movement. Using large cache blocks (e.g., 4KB) can improve performance by reducing metadata tracking overhead, but it also increases complexity in cache management and eviction policies. Striking the right balance between block size and metadata efficiency is crucial for maintaining system performance.


Another major challenge lies in operating system abstractions, as traditional demand paging mechanisms introduce significant overhead while moving data between DRAM and emerging memory. To fully leverage emerging technologies' high density and persistence, new memory management approaches are needed, including hardware-supported memory allocation and direct-access mechanisms. Additionally, asynchronous access models—such as user-level thread switching—can help hide latency by allowing workloads to continue execution while waiting for data movement between memory layers.


Lastly, memory access optimization is essential for integrating SOM into large-scale systems. Traditional TLB shootdowns, which ensure consistency in address translations, become a major bottleneck while scaling memory capacity. Efficient virtual memory translation mechanisms tailored for SOM are necessary to minimize translation overhead while maintaining high throughput. Addressing these challenges through hardware-software co-design will be critical in ensuring that emerging memory technologies can be effectively deployed in next-generation datacenters.


Conclusion

Emerging memory technologies, such as SOM, provide a cost-effective and scalable alternative to DRAM for modern datacenters. By adopting tiered memory architectures, where DRAM functions as a cache and SOM serves as high-density host memory, online services can maintain performance while significantly reducing memory costs. However, challenges in cache management, OS support, and bandwidth optimization should be addressed to fully realize the potential of these technologies. Future architectures should integrate hardware-software co-design to efficiently manage SOM alongside DRAM, ensuring seamless performance for next-generation datacenter workloads.

 

References

[1] M. Kim et al., "First Demonstration of Fully Integrated 16 nm Half-Pitch Selector Only Memory (SOM) for Emerging CXL Memory," 2024 IEEE Symposium on VLSI Technology and Circuits (VLSI Technology and Circuits), Honolulu, HI, USA, 2024, pp. 1-2, doi: 10.1109/VLSITechnologyandCir46783.2024.10631351.

[2] Michael Ferdman et al., “Clearing the clouds: a study of emerging scale-out workloads on modern hardware,” In Proceedings of the seventeenth international conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS XVII). Association for Computing Machinery, New York, NY, USA, 2012



Popular Insights

Previous Next List