[Summary] cMPI: Accelerating HPC Communication via CXL Memory Sharing
cMPI: Accelerating HPC Communication via CXL Memory Sharing
Message Passing Interface (MPI) is the standard programming model for high-performance computing (HPC) across distributed nodes. Traditionally, MPI relies on complex network interconnects (like Ethernet or InfiniBand) with extensive software stacks for cross-node communication.
This work presents cMPI, the first implementation to optimize MPI point-to-point communication (both one-sided and two-sided) by using Compute Express Link (CXL) memory sharing on a real hardware platform. cMPI transforms network-based communication into simple memory transactions within a shared CXL memory pool, effectively bypassing traditional network protocols.
Key Concepts and System Design
The cMPI approach is built upon a CXL pooled memory platform, which allows multiple hosts to access a shared memory region The cMPI library extends MPICH-4.2.3 and integrates a CXL Shared Memory (SHM) Arena for managing shareable data objects. This design avoids changes to user-facing MPI application code. As shown in Figure 1:

Figure 1: The overview of cMPI.
- -Two-sided Communication:Message queues are created and maintained directly with the CXL shared memory. Sender and receiver processes exchange data by performing enqueue and dequeue operations via memory load/store instructions.
- -One-sided Communication (RMA): Remote Memory Access windows are allocated in the CXL SHM. Origin processes perform MPI_Put() (write) and MPI_Get() (read) operations by directly accessing the target rank's data in the shared CXL memory, eliminating network transfers entirely.
A key challenge addressed by cMPI is memory coherence across nodes, which is solved using an efficient software-based cache coherence mechanism involving cache flushing and memory fences, as hardware-based solutions are less scalable.
Performance Evaluation
cMPI delivers substantial performance improvements, particularly for latency-sensitive applications with small message sizes (up to 16 KB).
- -Latency Advantage: CXL memory sharing achieves 7.2x to 8.1x lower latency compared to standard TCP-based interconnects commonly found in small-to-medium clusters.
- -Overall Performance: For small messages, cMPI outperforms TCP over a standard Ethernet NIC by up to 49x in latency and 72x in bandwidth.
- -Comparison with High-End Networks:
- CXL SHM outperforms TCP over a high-end SmartNICs (Mellanox CX-6 Dx) by up to 48x in latency for small messages.
- For messages larger than 16 KB, high-end SmartNICs with RDMA still offer better scalability and bandwidth due to offloaded network transfers that require less continuous CPU involvement.

Figure 2: Latency of one-sided MPI communication.
As shown in Figure 2, cMPI maintaining a consistently low latency for small messages, whereas traditional TCP latency remains in the hundreds of microseconds range.
Conclusion
cMPI demonstrates the significant potential of using CXL shared memory to accelerate MPI point-to-point communications across nodes. By leveraging direct memory access and bypassing complex network stacks, cMPI offers superior performance for latency-bound HPC workloads.
The Publication : Link

![[Summary] A Memristor-based In-Memory Computing SoC with Efficient Depthwise Convolution](https://mis-prod-koce-research-user-cdn-01-blob-ep.azureedge.net/web/blog/20260806/thumb_sEVmfv3K.20260806085519833.jpg)
![[Insights] Evolving Role of Emerging Memories in Next-Generation Computing](https://mis-prod-koce-research-user-cdn-01-blob-ep.azureedge.net/web/blog/20260806/thumb_NgJoNLYn.20260806080458980.jpg)
![[Summary] Electrical Characteristics of the 4F2 Vertical Gate (VG) DRAM integrated with Bit-Line Shielding (BLS) and Back Gate (BG) Transistor](https://mis-prod-koce-research-user-cdn-01-blob-ep.azureedge.net/web/blog/20260806/thumb_DOUurl5m.20260806080624234.png)