MegaMmap at SC24: Tiered Distributed Shared Memory for Data-Intensive Computing
Presented at SC24 International Conference for High Performance Computing, Networking, Storage, and Analysis November 17 to 22, 2024 • Atlanta, Georgia, USA
Available Now MegaMmap is released as open-source software. Visit the GitHub repository (opens in a new tab)
At SC24, Luke Logan presented MegaMmap, a software-based distributed shared memory (DSM) technology that blurs the boundary between memory and storage. Developed by Luke Logan, Anthony Kougkas, and Xian-He Sun, MegaMmap lets data-intensive applications work with datasets larger than available DRAM through memory tiering, prefetching, and coherence optimization.

Luke Logan presenting MegaMmap research at SC24
The Memory Crisis in Modern HPC
Scientific workloads such as cosmological simulation and machine learning increasingly process datasets larger than the DRAM of the nodes they run on. Out-of-core programming works around the limit, but it adds I/O code and complexity to every application.
The MegaMmap Solution
MegaMmap is a software distributed shared memory (DSM) system that presents a unified, byte-addressable interface spanning multiple storage tiers, from DRAM to NVMe, SSD, and HDD.
Key Technical Contributions
1. Infinite Memory Abstraction MegaMmap allows applications to present massive datasets as if they were in main memory, eliminating the need for explicit I/O management. A simple C++ vector-like interface enables developers to work with terabytes of data using familiar programming patterns.
2. Intelligent Tiering The system automatically manages data placement across heterogeneous storage based on access patterns, device performance characteristics, and application requirements. Hot data stays in fast DRAM, while cold data gracefully migrates to appropriate storage tiers.
3. Transactional Memory API Unlike traditional DSMs that must guess access patterns, MegaMmap allows applications to declare their intent through transactions. This enables:
- Optimized prefetching strategies
- Reduced coherence overhead
- Better data placement decisions
4. Intent-Aware Coherence MegaMmap provides workload-specific coherence optimizations for common HPC patterns:
- Read-only analytics (machine learning inference)
- Write-only simulations (scientific modeling)
- Mixed workloads (iterative algorithms)
Performance Results
The SC'24 paper evaluates MegaMmap on HPC applications and AI workloads and reports:
Machine Learning and AI Workloads
For KMeans clustering on large-scale cosmological datasets:
- 2.6x reduction in DRAM usage while maintaining competitive performance
- 45% less code compared to traditional out-of-core programming approaches
- As much as 2x faster than the Apache Spark implementations evaluated
Scientific Computing and Simulations
Gray-Scott reaction-diffusion modeling and complex scientific simulations demonstrated:
- Datasets larger than DRAM: processed datasets exceeding available physical DRAM
- At least 20% faster performance than the tiered I/O systems evaluated through grid size L=2688
- Scaling across memory hierarchies from gigabytes to terabytes of data
Big Data Analytics and Data-Intensive Computing
DBSCAN clustering on Gadget-4 cosmological simulation data performed competitively with the MPI-based implementation.
The Technology Behind MegaMmap: Architecture and Implementation Details
MegaMmap's software-based distributed memory architecture consists of the following components, which provide transparent memory virtualization and intelligent data management across heterogeneous storage hierarchies:

Distributed Caching and Memory Hierarchy Management
- Private Cache (pcache): Per-process DRAM cache providing ultra-low-latency access and reducing network overhead in distributed systems
- Shared Cache (scache): Distributed, tiered cache layer spanning all processes and storage tiers for efficient data sharing and locality optimization
- Asynchronous Operations: Intelligent overlap of computation with data movement to hide I/O latency and maximize hardware utilization
Intelligent Data Management and Placement Optimization
- Prefetcher: Predictive prefetching engine that anticipates future memory accesses based on declared transaction patterns and access history
- Data Organizer: Sophisticated data placement algorithm that positions data across storage tiers based on performance scores, access frequency, and latency requirements
- Persistent Integration: Transparent, automatic staging of data to/from various storage backends (HDD, SSD, NVMe) and persistent formats
Advanced Coherence Protocols and Memory Consistency
- Minimal Overhead: Optimized coherence mechanisms that avoid traditional distributed shared memory communication penalties
- Workload-Aware Optimization: Specialized coherence protocols tuned for HPC access patterns rather than generic cloud computing workloads
- Strong Consistency: Scheduling orders operations on the same page to provide read-after-write guarantees
Developer Experience: Simplified Programming Model for HPC
MegaMmap provides a simplified programming interface and abstractions for memory management. Traditional out-of-core computing requires explicit data partitioning, complex I/O orchestration, and synchronization logic. MegaMmap eliminates this complexity through a clean, familiar C++ vector-like API that abstracts away storage hierarchy details. Consider this practical KMeans clustering example:
// Create a shared vector from a Parquet file
mm::Vector<Point3D> pts("/points.parquet");
pts.BoundMemory(MEGABYTES(1)); // Limit to 1MB DRAM
pts.Pgas(rank, nprocs); // Partition across processes
// Begin read-only transaction
auto tx = pts.SeqTxBegin(pts.local_off(), pts.local_size(), MM_READ_ONLY);
float distance = 0;
for (Point3D p : tx) {
distance += pow(NearestCentroid(p, ks), 2);
}
This simple interface eliminates the complexity of manual data partitioning, I/O management, and memory synchronization that typically plague out-of-core applications.

Performance Validation and Benchmarking
Evaluation on a research cluster with hierarchical storage tiers (DRAM, NVMe, SSD, HDD):
Scalability Studies and Large-Scale Distributed Computing
- Scaling to 768 processes across nodes
- Competitive performance parity with conventional MPI-based distributed implementations
- As much as 2x faster than Apache Spark

Memory Efficiency and DRAM Optimization
- Data eviction and prefetching policies based on access patterns and memory pressure
Cost-Performance Analysis and Total Cost of Ownership
- 1.8x performance improvement when optimizing with NVMe storage tiering

Open Source and Community: Contributing to HPC Software Infrastructure
MegaMmap is now available as open-source software under an OSI-approved license, making its memory management techniques available to the HPC and research computing community.
The project includes documentation, reproducible benchmarks, and integration examples.
Get Started with MegaMmap: Installation, Documentation, and Support
Researchers, HPC system administrators, and developers interested in exploring MegaMmap and implementing memory-virtualized computing systems can:
- Access the open-source MegaMmap GitHub repository (opens in a new tab) featuring complete source code, API documentation, and implementation guides for distributed memory systems.
- Read the complete SC24 research paper (opens in a new tab) for detailed evaluation results, benchmarking methodology, technical insights, and comparison with alternative distributed memory approaches.
- Contact the lead author, Dr. Luke Logan, for collaboration and technical questions.
Acknowledgments
This research is supported by the National Science Foundation (NSF) (opens in a new tab) through grants CSSI-2104013 and Core-2313154, and by the U.S. Department of Energy (DOE) (opens in a new tab) Office of Science under Contract DE-SC0024593. The team also acknowledges the Chameleon Cloud (opens in a new tab) testbed for providing development infrastructure and resources that enabled this research in HPC and memory management systems.
