Skip to main content

MegaMmap at One Year: Memory Virtualization for Data-Intensive Computing

· 3 min read
Anthony Kougkas (opens in a new tab)
Co-Founder and Executive Director, Gnosis Research Center

One year ago, Luke Logan presented MegaMmap at SC'24, introducing a software distributed shared memory system that makes datasets larger than available DRAM accessible through familiar programming interfaces. The core problem it addresses has only grown more relevant since.

What MegaMmap Solved​

HPC applications, scientific simulations, and machine learning workloads routinely generate datasets that exceed available DRAM. The traditional solutions are buying more memory (expensive and unsustainable) or manually partitioning data across storage tiers (complex and error-prone). MegaMmap offered a third path: transparent memory virtualization that handles data placement automatically.

The system manages data across heterogeneous storage tiers, from DRAM to NVMe, SSD, and HDD, presenting applications with a unified memory abstraction. Three capabilities made this practical:

Transparent virtualization. Applications work with datasets as if they reside entirely in memory, using C++ vector-like interfaces. No explicit I/O management or partitioning logic required.

Workload-aware placement. Applications declare their access intent through transactions, and MegaMmap uses those hints to guide data placement across tiers, prefetching, and eviction.

Intent-based coherence. Instead of generic coherence protocols, MegaMmap provides workload-specific optimizations for read-only analytics, write-only simulations, and mixed access patterns.

SC'24 Results​

The results at SC'24 demonstrated that intelligent tiering can reduce memory requirements without sacrificing performance:

  • 2.6x DRAM reduction with competitive performance on ML clustering workloads
  • As much as 2x faster than Apache Spark on cosmological data analytics
  • 45% less code compared to manual out-of-core implementations
  • Competitive weak scaling to 768 processes

These numbers validated the central thesis: a well-designed memory virtualization layer can bridge the gap between available DRAM and dataset size without pushing complexity onto application developers.

Connecting to the GRC Portfolio​

MegaMmap reflects patterns established across multiple GRC projects:

Hermes taught us how to automatically manage heterogeneous storage hierarchies. MegaMmap extended those lessons from I/O buffering to full memory virtualization with distributed shared memory semantics. Learn more about Hermes.

LABIOS bridged HPC and Big Data storage paradigms. MegaMmap applies similar principles to bridge in-memory computing with storage-aware computing. Learn more about LABIOS.

ChronoLog provides time-ordered storage for activity and provenance data. MegaMmap addresses a different layer by expanding effective memory capacity across DRAM and storage. Learn more about ChronoLog.

Each project addresses a different facet of the data-centric computing challenge. Together, they form a coherent portfolio: intelligent systems that manage data movement across heterogeneous tiers so that applications and their developers don't have to.

What's Next​

As AI models grow larger and scientific datasets expand, the pressure on memory systems will intensify. MegaMmap's approach of transparent, workload-aware memory management positions it well for emerging challenges in LLM inference (where KV caches strain GPU memory), distributed training (where checkpointing I/O dominates), and scientific workflows (where multi-physics simulations exceed node-local DRAM).