Skip to main content

OptMem

Concurrency-aware models and control methods for modern memory hierarchies
Lifecycle
Completed
Period
2020–2024

OptMem developed performance models and cache-management methods that treat data locality and memory-access concurrency together. The project connected a recursive memory model with four adaptive systems for prefetching, cache partitioning, replacement, and holistic cache control.

Modeling Concurrent Memory Access​

Concurrent Average Memory Access Time (C-AMAT) extends the conventional AMAT model to systems with overlapping memory requests. It captures latency, locality, and concurrency in one recursive formulation across the memory hierarchy. OptMem generalized this model for hierarchical systems, including request splitting and merging at cache devices, and used it to guide architecture design and optimization.

Memory system viewed as multi-tree structure
A memory system represented as a multi-tree.

Concurrency-Aware Systems​

APAC: Adaptive Prefetching​

APAC (opens in a new tab) uses pure prefetch coverage, a metric designed for overlapping memory requests, to adjust prefetch aggressiveness at runtime. The ICCD 2020 paper reports an average 17.3% performance improvement over FDP for memory-intensive single-thread benchmarks and 8.5% higher IPC than FDP on a multicore system.

APAC adaptive prefetch framework adjusting aggressiveness
APAC adjusts prefetch aggressiveness with runtime metrics.

Premier: Shared Cache Partitioning​

Premier (opens in a new tab) uses pure misses per kilo instructions to measure cache efficiency under concurrent access. Its adaptive insertion and promotion policies approximate cache partitioning without rigidly dividing capacity. The ICCD 2021 paper reports 15.45% higher system performance and 10.91% better fairness than UCP in an eight-core system.

Premier cache partitioning PMPKI and CPI results
PMPKI and CPI across cache sizes in SPEC workloads.

CARE: Cache Replacement​

CARE (opens in a new tab) uses pure miss contribution to estimate the cost of each outstanding miss and guide cache replacement. The HPCA 2023 paper reports 10.3% higher IPC than LRU in a four-core system, 13.0% in an eight-core system, and 17.1% in a 16-core system.

CARE concurrency-aware cache management design overview
CARE uses concurrency-aware miss costs for cache management.

CHROME: Holistic Cache Control​

CHROME (opens in a new tab) jointly manages replacement, bypassing, and prefetching with concurrency-aware online reinforcement learning. It adapts decisions from program features and system feedback. The HPCA 2024 paper reports up to 13.7% higher performance than LRU in multicore systems.

CHROME holistic cache management design overview
CHROME coordinates cache replacement, bypassing, and prefetching.

The UniMCC project extends this line of research across a full memory-centric computing stack.

Florida International University
National Science Foundation

This material is based upon work supported by the National Science Foundation under Grant No. CCF-2008907, CCF-2008000. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation.