OptMem
OptMem developed performance models and cache-management methods that treat data locality and memory-access concurrency together. The project connected a recursive memory model with four adaptive systems for prefetching, cache partitioning, replacement, and holistic cache control.
Modeling Concurrent Memory Access
Concurrent Average Memory Access Time (C-AMAT) extends the conventional AMAT model to systems with overlapping memory requests. It captures latency, locality, and concurrency in one recursive formulation across the memory hierarchy. OptMem generalized this model for hierarchical systems, including request splitting and merging at cache devices, and used it to guide architecture design and optimization.

Concurrency-Aware Systems
APAC: Adaptive Prefetching
APAC (opens in a new tab) uses pure prefetch coverage, a metric designed for overlapping memory requests, to adjust prefetch aggressiveness at runtime. The ICCD 2020 paper reports an average 17.3% performance improvement over FDP for memory-intensive single-thread benchmarks and 8.5% higher IPC than FDP on a multicore system.

Premier: Shared Cache Partitioning
Premier (opens in a new tab) uses pure misses per kilo instructions to measure cache efficiency under concurrent access. Its adaptive insertion and promotion policies approximate cache partitioning without rigidly dividing capacity. The ICCD 2021 paper reports 15.45% higher system performance and 10.91% better fairness than UCP in an eight-core system.

CARE: Cache Replacement
CARE (opens in a new tab) uses pure miss contribution to estimate the cost of each outstanding miss and guide cache replacement. The HPCA 2023 paper reports 10.3% higher IPC than LRU in a four-core system, 13.0% in an eight-core system, and 17.1% in a 16-core system.

CHROME: Holistic Cache Control
CHROME (opens in a new tab) jointly manages replacement, bypassing, and prefetching with concurrency-aware online reinforcement learning. It adapts decisions from program features and system feedback. The HPCA 2024 paper reports up to 13.7% higher performance than LRU in multicore systems.

The UniMCC project extends this line of research across a full memory-centric computing stack.