Skip to main content

UniMCC

Cross-layer support for near-memory processing and disaggregated memory
Lifecycle
Active
Period
2023–Present

UniMCC is developing full-stack support for systems that combine near-memory data processors with disaggregated shared memory. The project aligns architecture, hardware and software interfaces, code generation, runtime support, and performance models so data-intensive applications can use those resources as one memory-centric system.

Published Systems​

CHROME: Concurrency-Aware Holistic Cache Management (HPCA'24)​

CHROME (opens in a new tab) coordinates cache replacement, bypassing, and prefetching through concurrency-aware online reinforcement learning. It adapts decisions from program features and system feedback. The HPCA 2024 paper reports up to 13.7% higher performance than LRU in multicore systems.

CHROME concurrency-aware cache management design overview
CHROME coordinates cache replacement, bypassing, and prefetching.

ACES: Adaptive and Concurrency-Aware Sparse Matrix Accelerator​

ACES (opens in a new tab) accelerates sparse matrix-matrix multiplication with an execution flow that adapts to each sparse pattern. A concurrency-aware global cache and a non-blocking buffer balance reuse, parallelism, and synchronization. The ASPLOS 2024 paper reports a 2.1x speedup over the accelerators it compared against.

ACES sparse matrix accelerator design overview
ACES combines adaptive execution with concurrency-aware cache control.

2026 Results​

Three published papers extend UniMCC's cross-layer approach. Zion develops a comprehensive, adaptive, and lightweight hardware prefetcher. I/O Analysis Is All You Need analyzes the I/O behavior of long-sequence attention to guide hardware and software co-design. I/O-Aware PIM Acceleration applies hybrid sparse attention to processing-in-memory acceleration for long-sequence LLM inference.

Compiler and Code Generation Work​

TrackFM: Compiler-Based Far Memory​

TrackFM, published at ASPLOS 2024 (opens in a new tab) by Brian Tauro, Brian Suchy, Simone Campanoni, Peter Dinda, and UniMCC co-PI Kyle Hale, is a compiler and runtime approach for running unmodified applications with far memory. LLVM passes inject access guards and connect applications to the AIFM runtime. Fast paths and loop chunking reduce redundant checks.

TrackFM compiler-based far memory system workflow
TrackFM transforms applications for a far-memory runtime.

CUDA Code Generation for KGE Score Functions​

A second active direction generates CUDA code for knowledge graph embedding score functions. It adds batched scalar-vector, vector-vector, and matrix-vector operators for fusion, then uses runtime inspection to cache unique data indices in shared memory.

University of Iowa
National Science Foundation

This material is based upon work supported by the National Science Foundation under Grant No. CNS-2310422, CNS-2310423. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation.