DeepIO
DeepIO develops data-path methods for scientific AI. Its systems profile training I/O, ingest scientific formats, move model state between concurrent training and inference, tune layered I/O stacks, and manage KV caches under concurrent LLM serving.
The work spans six published systems. DLIO (opens in a new tab) underpins MLPerf Storage and guided optimizations that reduced training time by up to 6.7x. Stimulus accelerated scientific data ingestion by up to 5.3x. TunIO reduced tuning time by up to 73% relative to H5Tuner. Viper reduced model update latency by approximately 9x with GPU-to-GPU transfer. UnboxKV characterized KV cache behavior under concurrent inference, and PKAS used predicted cache demand to deliver up to 7.34x higher throughput and 8x lower latency than the schedulers evaluated in its HPDC'26 paper.
The Data Path
Scientific workflows can generate training data, update a model, and serve inference concurrently. DeepIO addresses three I/O constraints across that cycle:
- Training I/O. Scientific formats such as HDF5, ADIOS2, and NetCDF introduce data movement and conversion costs in deep learning pipelines.
- Model synchronization. Checkpoint frequency must balance training overhead against the quality loss caused by stale inference models.
- Inference memory. Model weights and KV caches compete for GPU memory as concurrent requests arrive and expand during decoding.
Viper: Adaptive Model Transfer

Viper coordinates model updates between concurrent training and inference. It selects checkpoint times from predicted inference quality, then transfers model state through memory instead of the parallel file system.
Checkpoint scheduling. Viper learns the relationship between checkpoint timing and downstream inference accuracy. It uses that model to balance cumulative inference quality against training slowdown.
Memory-first transfer. GPU-to-GPU transfers bypass the storage stack. DRAM-to-DRAM transfers provide a fallback when GPU memory is constrained.


Results



The ICPP'24 paper (DOI: 10.1145/3673038.3673070 (opens in a new tab)) reports:
- Approximately 9x lower model update latency with GPU-to-GPU transfer (NT3: 12x, TC1: 9x, PtychoNN: 15x)
- Approximately 3x lower latency with DRAM-to-DRAM host transfer than the baseline
- Higher cumulative inference accuracy with predictor-derived checkpoint schedules than with epoch-based baselines



Published at ICPP'24. Read the paper (opens in a new tab)
UnboxKV: KV Cache Characterization Under Concurrency
UnboxKV studies how concurrent LLM inference requests use the KV cache. The study focuses on the GPU memory left after model weights are loaded and on the cache capacity consumed as requests progress through token generation.
Methodology. The study instruments vLLM to measure token throughput, KV cache access patterns, and forward-pass load balancing during prefill and decode across several benchmarks.
Findings. The paper compares batching strategies and two eviction policies: drop and recompute, and swap to host memory. It identifies specific opportunities for inference runtimes to reduce cache pressure under concurrency.
Published at IPDPS'25. Read the paper (opens in a new tab)
Best Poster Nominee, SC'24. The UnboxKV characterization tool was presented as a poster at SC'24 and nominated for the Best Poster award.
PKAS: Predictive KV-Cache-Aware Scheduling
PKAS turns KV cache measurements into an admission policy. Continuous-batching runtimes can accept prefill requests without reserving the cache capacity their decode phases will need, causing overflow, preemption, and recomputation.
PKAS simulates future KV cache utilization with a low-overhead technique and combines it with lightweight output-length prediction to decide which requests to admit. The HPDC'26 paper reports up to 7.34x higher throughput and 8x lower latency than the baseline schedulers it is compared against, across several models and workloads, with the largest gains on long-context workloads.
Published at HPDC 2026. Read the paper (opens in a new tab) · DOI (opens in a new tab)
DLIO: A Benchmark for Deep Learning I/O
DLIO is a data-centric benchmark for the I/O behavior of scientific deep learning workloads.
DLIO was built from I/O profiles of scientific deep learning workloads on the Theta supercomputer at Argonne, covering cosmology, particle physics, computer vision, and astrophysics. The CCGrid'21 paper reports that the benchmark reproduces the I/O behavior of those workloads with over 90% similarity, and that optimizations guided by it lowered training time by up to 6.7x.
The MLPerf Storage benchmark (MLCommons (opens in a new tab)) is built on DLIO. Industry has used it independently: a Dell Technologies talk at the SNIA Compute, Memory, and Storage Summit 2024 emulates ResNet-50 training with DLIO (slides (opens in a new tab), slide 7; video (opens in a new tab)).
Best Paper Award (First Prize), CCGrid 2021. Read the paper (opens in a new tab) · GitHub (opens in a new tab)
Stimulus: Scientific Data Ingestion for AI
Stimulus connects scientific formats such as HDF5, PnetCDF, ADIOS2, GNCF, and Silo to TensorFlow and PyTorch through its StimOps functions and StimPack abstraction.
On the Summit supercomputer at Oak Ridge, the CCGrid'22 paper reports speedups of 5.3x for Cosmic Tagger, 2.9x for Distributed FFN, and 1.9x for CosmoFlow, with ideal I/O scalability up to 768 GPUs.
Published at CCGrid 2022. Read the paper (opens in a new tab)
TunIO: I/O Stack Tuning
TunIO identifies high-impact parameters across libraries, middleware, and parallel file systems. It extracts I/O kernels from application source code and searches the configuration space with reinforcement learning and early stopping. The IPDPS'24 paper reports up to 73% less tuning time than H5Tuner at the same performance gain.
Published at IPDPS 2024. Read the paper (opens in a new tab)
Multi-GPU Inference
DeepIO is extending its inference work toward multi-GPU serving. Three systems are in development:
- DyTO (Dynamic Tensor Routing) studies how model shards and KV cache partitions should move among GPU memory, host memory, and NVMe as request patterns change.
- GPUcompress develops CUDA kernels that compress KV cache entries before transfer, freeing GPU memory for active requests.
- GPUDirect and NIXL benchmarking characterizes GPU-attached storage paths for AI workloads. Initial results appeared in an eScience'25 poster.
These directions connect DeepIO to IOWarp's GPU data infrastructure and MPI4AI's communication layer for distributed AI.
DeepIO is developed with Argonne National Laboratory and Lawrence Livermore National Laboratory. The Argonne Leadership Computing Facility provides computing resources.