Skip to main content

DeltaAI User Guide

Prerequisites

This guide assumes familiarity with Linux, SSH, and job scheduling basics. If you need a refresher, start with the HPC Foundations tutorials. For DeltaAI hardware specs and other GRC computing resources, see the Resources page.

Before You Start​

  • DeltaAI runs ARM aarch64 CPUs (NVIDIA Grace) with NVIDIA H100 GPUs (GH200 superchips). This is not x86.
  • General user SSH access requires an NCSA password and Duo MFA. SSH keys require an approved Science Gateway account.
  • Submit jobs through Slurm. Use srun for an interactive shell and sbatch for batch jobs. For interactive MPI, use salloc, then launch with srun.
  • Interactive sessions cost 2x the batch rate. Prefer batch jobs for real work.
  • Install user software with conda or use NCSA modules. Check compiler requirements before building.
  • Keep all work in /work/hdd/your_project/your_username/. Home directory is only 100 GB.

Introduction​

DeltaAI is a GPU-focused supercomputer operated by the National Center for Supercomputing Applications (NCSA) at the University of Illinois. It is accessed through the ACCESS (opens in a new tab) program, which provides compute allocations to research groups across the U.S. NCSA maintains the official DeltaAI user documentation (opens in a new tab); consult it for current system configuration and policies.

GRC researchers use DeltaAI for GPU-accelerated workloads including machine learning training, high-performance I/O benchmarking, and systems research. The cluster consists of 152 quad-GPU nodes based on the NVIDIA GH200 Grace Hopper Superchip architecture.

Confirm your ACCESS project, Slurm account, and storage project code with the PI or allocation manager. The examples below use placeholders for allocation-specific values.

Hardware Overview​

Each DeltaAI compute node contains four GH200 superchips. Each superchip pairs an NVIDIA Grace ARM CPU with an H100 GPU connected via NVLink-C2C.

ComponentSpecification
GPU4x NVIDIA GH200 (H100 96 GB HBM3 each)
CPU4x NVIDIA Grace (72 ARM cores each)
Cores/Node288 ARM Neoverse V2 cores
Memory480 GB LPDDR5 (CPU) + 384 GB HBM3 (GPU)
Local Storage3.9 TB NVMe SSD
Network4x 200 GbE HPE Slingshot 11
Total Nodes152

Hardware values follow NCSA's system architecture guide (opens in a new tab).

GPU Specifications​

PropertyValue
ArchitectureHopper (SM 9.0)
VRAM per GPU96 GB HBM3
GPUs per Node4
Total Cluster GPUs608

Architecture Notes​

DeltaAI uses ARM aarch64 processors, not x86. This has several practical consequences:

  • Use aarch64 builds of tools and libraries. An x86_64 executable cannot run natively on a Grace CPU and may report "Exec format error".
  • Check the memory page size with getconf PAGESIZE. Tools whose allocator does not support the reported page size may fail with "Unsupported system page size."
  • Check the operating system with cat /etc/os-release. Shell detection scripts should accept the local $OSTYPE value or use uname -s.
  • NCSA provides the Cray compiler wrappers cc, CC, and ftn. Use a programming environment appropriate for your project; see NCSA's compiler guidance (opens in a new tab).

Accessing the Cluster​

NCSA Account Setup​

Before you can access DeltaAI, you must be added to the ACCESS allocation. Contact the PI or an allocation manager to add your NCSA username to the project.

Required accounts:

  1. An ACCESS account (opens in a new tab), linked to your institution
  2. An NCSA Identity (opens in a new tab) linked to the allocation; check your ACCESS profile for the NCSA username
  3. NCSA Duo MFA (opens in a new tab) enrollment for two-factor authentication

SSH Configuration​

General user access requires an NCSA password and Duo MFA. SSH keys are reserved for approved Science Gateway accounts; see NCSA's login instructions (opens in a new tab).

Add the following to your ~/.ssh/config:

Host delta-ai
HostName dtai-login.delta.ncsa.illinois.edu
User your_username
PreferredAuthentications keyboard-interactive,password
ServerAliveInterval 60
ServerAliveCountMax 3
ForwardAgent no

Then connect with:

ssh delta-ai

You will be prompted for your NCSA password, then a Duo verification (push notification, phone call, or recovery code).

Login Node Persistence​

DeltaAI has four login nodes. The delta-ai hostname round-robins between them. If you start a tmux session on one login node, you must reconnect to the same node to reattach.

Pin your SSH connection to a specific node:

Host delta-ai-4
HostName gh-login04.delta.ncsa.illinois.edu
User your_username
PreferredAuthentications keyboard-interactive,password
ServerAliveInterval 60
ServerAliveCountMax 3
ForwardAgent no

tmux Workflow​

Use tmux to keep a session running between SSH connections:

  1. SSH to your pinned login node: ssh delta-ai-4
  2. Start a tmux session: tmux new -s work
  3. Do your work inside tmux
  4. Detach with Ctrl-B D when done
  5. Reattach later: SSH to the same node, then tmux attach -t work

If you are working from a local workstation, consider running a local tmux session with the SSH connection inside it. This gives you two layers of persistence: the local tmux survives terminal crashes, and the remote tmux survives SSH disconnections.

Login Node Limitations​

Login nodes are shared and have no GPUs. NCSA may terminate long-running or memory-intensive processes.

Use login nodes only for: editing files, submitting jobs, installing software, light compilation, and file transfers. All heavy computation must happen on compute nodes via Slurm.

Storage Layout​

MountPathQuotaPurpose
HOME/u/your_username100 GBSmall configs, SSH keys, bashrc
WORK-HDD/work/hdd/your_project/your_username1 TB shared per allocationPrimary workspace
WORK-NVME/work/nvme/your_project/1 TB shared per allocationFast workspace
PROJECTS/projects/your_project/500 GB shared per allocationShared project data
TMP/tmp3.9 TBNode-local burst buffer

These are NCSA's default quotas (opens in a new tab). Run quota to check the limits assigned to your allocation.

warning

Keep large builds, datasets, and conda environments in WORK storage to conserve the HOME quota. NCSA permits small scripts and software in HOME, but advises against using it for job I/O.

The /tmp directory on compute nodes is a fast local NVMe scratch space, but it is purged when your job ends. Copy any results you need to persistent storage before the job completes.

Check your quota with:

quota

Job Scheduling with Slurm​

Slurm is the job scheduler on DeltaAI. You request compute resources, Slurm allocates them when available, and your code runs on the assigned nodes.

Core Concepts​

Account: Your billing identifier, assigned to the ACCESS allocation. Replace your_account with the value supplied by the PI or allocation manager.

Partition: The queue your job enters. DeltaAI has two:

PartitionUseBilling Rate
ghx4Batch jobs1 SU per GPU-hour
ghx4-interactiveInteractive shells2 SU per GPU-hour

Service Unit (SU): One SU corresponds to one GH200 superchip reserved for one hour at the batch rate. Charges depend on the GPUs, CPU cores, and CPU memory reserved, even when your code does not use all of them. One superchip supplies 72 CPU cores and about 110 GB of usable CPU memory. A full node reserved for one hour costs 4 SU; the interactive partition applies a 2x charge factor. See NCSA's job accounting guide (opens in a new tab).

Interactive Sessions​

Use srun for interactive access to compute nodes:

srun --account=your_account --partition=ghx4-interactive \
--nodes=1 --gpus-per-node=1 --cpus-per-task=16 \
--mem=64G --time=00:30:00 --pty bash

Flag breakdown:

FlagMeaning
--accountBilling account
--partitionJob queue (interactive = 2x cost)
--nodes=1Number of nodes
--gpus-per-node=1GPUs to allocate (1 to 4)
--cpus-per-task=16CPU cores (72 per superchip, 288 per node)
--mem=64GMemory limit
--time=00:30:00Wall time (HH:MM:SS)
--pty bashOpen an interactive shell
info

If the system MOTD mentions a reservation requirement, add --reservation=update to your command. This routes your job to nodes running the latest OS image.

Batch Jobs​

Write a Slurm batch script and submit it with sbatch:

#!/bin/bash
#SBATCH --account=your_account
#SBATCH --partition=ghx4
#SBATCH --job-name=my-training
#SBATCH --nodes=1
#SBATCH --gpus-per-node=1
#SBATCH --cpus-per-task=16
#SBATCH --mem=64G
#SBATCH --time=04:00:00
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err

set -euo pipefail

source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh
conda activate your_env

cd /work/hdd/your_project/your_username
python train.py --epochs 100

Submit:

mkdir -p logs
sbatch train.slurm

The %x in the output path expands to the job name, %j to the job ID.

Job Monitoring​

# Check your jobs
squeue -u $USER

# Check all jobs in the partition
squeue -p ghx4

# Detailed job info
scontrol show job JOBID

# Node availability
sinfo -p ghx4

# Cancel a job
scancel JOBID

Multi-GPU Jobs​

Request multiple GPUs on a single node:

#SBATCH --gpus-per-node=4
#SBATCH --cpus-per-task=64

For PyTorch distributed training, use torchrun:

torchrun --nproc_per_node=4 train.py

For multi-node PyTorch jobs, reserve one launcher task per node:

#SBATCH --nodes=2
#SBATCH --gpus-per-node=4
#SBATCH --ntasks-per-node=1

Follow NCSA's multi-node PyTorch example (opens in a new tab) to configure srun, torchrun, and rendezvous across those nodes. The single-node command above does not configure multi-node training.

For MPI in a batch job, launch with srun. An interactive shell started by srun --pty cannot start another srun or mpirun inside it. For interactive MPI, obtain an allocation with salloc, then launch from the allocation shell:

salloc --account=your_account --partition=ghx4-interactive \
--nodes=1 --ntasks=4 --gpus-per-node=1 --time=00:30:00
srun ./my_mpi_program
exit

See NCSA's interactive MPI instructions (opens in a new tab).

Billing and Budget Planning​

ScenarioGPUsHoursRateCost (SU)
Interactive debugging10.52x1
Single-GPU training (batch)141x4
Full-node training (batch)441x16
Multi-node (2 nodes, batch)841x32

Check your remaining balance:

accounts

Software Environment​

Module System​

DeltaAI uses Lmod for system-provided software modules:

# List available modules
module avail

# Search for a module
module spider cuda

# Load a module
module load cudatoolkit

# Show loaded modules
module list
warning

Available modules differ between login nodes and compute nodes. Always verify module availability on the node type where you plan to use them.

Conda Setup​

User software is managed through conda. If miniconda is not yet installed:

curl -L -o /tmp/mc.sh https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-aarch64.sh
bash /tmp/mc.sh -b -p /work/hdd/your_project/your_username/miniconda3
source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh

Create and activate an environment:

conda create -n myenv -y python
conda activate myenv
conda install -c conda-forge numpy scipy pandas

Add to your ~/.bashrc to auto-source conda on login:

source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh

Compiler Selection​

NCSA recommends its Cray compiler wrappers and programming-environment modules. The IOWarp example below uses GCC 13 directly. Check that both paths exist before using that example:

# Select GCC 13 explicitly for this build
cmake -DCMAKE_C_COMPILER=/usr/bin/gcc-13 \
-DCMAKE_CXX_COMPILER=/usr/bin/g++-13 \
-B build

GCC 7 has experimental C++17 support (opens in a new tab), but may lack features a project requires. Consult your project's compiler requirements and NCSA's programming environment guide (opens in a new tab).

Python and PyTorch​

Python is available through conda. For an aarch64 PyTorch environment targeting CUDA 12.8, use a compatible driver and the CUDA-specific index:

conda activate myenv
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128

Verify on a compute node (GPUs are not available on login nodes):

srun --account=your_account --partition=ghx4-interactive \
--gpus-per-node=1 --time=00:05:00 --pty bash -c \
'source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh && conda activate myenv && \
python -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"'

Node.js​

Install via conda for aarch64 compatibility:

conda install -c conda-forge nodejs
node --version
npm --version

CMake with Conda Dependencies​

Start with -DCMAKE_PREFIX_PATH="$CONDA_PREFIX" so CMake can find conda-installed packages. If a dependency is still not found, explicit include and library paths may help:

cmake \
-DCMAKE_C_COMPILER=/usr/bin/gcc-13 \
-DCMAKE_CXX_COMPILER=/usr/bin/g++-13 \
-DCMAKE_C_FLAGS="-I$CONDA_PREFIX/include" \
-DCMAKE_CXX_FLAGS="-I$CONDA_PREFIX/include" \
-DCMAKE_EXE_LINKER_FLAGS="-L$CONDA_PREFIX/lib" \
-DCMAKE_SHARED_LINKER_FLAGS="-L$CONDA_PREFIX/lib" \
-DCMAKE_PREFIX_PATH=$CONDA_PREFIX \
-B build -G Ninja

Common Gotchas​

ARM aarch64 and memory page size

DeltaAI uses NVIDIA Grace CPUs (ARM aarch64). Check memory page size with getconf PAGESIZE when troubleshooting an allocator error:

  • Allocator compatibility: A jemalloc build can fail with "Unsupported system page size" if it does not support the node's page size.
  • x86 binaries fail with "Exec format error" or produce no output. Always verify that binaries are compiled for aarch64.
  • Workaround: Use conda packages (which provide native aarch64 builds) or compile from source on DeltaAI.
OSTYPE detection scripts

Some detection scripts accept only linux-gnu. If the local $OSTYPE value is linux, those scripts may reject the platform.

Fix: Patch detection scripts to accept both linux and linux-gnu, or use uname -s instead of $OSTYPE.

msgpack CMake naming mismatch

Some msgpack-cxx packages install msgpack-cxx-config.cmake, while a consuming project may call find_package(msgpack) and expect msgpackConfig.cmake. Check the installed files and the project's CMake code before applying this workaround.

Fix: Create symlinks in your conda environment:

mkdir -p $CONDA_PREFIX/lib/cmake/msgpack
ln -sf $CONDA_PREFIX/lib/cmake/msgpack-cxx/msgpack-cxx-config.cmake \
$CONDA_PREFIX/lib/cmake/msgpack/msgpackConfig.cmake
ln -sf $CONDA_PREFIX/lib/cmake/msgpack-cxx/msgpack-cxx-config-version.cmake \
$CONDA_PREFIX/lib/cmake/msgpack/msgpackConfigVersion.cmake
ln -sf $CONDA_PREFIX/lib/cmake/msgpack-cxx/msgpack-cxx-targets.cmake \
$CONDA_PREFIX/lib/cmake/msgpack/msgpack-cxx-targets.cmake
No io_uring support

Check the compute node's kernel and your project's requirements before enabling io_uring. An unsupported configuration may fail at build time or runtime.

Fix: Disable io_uring at build time. For IOWarp, pass -DWRP_CORE_ENABLE_IO_URING=OFF to CMake.

Reservation flag during system updates

During system updates, compute nodes may be reimaged and placed in a Slurm reservation. The login MOTD will indicate when this is required.

Fix: Add --reservation=update to your srun or sbatch commands. Remove it once the MOTD no longer mentions the reservation.

Building IOWarp on DeltaAI​

The following build example uses conda dependencies and explicit GCC 13 paths. Check those paths and the CMake options against the CLIO Core revision you are building; this guide does not establish that the recipe works with every release.

Install Dependencies​

source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh
conda create -n iowarp -y python
conda activate iowarp
conda install -y -c conda-forge \
boost-cpp yaml-cpp zeromq cereal catch2 ninja cmake pkg-config \
msgpack-cxx msgpack-c hdf5

Apply the msgpack CMake workaround only if the installed package and the project's find_package call use different names (see Common Gotchas above).

Clone and Build​

cd /work/hdd/your_project/your_username
git clone --recurse-submodules https://github.com/iowarp/clio-core.git
cd clio-core

If install.sh rejects the platform, inspect its operating-system check. Update the check to accept the value reported by the node, or use uname -s; avoid replacing unrelated strings throughout the script.

Configure with CMake:

cmake \
-DCMAKE_BUILD_TYPE=Release \
-DCMAKE_C_COMPILER=/usr/bin/gcc-13 \
-DCMAKE_CXX_COMPILER=/usr/bin/g++-13 \
-DCMAKE_C_FLAGS="-I$CONDA_PREFIX/include" \
-DCMAKE_CXX_FLAGS="-I$CONDA_PREFIX/include" \
-DCMAKE_EXE_LINKER_FLAGS="-L$CONDA_PREFIX/lib" \
-DCMAKE_SHARED_LINKER_FLAGS="-L$CONDA_PREFIX/lib" \
-DCMAKE_PREFIX_PATH=$CONDA_PREFIX \
-DCMAKE_INSTALL_PREFIX=$CONDA_PREFIX \
-DWRP_CORE_ENABLE_RUNTIME=ON \
-DWRP_CORE_ENABLE_CTE=ON \
-DWRP_CORE_ENABLE_CAE=ON \
-DWRP_CORE_ENABLE_CEE=ON \
-DWRP_CORE_ENABLE_TESTS=OFF \
-DWRP_CORE_ENABLE_BENCHMARKS=OFF \
-DWRP_CORE_ENABLE_PYTHON=OFF \
-DWRP_CORE_ENABLE_MPI=OFF \
-DWRP_CORE_ENABLE_IO_URING=OFF \
-DWRP_CORE_ENABLE_ZMQ=ON \
-DWRP_CORE_ENABLE_CEREAL=ON \
-DWRP_CORE_ENABLE_HDF5=ON \
-Wno-dev -B build -G Ninja

Build and install:

cmake --build build -j16
cmake --install build

Verify Installation​

ls $CONDA_PREFIX/iowarp_core/bin/chimaera
$CONDA_PREFIX/iowarp_core/bin/chimaera --help

Web Interfaces​

Acknowledging DeltaAI​

Use NCSA's current acknowledgment instructions (opens in a new tab) for publications that use DeltaAI. Include the allocation that supported the work and the acknowledgment required by its allocation program.

See the FAQ or the Job Templates page for Slurm script examples.