DeltaAI User Guide
This guide assumes familiarity with Linux, SSH, and job scheduling basics. If you need a refresher, start with the HPC Foundations tutorials. For DeltaAI hardware specs and other GRC computing resources, see the Resources page.
Before You Start
- DeltaAI runs ARM aarch64 CPUs (NVIDIA Grace) with NVIDIA H100 GPUs (GH200 superchips). This is not x86.
- General user SSH access requires an NCSA password and Duo MFA. SSH keys require an approved Science Gateway account.
- Submit jobs through Slurm. Use
srunfor an interactive shell andsbatchfor batch jobs. For interactive MPI, usesalloc, then launch withsrun. - Interactive sessions cost 2x the batch rate. Prefer batch jobs for real work.
- Install user software with conda or use NCSA modules. Check compiler requirements before building.
- Keep all work in
/work/hdd/your_project/your_username/. Home directory is only 100 GB.
Introduction
DeltaAI is a GPU-focused supercomputer operated by the National Center for Supercomputing Applications (NCSA) at the University of Illinois. It is accessed through the ACCESS (opens in a new tab) program, which provides compute allocations to research groups across the U.S. NCSA maintains the official DeltaAI user documentation (opens in a new tab); consult it for current system configuration and policies.
GRC researchers use DeltaAI for GPU-accelerated workloads including machine learning training, high-performance I/O benchmarking, and systems research. The cluster consists of 152 quad-GPU nodes based on the NVIDIA GH200 Grace Hopper Superchip architecture.
Confirm your ACCESS project, Slurm account, and storage project code with the PI or allocation manager. The examples below use placeholders for allocation-specific values.
Hardware Overview
Each DeltaAI compute node contains four GH200 superchips. Each superchip pairs an NVIDIA Grace ARM CPU with an H100 GPU connected via NVLink-C2C.
| Component | Specification |
|---|---|
| GPU | 4x NVIDIA GH200 (H100 96 GB HBM3 each) |
| CPU | 4x NVIDIA Grace (72 ARM cores each) |
| Cores/Node | 288 ARM Neoverse V2 cores |
| Memory | 480 GB LPDDR5 (CPU) + 384 GB HBM3 (GPU) |
| Local Storage | 3.9 TB NVMe SSD |
| Network | 4x 200 GbE HPE Slingshot 11 |
| Total Nodes | 152 |
Hardware values follow NCSA's system architecture guide (opens in a new tab).
GPU Specifications
| Property | Value |
|---|---|
| Architecture | Hopper (SM 9.0) |
| VRAM per GPU | 96 GB HBM3 |
| GPUs per Node | 4 |
| Total Cluster GPUs | 608 |
Architecture Notes
DeltaAI uses ARM aarch64 processors, not x86. This has several practical consequences:
- Use aarch64 builds of tools and libraries. An x86_64 executable cannot run natively on a Grace CPU and may report "Exec format error".
- Check the memory page size with
getconf PAGESIZE. Tools whose allocator does not support the reported page size may fail with "Unsupported system page size." - Check the operating system with
cat /etc/os-release. Shell detection scripts should accept the local$OSTYPEvalue or useuname -s. - NCSA provides the Cray compiler wrappers
cc,CC, andftn. Use a programming environment appropriate for your project; see NCSA's compiler guidance (opens in a new tab).
Accessing the Cluster
NCSA Account Setup
Before you can access DeltaAI, you must be added to the ACCESS allocation. Contact the PI or an allocation manager to add your NCSA username to the project.
Required accounts:
- An ACCESS account (opens in a new tab), linked to your institution
- An NCSA Identity (opens in a new tab) linked to the allocation; check your ACCESS profile for the NCSA username
- NCSA Duo MFA (opens in a new tab) enrollment for two-factor authentication
SSH Configuration
General user access requires an NCSA password and Duo MFA. SSH keys are reserved for approved Science Gateway accounts; see NCSA's login instructions (opens in a new tab).
Add the following to your ~/.ssh/config:
Host delta-ai
HostName dtai-login.delta.ncsa.illinois.edu
User your_username
PreferredAuthentications keyboard-interactive,password
ServerAliveInterval 60
ServerAliveCountMax 3
ForwardAgent no
Then connect with:
ssh delta-ai
You will be prompted for your NCSA password, then a Duo verification (push notification, phone call, or recovery code).
Login Node Persistence
DeltaAI has four login nodes. The delta-ai hostname round-robins between them. If you start a tmux session on one login node, you must reconnect to the same node to reattach.
Pin your SSH connection to a specific node:
Host delta-ai-4
HostName gh-login04.delta.ncsa.illinois.edu
User your_username
PreferredAuthentications keyboard-interactive,password
ServerAliveInterval 60
ServerAliveCountMax 3
ForwardAgent no
tmux Workflow
Use tmux to keep a session running between SSH connections:
- SSH to your pinned login node:
ssh delta-ai-4 - Start a tmux session:
tmux new -s work - Do your work inside tmux
- Detach with
Ctrl-B Dwhen done - Reattach later: SSH to the same node, then
tmux attach -t work
If you are working from a local workstation, consider running a local tmux session with the SSH connection inside it. This gives you two layers of persistence: the local tmux survives terminal crashes, and the remote tmux survives SSH disconnections.
Login Node Limitations
Login nodes are shared and have no GPUs. NCSA may terminate long-running or memory-intensive processes.
Use login nodes only for: editing files, submitting jobs, installing software, light compilation, and file transfers. All heavy computation must happen on compute nodes via Slurm.
Storage Layout
| Mount | Path | Quota | Purpose |
|---|---|---|---|
| HOME | /u/your_username | 100 GB | Small configs, SSH keys, bashrc |
| WORK-HDD | /work/hdd/your_project/your_username | 1 TB shared per allocation | Primary workspace |
| WORK-NVME | /work/nvme/your_project/ | 1 TB shared per allocation | Fast workspace |
| PROJECTS | /projects/your_project/ | 500 GB shared per allocation | Shared project data |
| TMP | /tmp | 3.9 TB | Node-local burst buffer |
These are NCSA's default quotas (opens in a new tab). Run quota to check the limits assigned to your allocation.
Keep large builds, datasets, and conda environments in WORK storage to conserve the HOME quota. NCSA permits small scripts and software in HOME, but advises against using it for job I/O.
The /tmp directory on compute nodes is a fast local NVMe scratch space, but it is purged when your job ends. Copy any results you need to persistent storage before the job completes.
Check your quota with:
quota
Job Scheduling with Slurm
Slurm is the job scheduler on DeltaAI. You request compute resources, Slurm allocates them when available, and your code runs on the assigned nodes.
Core Concepts
Account: Your billing identifier, assigned to the ACCESS allocation. Replace your_account with the value supplied by the PI or allocation manager.
Partition: The queue your job enters. DeltaAI has two:
| Partition | Use | Billing Rate |
|---|---|---|
ghx4 | Batch jobs | 1 SU per GPU-hour |
ghx4-interactive | Interactive shells | 2 SU per GPU-hour |
Service Unit (SU): One SU corresponds to one GH200 superchip reserved for one hour at the batch rate. Charges depend on the GPUs, CPU cores, and CPU memory reserved, even when your code does not use all of them. One superchip supplies 72 CPU cores and about 110 GB of usable CPU memory. A full node reserved for one hour costs 4 SU; the interactive partition applies a 2x charge factor. See NCSA's job accounting guide (opens in a new tab).
Interactive Sessions
Use srun for interactive access to compute nodes:
srun --account=your_account --partition=ghx4-interactive \
--nodes=1 --gpus-per-node=1 --cpus-per-task=16 \
--mem=64G --time=00:30:00 --pty bash
Flag breakdown:
| Flag | Meaning |
|---|---|
--account | Billing account |
--partition | Job queue (interactive = 2x cost) |
--nodes=1 | Number of nodes |
--gpus-per-node=1 | GPUs to allocate (1 to 4) |
--cpus-per-task=16 | CPU cores (72 per superchip, 288 per node) |
--mem=64G | Memory limit |
--time=00:30:00 | Wall time (HH:MM:SS) |
--pty bash | Open an interactive shell |
If the system MOTD mentions a reservation requirement, add --reservation=update to your command. This routes your job to nodes running the latest OS image.
Batch Jobs
Write a Slurm batch script and submit it with sbatch:
#!/bin/bash
#SBATCH --account=your_account
#SBATCH --partition=ghx4
#SBATCH --job-name=my-training
#SBATCH --nodes=1
#SBATCH --gpus-per-node=1
#SBATCH --cpus-per-task=16
#SBATCH --mem=64G
#SBATCH --time=04:00:00
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err
set -euo pipefail
source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh
conda activate your_env
cd /work/hdd/your_project/your_username
python train.py --epochs 100
Submit:
mkdir -p logs
sbatch train.slurm
The %x in the output path expands to the job name, %j to the job ID.
Job Monitoring
# Check your jobs
squeue -u $USER
# Check all jobs in the partition
squeue -p ghx4
# Detailed job info
scontrol show job JOBID
# Node availability
sinfo -p ghx4
# Cancel a job
scancel JOBID
Multi-GPU Jobs
Request multiple GPUs on a single node:
#SBATCH --gpus-per-node=4
#SBATCH --cpus-per-task=64
For PyTorch distributed training, use torchrun:
torchrun --nproc_per_node=4 train.py
For multi-node PyTorch jobs, reserve one launcher task per node:
#SBATCH --nodes=2
#SBATCH --gpus-per-node=4
#SBATCH --ntasks-per-node=1
Follow NCSA's multi-node PyTorch example (opens in a new tab) to configure srun, torchrun, and rendezvous across those nodes. The single-node command above does not configure multi-node training.
For MPI in a batch job, launch with srun. An interactive shell started by srun --pty cannot start another srun or mpirun inside it. For interactive MPI, obtain an allocation with salloc, then launch from the allocation shell:
salloc --account=your_account --partition=ghx4-interactive \
--nodes=1 --ntasks=4 --gpus-per-node=1 --time=00:30:00
srun ./my_mpi_program
exit
See NCSA's interactive MPI instructions (opens in a new tab).
Billing and Budget Planning
| Scenario | GPUs | Hours | Rate | Cost (SU) |
|---|---|---|---|---|
| Interactive debugging | 1 | 0.5 | 2x | 1 |
| Single-GPU training (batch) | 1 | 4 | 1x | 4 |
| Full-node training (batch) | 4 | 4 | 1x | 16 |
| Multi-node (2 nodes, batch) | 8 | 4 | 1x | 32 |
Check your remaining balance:
accounts
Software Environment
Module System
DeltaAI uses Lmod for system-provided software modules:
# List available modules
module avail
# Search for a module
module spider cuda
# Load a module
module load cudatoolkit
# Show loaded modules
module list
Available modules differ between login nodes and compute nodes. Always verify module availability on the node type where you plan to use them.
Conda Setup
User software is managed through conda. If miniconda is not yet installed:
curl -L -o /tmp/mc.sh https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-aarch64.sh
bash /tmp/mc.sh -b -p /work/hdd/your_project/your_username/miniconda3
source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh
Create and activate an environment:
conda create -n myenv -y python
conda activate myenv
conda install -c conda-forge numpy scipy pandas
Add to your ~/.bashrc to auto-source conda on login:
source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh
Compiler Selection
NCSA recommends its Cray compiler wrappers and programming-environment modules. The IOWarp example below uses GCC 13 directly. Check that both paths exist before using that example:
# Select GCC 13 explicitly for this build
cmake -DCMAKE_C_COMPILER=/usr/bin/gcc-13 \
-DCMAKE_CXX_COMPILER=/usr/bin/g++-13 \
-B build
GCC 7 has experimental C++17 support (opens in a new tab), but may lack features a project requires. Consult your project's compiler requirements and NCSA's programming environment guide (opens in a new tab).
Python and PyTorch
Python is available through conda. For an aarch64 PyTorch environment targeting CUDA 12.8, use a compatible driver and the CUDA-specific index:
conda activate myenv
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
Verify on a compute node (GPUs are not available on login nodes):
srun --account=your_account --partition=ghx4-interactive \
--gpus-per-node=1 --time=00:05:00 --pty bash -c \
'source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh && conda activate myenv && \
python -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"'
Node.js
Install via conda for aarch64 compatibility:
conda install -c conda-forge nodejs
node --version
npm --version
CMake with Conda Dependencies
Start with -DCMAKE_PREFIX_PATH="$CONDA_PREFIX" so CMake can find conda-installed packages. If a dependency is still not found, explicit include and library paths may help:
cmake \
-DCMAKE_C_COMPILER=/usr/bin/gcc-13 \
-DCMAKE_CXX_COMPILER=/usr/bin/g++-13 \
-DCMAKE_C_FLAGS="-I$CONDA_PREFIX/include" \
-DCMAKE_CXX_FLAGS="-I$CONDA_PREFIX/include" \
-DCMAKE_EXE_LINKER_FLAGS="-L$CONDA_PREFIX/lib" \
-DCMAKE_SHARED_LINKER_FLAGS="-L$CONDA_PREFIX/lib" \
-DCMAKE_PREFIX_PATH=$CONDA_PREFIX \
-B build -G Ninja
Common Gotchas
ARM aarch64 and memory page size
DeltaAI uses NVIDIA Grace CPUs (ARM aarch64). Check memory page size with getconf PAGESIZE when troubleshooting an allocator error:
- Allocator compatibility: A jemalloc build can fail with "Unsupported system page size" if it does not support the node's page size.
- x86 binaries fail with "Exec format error" or produce no output. Always verify that binaries are compiled for aarch64.
- Workaround: Use conda packages (which provide native aarch64 builds) or compile from source on DeltaAI.
OSTYPE detection scripts
Some detection scripts accept only linux-gnu. If the local $OSTYPE value is linux, those scripts may reject the platform.
Fix: Patch detection scripts to accept both linux and linux-gnu, or use uname -s instead of $OSTYPE.
msgpack CMake naming mismatch
Some msgpack-cxx packages install msgpack-cxx-config.cmake, while a consuming project may call find_package(msgpack) and expect msgpackConfig.cmake. Check the installed files and the project's CMake code before applying this workaround.
Fix: Create symlinks in your conda environment:
mkdir -p $CONDA_PREFIX/lib/cmake/msgpack
ln -sf $CONDA_PREFIX/lib/cmake/msgpack-cxx/msgpack-cxx-config.cmake \
$CONDA_PREFIX/lib/cmake/msgpack/msgpackConfig.cmake
ln -sf $CONDA_PREFIX/lib/cmake/msgpack-cxx/msgpack-cxx-config-version.cmake \
$CONDA_PREFIX/lib/cmake/msgpack/msgpackConfigVersion.cmake
ln -sf $CONDA_PREFIX/lib/cmake/msgpack-cxx/msgpack-cxx-targets.cmake \
$CONDA_PREFIX/lib/cmake/msgpack/msgpack-cxx-targets.cmake
No io_uring support
Check the compute node's kernel and your project's requirements before enabling io_uring. An unsupported configuration may fail at build time or runtime.
Fix: Disable io_uring at build time. For IOWarp, pass -DWRP_CORE_ENABLE_IO_URING=OFF to CMake.
Reservation flag during system updates
During system updates, compute nodes may be reimaged and placed in a Slurm reservation. The login MOTD will indicate when this is required.
Fix: Add --reservation=update to your srun or sbatch commands. Remove it once the MOTD no longer mentions the reservation.
Building IOWarp on DeltaAI
The following build example uses conda dependencies and explicit GCC 13 paths. Check those paths and the CMake options against the CLIO Core revision you are building; this guide does not establish that the recipe works with every release.
Install Dependencies
source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh
conda create -n iowarp -y python
conda activate iowarp
conda install -y -c conda-forge \
boost-cpp yaml-cpp zeromq cereal catch2 ninja cmake pkg-config \
msgpack-cxx msgpack-c hdf5
Apply the msgpack CMake workaround only if the installed package and the project's find_package call use different names (see Common Gotchas above).
Clone and Build
cd /work/hdd/your_project/your_username
git clone --recurse-submodules https://github.com/iowarp/clio-core.git
cd clio-core
If install.sh rejects the platform, inspect its operating-system check. Update the check to accept the value reported by the node, or use uname -s; avoid replacing unrelated strings throughout the script.
Configure with CMake:
cmake \
-DCMAKE_BUILD_TYPE=Release \
-DCMAKE_C_COMPILER=/usr/bin/gcc-13 \
-DCMAKE_CXX_COMPILER=/usr/bin/g++-13 \
-DCMAKE_C_FLAGS="-I$CONDA_PREFIX/include" \
-DCMAKE_CXX_FLAGS="-I$CONDA_PREFIX/include" \
-DCMAKE_EXE_LINKER_FLAGS="-L$CONDA_PREFIX/lib" \
-DCMAKE_SHARED_LINKER_FLAGS="-L$CONDA_PREFIX/lib" \
-DCMAKE_PREFIX_PATH=$CONDA_PREFIX \
-DCMAKE_INSTALL_PREFIX=$CONDA_PREFIX \
-DWRP_CORE_ENABLE_RUNTIME=ON \
-DWRP_CORE_ENABLE_CTE=ON \
-DWRP_CORE_ENABLE_CAE=ON \
-DWRP_CORE_ENABLE_CEE=ON \
-DWRP_CORE_ENABLE_TESTS=OFF \
-DWRP_CORE_ENABLE_BENCHMARKS=OFF \
-DWRP_CORE_ENABLE_PYTHON=OFF \
-DWRP_CORE_ENABLE_MPI=OFF \
-DWRP_CORE_ENABLE_IO_URING=OFF \
-DWRP_CORE_ENABLE_ZMQ=ON \
-DWRP_CORE_ENABLE_CEREAL=ON \
-DWRP_CORE_ENABLE_HDF5=ON \
-Wno-dev -B build -G Ninja
Build and install:
cmake --build build -j16
cmake --install build
Verify Installation
ls $CONDA_PREFIX/iowarp_core/bin/chimaera
$CONDA_PREFIX/iowarp_core/bin/chimaera --help
Web Interfaces
- DeltaAI Open OnDemand: https://gh-ondemand.delta.ncsa.illinois.edu/ (opens in a new tab)
- ACCESS Portal: https://allocations.access-ci.org (opens in a new tab)
- NCSA Identity Management: https://identity.ncsa.illinois.edu (opens in a new tab)
- NCSA Duo MFA: https://duo.security.ncsa.illinois.edu (opens in a new tab)
- Support: http://help.ncsa.illinois.edu (opens in a new tab)
Acknowledging DeltaAI
Use NCSA's current acknowledgment instructions (opens in a new tab) for publications that use DeltaAI. Include the allocation that supported the work and the acknowledgment required by its allocation program.
Related Guides
See the FAQ or the Job Templates page for Slurm script examples.