Skip to main content

Job Templates

Slurm script examples for common DeltaAI workloads. Replace the account, username, environment, and script paths before submitting. Check NCSA's current job policies (opens in a new tab) and billing rules (opens in a new tab).

Before submitting any batch job, create the logs directory:

mkdir -p logs

Interactive GPU Session​

For quick debugging, testing, and exploration. Launches an interactive shell on a compute node with GPU access.

This is not a batch script. Run it directly:

srun --account=your_account --partition=ghx4-interactive \
--nodes=1 --gpus-per-node=1 --cpus-per-task=16 \
--mem=64G --time=00:30:00 --pty bash

Cost: approximately 1 SU per 30 minutes (2x interactive rate).

Once on the compute node, activate your environment and verify GPU access:

source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh
conda activate your_env
nvidia-smi
python -c "import torch; print(torch.cuda.is_available())"

To request more GPUs or time:

# 2 GPUs for 1 hour
srun --account=your_account --partition=ghx4-interactive \
--nodes=1 --gpus-per-node=2 --cpus-per-task=32 \
--mem=128G --time=01:00:00 --pty bash

Single-GPU Training (Batch)​

Training job on one GPU with conda activation and logging. The signal handler is a placeholder; it does not save a training checkpoint. Implement checkpointing in your training code before relying on recovery after a timeout.

#!/bin/bash
#SBATCH --account=your_account
#SBATCH --partition=ghx4
#SBATCH --job-name=gpu-train
#SBATCH --nodes=1
#SBATCH --gpus-per-node=1
#SBATCH --cpus-per-task=16
#SBATCH --mem=64G
#SBATCH --time=04:00:00
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err

set -euo pipefail

# Configuration (edit these)
TRAIN_SCRIPT="train.py"
TRAIN_ARGS=""
CONDA_ENV="your_env"
WORK_DIR="/work/hdd/your_project/your_username"

# Signal handler placeholder
# Implement checkpointing for your training process here.
cleanup() {
echo "[$(date)] SIGTERM received."
# Add checkpoint logic, for example:
# python save_checkpoint.py --output $WORK_DIR/checkpoints/
echo "[$(date)] Cleanup complete."
exit 0
}
trap cleanup SIGTERM

# Setup
echo "[$(date)] Job $SLURM_JOB_ID started on $(hostname)"

source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh
conda activate "$CONDA_ENV"

cd "$WORK_DIR"
mkdir -p logs checkpoints

nvidia-smi --query-gpu=name,memory.total --format=csv,noheader

# Run
echo "[$(date)] Starting: $TRAIN_SCRIPT $TRAIN_ARGS"
python "$TRAIN_SCRIPT" $TRAIN_ARGS
echo "[$(date)] Completed."

Cost: 4 SU (1 GPU, 4 hours, 1x batch rate).

Submit with:

sbatch train.slurm

Multi-GPU Training (Batch)​

Full-node job with all 4 GH200 GPUs. Uses torchrun for PyTorch distributed training.

#!/bin/bash
#SBATCH --account=your_account
#SBATCH --partition=ghx4
#SBATCH --job-name=multi-gpu
#SBATCH --nodes=1
#SBATCH --gpus-per-node=4
#SBATCH --cpus-per-task=64
#SBATCH --mem=240G
#SBATCH --time=04:00:00
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err

set -euo pipefail

# Configuration (edit these)
TRAIN_SCRIPT="train.py"
TRAIN_ARGS=""
CONDA_ENV="your_env"
WORK_DIR="/work/hdd/your_project/your_username"
NUM_GPUS=4

# NCCL logging. Keep the site communication defaults.
export NCCL_DEBUG=WARN

# Signal handler placeholder
cleanup() {
echo "[$(date)] SIGTERM received. Exiting."
exit 0
}
trap cleanup SIGTERM

# Setup
echo "[$(date)] Job $SLURM_JOB_ID on $(hostname), $NUM_GPUS GPUs"

source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh
conda activate "$CONDA_ENV"

cd "$WORK_DIR"
mkdir -p logs checkpoints

nvidia-smi --query-gpu=index,name,memory.total --format=csv,noheader
nvidia-smi topo -m 2>/dev/null || true

# Multi-GPU launch
echo "[$(date)] Launching torchrun with $NUM_GPUS GPUs"
torchrun \
--nproc_per_node=$NUM_GPUS \
--master_port=29500 \
"$TRAIN_SCRIPT" $TRAIN_ARGS

echo "[$(date)] Completed."

Cost: 16 SU (4 GPUs, 4 hours, 1x batch rate).

info

Your training script must support PyTorch's distributed training API (torch.distributed). At minimum, it should call torch.distributed.init_process_group() and use DistributedDataParallel. See the PyTorch DDP tutorial (opens in a new tab) for details.

CPU-Only Job (Batch)​

For preprocessing, file operations, or compilation that does not use a GPU. DeltaAI's smallest allocatable unit is one GH200 superchip, so this example reserves one even though the script runs CPU code.

#!/bin/bash
#SBATCH --account=your_account
#SBATCH --partition=ghx4
#SBATCH --job-name=cpu-work
#SBATCH --nodes=1
#SBATCH --gpus-per-node=1
#SBATCH --cpus-per-task=32
#SBATCH --mem=64G
#SBATCH --time=02:00:00
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err

set -euo pipefail

# Configuration (edit these)
RUN_SCRIPT="process.py"
RUN_ARGS=""
CONDA_ENV="your_env"
WORK_DIR="/work/hdd/your_project/your_username"

# Setup
echo "[$(date)] Job $SLURM_JOB_ID on $(hostname) (CPU-only)"

source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh
conda activate "$CONDA_ENV"

cd "$WORK_DIR"
mkdir -p logs

# Run
echo "[$(date)] Running: $RUN_SCRIPT $RUN_ARGS"
python "$RUN_SCRIPT" $RUN_ARGS
echo "[$(date)] Completed."

Cost: 2 SU for two hours at the batch rate, provided the CPU and memory request fits within one GH200 superchip. Charges apply to reserved resources even when the GPU is idle.

Customization Notes​

Reservation flag: If the system MOTD mentions --reservation=update, add this line to the #SBATCH directives:

#SBATCH --reservation=update

Email notifications: Add these directives to receive email when jobs start, end, or fail:

#SBATCH --mail-type=BEGIN,END,FAIL
#SBATCH --mail-user=your_email@iit.edu

Job arrays: For parameter sweeps, use Slurm job arrays:

#SBATCH --array=0-9

PARAM_FILE="params.txt"
PARAM=$(sed -n "$((SLURM_ARRAY_TASK_ID + 1))p" "$PARAM_FILE")
python train.py --config "$PARAM"

Time limits: Adjust --time based on measured runtime. A shorter request may fit a backfill window, but does not guarantee an earlier start. Allow time to save results before the limit expires.