Job Templates
Slurm script examples for common DeltaAI workloads. Replace the account, username, environment, and script paths before submitting. Check NCSA's current job policies (opens in a new tab) and billing rules (opens in a new tab).
Before submitting any batch job, create the logs directory:
mkdir -p logs
Interactive GPU Session
For quick debugging, testing, and exploration. Launches an interactive shell on a compute node with GPU access.
This is not a batch script. Run it directly:
srun --account=your_account --partition=ghx4-interactive \
--nodes=1 --gpus-per-node=1 --cpus-per-task=16 \
--mem=64G --time=00:30:00 --pty bash
Cost: approximately 1 SU per 30 minutes (2x interactive rate).
Once on the compute node, activate your environment and verify GPU access:
source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh
conda activate your_env
nvidia-smi
python -c "import torch; print(torch.cuda.is_available())"
To request more GPUs or time:
# 2 GPUs for 1 hour
srun --account=your_account --partition=ghx4-interactive \
--nodes=1 --gpus-per-node=2 --cpus-per-task=32 \
--mem=128G --time=01:00:00 --pty bash
Single-GPU Training (Batch)
Training job on one GPU with conda activation and logging. The signal handler is a placeholder; it does not save a training checkpoint. Implement checkpointing in your training code before relying on recovery after a timeout.
#!/bin/bash
#SBATCH --account=your_account
#SBATCH --partition=ghx4
#SBATCH --job-name=gpu-train
#SBATCH --nodes=1
#SBATCH --gpus-per-node=1
#SBATCH --cpus-per-task=16
#SBATCH --mem=64G
#SBATCH --time=04:00:00
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err
set -euo pipefail
# Configuration (edit these)
TRAIN_SCRIPT="train.py"
TRAIN_ARGS=""
CONDA_ENV="your_env"
WORK_DIR="/work/hdd/your_project/your_username"
# Signal handler placeholder
# Implement checkpointing for your training process here.
cleanup() {
echo "[$(date)] SIGTERM received."
# Add checkpoint logic, for example:
# python save_checkpoint.py --output $WORK_DIR/checkpoints/
echo "[$(date)] Cleanup complete."
exit 0
}
trap cleanup SIGTERM
# Setup
echo "[$(date)] Job $SLURM_JOB_ID started on $(hostname)"
source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh
conda activate "$CONDA_ENV"
cd "$WORK_DIR"
mkdir -p logs checkpoints
nvidia-smi --query-gpu=name,memory.total --format=csv,noheader
# Run
echo "[$(date)] Starting: $TRAIN_SCRIPT $TRAIN_ARGS"
python "$TRAIN_SCRIPT" $TRAIN_ARGS
echo "[$(date)] Completed."
Cost: 4 SU (1 GPU, 4 hours, 1x batch rate).
Submit with:
sbatch train.slurm
Multi-GPU Training (Batch)
Full-node job with all 4 GH200 GPUs. Uses torchrun for PyTorch distributed training.
#!/bin/bash
#SBATCH --account=your_account
#SBATCH --partition=ghx4
#SBATCH --job-name=multi-gpu
#SBATCH --nodes=1
#SBATCH --gpus-per-node=4
#SBATCH --cpus-per-task=64
#SBATCH --mem=240G
#SBATCH --time=04:00:00
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err
set -euo pipefail
# Configuration (edit these)
TRAIN_SCRIPT="train.py"
TRAIN_ARGS=""
CONDA_ENV="your_env"
WORK_DIR="/work/hdd/your_project/your_username"
NUM_GPUS=4
# NCCL logging. Keep the site communication defaults.
export NCCL_DEBUG=WARN
# Signal handler placeholder
cleanup() {
echo "[$(date)] SIGTERM received. Exiting."
exit 0
}
trap cleanup SIGTERM
# Setup
echo "[$(date)] Job $SLURM_JOB_ID on $(hostname), $NUM_GPUS GPUs"
source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh
conda activate "$CONDA_ENV"
cd "$WORK_DIR"
mkdir -p logs checkpoints
nvidia-smi --query-gpu=index,name,memory.total --format=csv,noheader
nvidia-smi topo -m 2>/dev/null || true
# Multi-GPU launch
echo "[$(date)] Launching torchrun with $NUM_GPUS GPUs"
torchrun \
--nproc_per_node=$NUM_GPUS \
--master_port=29500 \
"$TRAIN_SCRIPT" $TRAIN_ARGS
echo "[$(date)] Completed."
Cost: 16 SU (4 GPUs, 4 hours, 1x batch rate).
Your training script must support PyTorch's distributed training API (torch.distributed). At minimum, it should call torch.distributed.init_process_group() and use DistributedDataParallel. See the PyTorch DDP tutorial (opens in a new tab) for details.
CPU-Only Job (Batch)
For preprocessing, file operations, or compilation that does not use a GPU. DeltaAI's smallest allocatable unit is one GH200 superchip, so this example reserves one even though the script runs CPU code.
#!/bin/bash
#SBATCH --account=your_account
#SBATCH --partition=ghx4
#SBATCH --job-name=cpu-work
#SBATCH --nodes=1
#SBATCH --gpus-per-node=1
#SBATCH --cpus-per-task=32
#SBATCH --mem=64G
#SBATCH --time=02:00:00
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err
set -euo pipefail
# Configuration (edit these)
RUN_SCRIPT="process.py"
RUN_ARGS=""
CONDA_ENV="your_env"
WORK_DIR="/work/hdd/your_project/your_username"
# Setup
echo "[$(date)] Job $SLURM_JOB_ID on $(hostname) (CPU-only)"
source /work/hdd/your_project/your_username/miniconda3/etc/profile.d/conda.sh
conda activate "$CONDA_ENV"
cd "$WORK_DIR"
mkdir -p logs
# Run
echo "[$(date)] Running: $RUN_SCRIPT $RUN_ARGS"
python "$RUN_SCRIPT" $RUN_ARGS
echo "[$(date)] Completed."
Cost: 2 SU for two hours at the batch rate, provided the CPU and memory request fits within one GH200 superchip. Charges apply to reserved resources even when the GPU is idle.
Customization Notes
Reservation flag: If the system MOTD mentions --reservation=update, add this line to the #SBATCH directives:
#SBATCH --reservation=update
Email notifications: Add these directives to receive email when jobs start, end, or fail:
#SBATCH --mail-type=BEGIN,END,FAIL
#SBATCH --mail-user=your_email@iit.edu
Job arrays: For parameter sweeps, use Slurm job arrays:
#SBATCH --array=0-9
PARAM_FILE="params.txt"
PARAM=$(sed -n "$((SLURM_ARRAY_TASK_ID + 1))p" "$PARAM_FILE")
python train.py --config "$PARAM"
Time limits: Adjust --time based on measured runtime. A shorter request may fit a backfill window, but does not guarantee an earlier start. Allow time to save results before the limit expires.