DeltaAI FAQ
Access and Authentication
How do I get access to DeltaAI?
You need to be added to the ACCESS allocation by the PI or an allocation manager. Once added, you will receive an NCSA account. You must then enroll in NCSA Duo MFA at https://duo.security.ncsa.illinois.edu (opens in a new tab) before you can log in.
Can I use SSH keys to log in?
General users must use an NCSA password and Duo MFA. NCSA allows SSH keys only for approved Science Gateway accounts; see its login instructions (opens in a new tab).
How do I generate Duo recovery codes?
Visit https://duo.security.ncsa.illinois.edu (opens in a new tab) and check the recovery options available for your account. If you cannot complete MFA, contact NCSA support. Recovery codes do not remove the requirement for interactive authentication.
I keep getting disconnected. How do I maintain my session?
Use tmux on the login node. Start a session with tmux new -s work, and reattach after reconnecting with tmux attach -t work. Always SSH to the same login node (e.g., gh-login04) since tmux sessions do not roam between nodes.
Add ServerAliveInterval 60 to your SSH config to send keepalive packets. This can help detect a broken connection, but cannot prevent every timeout or network interruption.
Which login node should I use?
Any of the four login nodes works (gh-login01 through gh-login04). If you plan to use tmux, pick one and always reconnect to the same one. The dtai-login.delta.ncsa.illinois.edu hostname round-robins between all four.
Jobs and Scheduling
What Slurm account do I use?
Use the account assigned to your ACCESS allocation. Confirm the current value with the PI or allocation manager, or check the accounts available to you with accounts.
Why do interactive sessions cost 2x?
NCSA assigns a 2x charge factor to ghx4-interactive, which is intended for short debugging and prototyping sessions. It assigns a 1x factor to ghx4. Charges apply to reserved resources, including idle time; see NCSA's partition policies (opens in a new tab).
What is the maximum wall time for a job?
Check current limits with scontrol show partition ghx4. Typical limits vary by partition and may change during system updates.
How many GPUs can I request?
Each node has 4 GPUs. You can request 1 to 4 GPUs per node. For multi-node jobs, you can request multiple nodes with up to 4 GPUs each. Partition and allocation limits also apply. NCSA currently limits each user to one queued or running ghx4-interactive job; check current limits (opens in a new tab) before submitting.
What happens if I exceed my allocation?
Jobs may remain pending with QOSGrpBillingMinutes when the allocation cannot cover their requested resources and duration. Ask the PI to request a supplement through the allocation program; see NCSA's allocation instructions (opens in a new tab).
What does --reservation=update mean?
During system updates, NCSA reimages compute nodes and places them in a Slurm reservation. The --reservation=update flag routes your job to nodes running the new image. Check the login MOTD for current reservation requirements. If the MOTD does not mention a reservation, you do not need the flag.
My job is stuck in PENDING. What do I check?
Run squeue -u $USER and look at the REASON column:
| Reason | Meaning | Action |
|---|---|---|
Priority | Other jobs have higher priority | Wait, or reduce resource request |
Resources | Requested resources not yet available | Wait, or request fewer GPUs/nodes |
ReqNodeNotAvail | Requested nodes are down or reserved | Add --reservation=update if needed |
QOSGrpBillingMinutes | Allocation balance cannot cover the request | Check accounts; ask the PI about a supplement |
How do I use srun instead of mpirun?
In an MPI batch job, use srun to launch within the Slurm allocation:
srun --ntasks=4 ./my_program
Inside a batch script, srun inherits the job's resource allocation. For interactive MPI, use salloc first, then srun; do not nest a launch inside an interactive shell started by srun --pty. See NCSA's interactive MPI instructions (opens in a new tab).
Software
Why does my build fail with C++ errors?
Check your project's compiler requirements and the loaded programming environment. GCC 7 has experimental C++17 support, but may lack required features. The IOWarp guide uses GCC 13 explicitly; if those paths are available, configure with:
cmake -DCMAKE_C_COMPILER=/usr/bin/gcc-13 \
-DCMAKE_CXX_COMPILER=/usr/bin/g++-13 \
-B build
Can I install software with apt or yum?
No. You do not have root access. Use conda to install packages into your user environment:
conda install -c conda-forge package_name
Why does my tool crash with "Unsupported system page size"?
Run getconf PAGESIZE to check the node's page size. A tool built with jemalloc may fail if its allocator build does not support that size.
Workaround: Use conda packages (which are compiled for the target platform) or rebuild the tool from source with the system allocator.
How do I install PyTorch with GPU support?
Use a CUDA-enabled aarch64 build compatible with the node's driver. NCSA provides a python/miniforge3_pytorch module; see its software environment guide (opens in a new tab). For a custom environment targeting CUDA 12.8, use the CUDA-specific index:
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
Verify GPU support on a compute node (login nodes have no GPUs):
python -c "import torch; print(torch.cuda.is_available())"
Modules show different packages on login vs compute nodes. Why?
The module system may have different configurations on login and compute nodes. Always verify module availability on the node type where you plan to use the software. Run module avail or module spider package_name from within a Slurm job to see what is available on compute nodes.
Storage
Where should I store my data?
Use /work/hdd/your_project/your_username/ for project files, builds, data, and outputs. NCSA's default WORK-HDD quota is 1 TB shared across the allocation; run quota to check your assigned limit.
Keep small configuration files and scripts in your home directory (/u/your_username/). Its default quota is 100 GB. Put large builds and datasets in WORK storage, and avoid job I/O in HOME.
What is the difference between WORK-HDD and WORK-NVME?
WORK-HDD uses spinning disks; WORK-NVME uses faster NVMe storage. NCSA lists a default quota of 1 TB per allocation for each, shared among project members. Neither provides snapshots. See NCSA's data-management guide (opens in a new tab) and run quota for your limits.
How do I check my quota?
quota
Is /tmp persistent?
No. /tmp on compute nodes is a fast node-local NVMe burst buffer (3.9 TB). It is purged when your job ends. Copy results to /work/hdd/your_project/your_username/ before the job completes.
Troubleshooting
"Exec format error" when running a binary
The binary was compiled for x86_64, not aarch64. Recompile from source on DeltaAI or find an aarch64 build.
Check architecture:
file /path/to/binary
# Expected: ELF 64-bit LSB executable, ARM aarch64
"error while loading shared libraries"
The dynamic loader could not find a required library. If the library exists in your conda environment, add that directory to the runtime search path:
export LD_LIBRARY_PATH="$CONDA_PREFIX/lib${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}"
For a persistent runtime path, configure CMake's build or install RPATH for the library directory. A link-time -L flag alone does not tell the dynamic loader where to find the library when the executable runs.
CMake cannot find a package installed via conda
Conda installs packages under $CONDA_PREFIX. Pass this to CMake:
cmake -DCMAKE_PREFIX_PATH=$CONDA_PREFIX -B build
For header-only or non-standard packages, you may also need:
-DCMAKE_C_FLAGS="-I$CONDA_PREFIX/include"
-DCMAKE_CXX_FLAGS="-I$CONDA_PREFIX/include"
nvidia-smi shows "no devices found" or is not available
Check whether you are on a login node. GPUs are accessible through compute-node allocations. If you are on a login node, start an interactive session:
srun --account=your_account --partition=ghx4-interactive \
--gpus-per-node=1 --time=00:05:00 --pty bash
Then run nvidia-smi from the compute node shell.
My conda environment is too large for HOME
Export the environments you need, install conda at /work/hdd/your_project/your_username/miniconda3, and recreate them there. Follow the user guide and update shell profiles and job scripts to source that installation directly. NCSA advises against symlinks from HOME to WORK or PROJECTS; moving an existing installation can also leave absolute paths pointing to its old location.
Support
- NCSA Help Desk: http://help.ncsa.illinois.edu (opens in a new tab) or help@ncsa.illinois.edu
- DeltaAI Documentation: https://docs.ncsa.illinois.edu/systems/deltaai/en/latest/ (opens in a new tab)
- ACCESS Support: https://support.access-ci.org (opens in a new tab)