Skip to content

GPU Jobs

This page explains how to run GPU-accelerated computations on the XLence cluster.


Requesting GPU Resources

To request GPU resources, use the --gres=gpu:N option where N is the number of GPUs.

#SBATCH --gres=gpu:1  # Request 1 GPU
#SBATCH --gres=gpu:2  # Request 2 GPUs

Important: CPU Requirements for GPU Jobs

Many modern GPU-accelerated software packages (molecular dynamics, deep learning, etc.) require only 1 CPU core when running on GPU, as most computation happens on the GPU itself.

However, not all software follows this pattern—some may benefit from multiple CPU cores even when using GPUs.

Always check your software documentation to determine the optimal CPU/GPU configuration.


Common Configurations

Single-Core GPU Job (Most Common)

This is the typical configuration for many MD simulations and ML applications:

#!/bin/bash
#SBATCH --job-name=gpu_job
#SBATCH --partition=normal
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=1      # Many GPU programs need only 1 CPU
#SBATCH --gres=gpu:1            # Request 1 GPU
#SBATCH --mem=16G
#SBATCH --time=24:00:00
#SBATCH --output=gpu_job_%j.out

module purge
module load software/version

your_gpu_program input.txt

Multi-Core GPU Job

For software that benefits from CPU parallelism alongside GPU acceleration:

#!/bin/bash
#SBATCH --job-name=gpu_job_multicore
#SBATCH --partition=normal
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4       # Request multiple CPUs if needed
#SBATCH --gres=gpu:1            # Request 1 GPU
#SBATCH --mem=32G
#SBATCH --time=24:00:00
#SBATCH --output=gpu_job_%j.out

module purge
module load software/version

your_gpu_program --threads 4 input.txt

Multi-GPU Job

Some applications can use multiple GPUs:

#!/bin/bash
#SBATCH --job-name=multi_gpu
#SBATCH --partition=normal
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=2       # Often 1 CPU per GPU
#SBATCH --gres=gpu:2            # Request 2 GPUs
#SBATCH --mem=32G
#SBATCH --time=24:00:00
#SBATCH --output=multi_gpu_%j.out

module purge
module load software/version

your_gpu_program --gpus 2 input.txt

GPU-Specific Examples

GROMACS with GPU

#!/bin/bash
#SBATCH --job-name=gromacs_gpu
#SBATCH --partition=normal
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8       # GROMACS benefits from multiple CPU cores
#SBATCH --gres=gpu:1
#SBATCH --mem=16G
#SBATCH --time=48:00:00
#SBATCH --output=gromacs_gpu_%j.out

module purge
module load gromacs/2025.0

# GROMACS automatically detects and uses GPUs
gmx_mpi mdrun -s md.tpr -deffnm md -ntomp 8 -gpu_id 0

AMBER with GPU

#!/bin/bash
#SBATCH --job-name=amber_gpu
#SBATCH --partition=normal
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=1       # AMBER GPU code uses 1 CPU
#SBATCH --gres=gpu:1
#SBATCH --mem=8G
#SBATCH --time=24:00:00
#SBATCH --output=amber_gpu_%j.out

module purge
module load amber

# Run AMBER GPU MD
pmemd.cuda -O -i md.in -o md.out -p system.prmtop -c system.inpcrd -r md.rst -x md.nc

NAMD with GPU

#!/bin/bash
#SBATCH --job-name=namd_gpu
#SBATCH --partition=normal
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=1       # Single-node NAMD uses 1 CPU with GPU
#SBATCH --gres=gpu:1
#SBATCH --mem=8G
#SBATCH --time=24:00:00
#SBATCH --output=namd_gpu_%j.out

module purge
module load namd/3.0

# NAMD automatically detects GPUs
namd3 +p1 +devices 0 simulation.conf > output.log

PyTorch / TensorFlow (Deep Learning)

#!/bin/bash
#SBATCH --job-name=pytorch_training
#SBATCH --partition=normal
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4       # Data loading often benefits from multiple CPUs
#SBATCH --gres=gpu:1
#SBATCH --mem=32G
#SBATCH --time=48:00:00
#SBATCH --output=pytorch_%j.out

module purge
module load miniforge3/20250911

# Activate conda environment with PyTorch
source activate pytorch_env

# Train model
python train.py --gpu 0 --epochs 100

Checking GPU Availability

From Login Node

# Check which GPUs are available on nodes
sinfo -o "%N %G"

From Compute Node (Interactive Session)

# Request interactive session with GPU
srun --partition=normal --gres=gpu:1 --pty bash

# Check GPU with nvidia-smi
nvidia-smi

Within a Job Script

#!/bin/bash
#SBATCH --gres=gpu:1

# Check which GPU is assigned
nvidia-smi

# Show GPU usage during computation
your_gpu_program input.txt

GPU Best Practices

1. Check Software GPU Requirements

Before running GPU jobs, verify: - Does your software support GPU acceleration? - Which GPU libraries are required (CUDA, cuDNN, etc.)? - How many CPUs should be used with GPUs?

Consult the software's module documentation: module help software/version

2. Monitor GPU Utilization

Use nvidia-smi to check if your GPU is being utilized efficiently:

# In your job script or interactive session
watch -n 1 nvidia-smi

If GPU utilization is low, your software may not be properly configured for GPU use.

3. Respect Local Conventions

Remember to follow the CPU/GPU conventions: - Most GPU jobs need only 1 CPU per GPU - This ensures fair resource sharing

4. Test GPU vs CPU Performance

Not all problems benefit from GPU acceleration. For small systems or short calculations, CPU-only may be faster due to GPU setup overhead.

Always test to confirm GPU provides speedup for your specific use case.

5. Clean Up After GPU Jobs

Ensure your job properly releases GPU resources:

# At the end of your script
nvidia-smi  # Verify GPU is released

Troubleshooting GPU Jobs

GPU Not Detected

Problem: Job runs but doesn't use GPU

Solutions: - Verify #SBATCH --gres=gpu:1 is in your script - Check software is GPU-enabled version: module show software/version - Ensure CUDA libraries are loaded (usually automatic with modules) - Check software command-line flags for GPU activation

Out of GPU Memory

Problem: Job fails with "out of memory" error

Solutions: - Reduce batch size (for ML applications) - Use smaller system (for MD simulations) - Request different GPU with more memory - Use CPU-only for this calculation

GPU Conflict

Problem: Job fails with "GPU already in use"

Solutions: - Check another job isn't using the GPU: nvidia-smi - Contact administrators if issue persists - GPU should be exclusively allocated by Slurm

Poor GPU Performance

Problem: GPU job runs but is slow

Check:

# Monitor GPU utilization
nvidia-smi

# If utilization is low (<50%), possible causes:
# - Software not properly GPU-accelerated
# - System too small to benefit from GPU
# - CPU bottleneck (may need more CPUs)
# - I/O bottleneck (data transfer to GPU)

Support

For help with GPU jobs:


See Also: - Submitting Jobs - General job submission guide - Monitoring & Troubleshooting - Track and debug jobs - Job Scheduling Overview - Cluster policies and conventions


Last Updated: October 2025