GPU Jobs
This page explains how to run GPU-accelerated computations on the XLence cluster.
Requesting GPU Resources
To request GPU resources, use the --gres=gpu:N option where N is the number of GPUs.
#SBATCH --gres=gpu:1 # Request 1 GPU
#SBATCH --gres=gpu:2 # Request 2 GPUs
Important: CPU Requirements for GPU Jobs
Many modern GPU-accelerated software packages (molecular dynamics, deep learning, etc.) require only 1 CPU core when running on GPU, as most computation happens on the GPU itself.
However, not all software follows this pattern—some may benefit from multiple CPU cores even when using GPUs.
Always check your software documentation to determine the optimal CPU/GPU configuration.
Common Configurations
Single-Core GPU Job (Most Common)
This is the typical configuration for many MD simulations and ML applications:
#!/bin/bash
#SBATCH --job-name=gpu_job
#SBATCH --partition=normal
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=1 # Many GPU programs need only 1 CPU
#SBATCH --gres=gpu:1 # Request 1 GPU
#SBATCH --mem=16G
#SBATCH --time=24:00:00
#SBATCH --output=gpu_job_%j.out
module purge
module load software/version
your_gpu_program input.txt
Multi-Core GPU Job
For software that benefits from CPU parallelism alongside GPU acceleration:
#!/bin/bash
#SBATCH --job-name=gpu_job_multicore
#SBATCH --partition=normal
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4 # Request multiple CPUs if needed
#SBATCH --gres=gpu:1 # Request 1 GPU
#SBATCH --mem=32G
#SBATCH --time=24:00:00
#SBATCH --output=gpu_job_%j.out
module purge
module load software/version
your_gpu_program --threads 4 input.txt
Multi-GPU Job
Some applications can use multiple GPUs:
#!/bin/bash
#SBATCH --job-name=multi_gpu
#SBATCH --partition=normal
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=2 # Often 1 CPU per GPU
#SBATCH --gres=gpu:2 # Request 2 GPUs
#SBATCH --mem=32G
#SBATCH --time=24:00:00
#SBATCH --output=multi_gpu_%j.out
module purge
module load software/version
your_gpu_program --gpus 2 input.txt
GPU-Specific Examples
GROMACS with GPU
#!/bin/bash
#SBATCH --job-name=gromacs_gpu
#SBATCH --partition=normal
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8 # GROMACS benefits from multiple CPU cores
#SBATCH --gres=gpu:1
#SBATCH --mem=16G
#SBATCH --time=48:00:00
#SBATCH --output=gromacs_gpu_%j.out
module purge
module load gromacs/2025.0
# GROMACS automatically detects and uses GPUs
gmx_mpi mdrun -s md.tpr -deffnm md -ntomp 8 -gpu_id 0
AMBER with GPU
#!/bin/bash
#SBATCH --job-name=amber_gpu
#SBATCH --partition=normal
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=1 # AMBER GPU code uses 1 CPU
#SBATCH --gres=gpu:1
#SBATCH --mem=8G
#SBATCH --time=24:00:00
#SBATCH --output=amber_gpu_%j.out
module purge
module load amber
# Run AMBER GPU MD
pmemd.cuda -O -i md.in -o md.out -p system.prmtop -c system.inpcrd -r md.rst -x md.nc
NAMD with GPU
#!/bin/bash
#SBATCH --job-name=namd_gpu
#SBATCH --partition=normal
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=1 # Single-node NAMD uses 1 CPU with GPU
#SBATCH --gres=gpu:1
#SBATCH --mem=8G
#SBATCH --time=24:00:00
#SBATCH --output=namd_gpu_%j.out
module purge
module load namd/3.0
# NAMD automatically detects GPUs
namd3 +p1 +devices 0 simulation.conf > output.log
PyTorch / TensorFlow (Deep Learning)
#!/bin/bash
#SBATCH --job-name=pytorch_training
#SBATCH --partition=normal
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4 # Data loading often benefits from multiple CPUs
#SBATCH --gres=gpu:1
#SBATCH --mem=32G
#SBATCH --time=48:00:00
#SBATCH --output=pytorch_%j.out
module purge
module load miniforge3/20250911
# Activate conda environment with PyTorch
source activate pytorch_env
# Train model
python train.py --gpu 0 --epochs 100
Checking GPU Availability
From Login Node
# Check which GPUs are available on nodes
sinfo -o "%N %G"
From Compute Node (Interactive Session)
# Request interactive session with GPU
srun --partition=normal --gres=gpu:1 --pty bash
# Check GPU with nvidia-smi
nvidia-smi
Within a Job Script
#!/bin/bash
#SBATCH --gres=gpu:1
# Check which GPU is assigned
nvidia-smi
# Show GPU usage during computation
your_gpu_program input.txt
GPU Best Practices
1. Check Software GPU Requirements
Before running GPU jobs, verify: - Does your software support GPU acceleration? - Which GPU libraries are required (CUDA, cuDNN, etc.)? - How many CPUs should be used with GPUs?
Consult the software's module documentation: module help software/version
2. Monitor GPU Utilization
Use nvidia-smi to check if your GPU is being utilized efficiently:
# In your job script or interactive session
watch -n 1 nvidia-smi
If GPU utilization is low, your software may not be properly configured for GPU use.
3. Respect Local Conventions
Remember to follow the CPU/GPU conventions: - Most GPU jobs need only 1 CPU per GPU - This ensures fair resource sharing
4. Test GPU vs CPU Performance
Not all problems benefit from GPU acceleration. For small systems or short calculations, CPU-only may be faster due to GPU setup overhead.
Always test to confirm GPU provides speedup for your specific use case.
5. Clean Up After GPU Jobs
Ensure your job properly releases GPU resources:
# At the end of your script
nvidia-smi # Verify GPU is released
Troubleshooting GPU Jobs
GPU Not Detected
Problem: Job runs but doesn't use GPU
Solutions:
- Verify #SBATCH --gres=gpu:1 is in your script
- Check software is GPU-enabled version: module show software/version
- Ensure CUDA libraries are loaded (usually automatic with modules)
- Check software command-line flags for GPU activation
Out of GPU Memory
Problem: Job fails with "out of memory" error
Solutions: - Reduce batch size (for ML applications) - Use smaller system (for MD simulations) - Request different GPU with more memory - Use CPU-only for this calculation
GPU Conflict
Problem: Job fails with "GPU already in use"
Solutions:
- Check another job isn't using the GPU: nvidia-smi
- Contact administrators if issue persists
- GPU should be exclusively allocated by Slurm
Poor GPU Performance
Problem: GPU job runs but is slow
Check:
# Monitor GPU utilization
nvidia-smi
# If utilization is low (<50%), possible causes:
# - Software not properly GPU-accelerated
# - System too small to benefit from GPU
# - CPU bottleneck (may need more CPUs)
# - I/O bottleneck (data transfer to GPU)
Support
For help with GPU jobs:
- Uliano Guerrini: uliano.guerrini@unimi.it
- Omar Ben Mariem: omar.benmariem@unimi.it
See Also: - Submitting Jobs - General job submission guide - Monitoring & Troubleshooting - Track and debug jobs - Job Scheduling Overview - Cluster policies and conventions
Last Updated: October 2025