Job Scheduling with Slurm
The XLence cluster uses Slurm (Simple Linux Utility for Resource Management) as the workload manager for submitting and managing computational jobs.
Overview
Slurm is a job scheduler that: - Allocates resources (CPUs, GPUs, memory) on compute nodes - Queues jobs when resources are busy - Manages job priorities and fair resource sharing - Monitors job execution and resource usage
Basic Workflow
- Prepare your job: Create a job script or command
- Submit to Slurm: Use
sbatch(batch jobs) orsrun(interactive) - Monitor job: Check status with
squeue - Retrieve results: Access output files when job completes
See the detailed guides: - Submitting Jobs - How to submit batch and interactive jobs - GPU Jobs - Running GPU-accelerated computations - Monitoring & Troubleshooting - Track and debug your jobs
Partitions
The cluster has two partitions:
normal (default)
- 5 compute nodes
- 10 CPUs per node
- 128 GB RAM per node
- 2 GPUs per node
ngs
- 1 compute node
- 18 CPUs
- 128 GB RAM
- 1 GPU
Specify partition in your job script:
#SBATCH --partition=normal # or --partition=ngs
Local Conventions (Important!)
To maximize cluster utilization and ensure fair access to both CPU and GPU resources, we ask all users to follow these guidelines:
CPU Job Limits by Partition
Partition normal (10 CPUs, 2 GPUs per node):
- When submitting CPU-only jobs, request maximum 8 CPUs per job
- This leaves 2 CPUs available (1 per GPU) for GPU jobs
Partition ngs (18 CPUs, 1 GPU):
- When submitting CPU-only jobs, request maximum 17 CPUs per job
- This leaves 1 CPU available for GPU jobs
Why This Matters
Many modern GPU-accelerated programs (molecular dynamics, deep learning, etc.) require only 1 CPU core per GPU, as most computation happens on the GPU itself. By reserving 1 CPU per GPU on each node, we ensure that GPU jobs can run even when CPU resources are heavily utilized, maximizing overall cluster throughput.
Examples Following Conventions
# CPU job on normal partition (leaves room for 2 GPU jobs)
#SBATCH --partition=normal
#SBATCH --cpus-per-task=8 # Maximum recommended for CPU jobs
# CPU job on ngs partition (leaves room for 1 GPU job)
#SBATCH --partition=ngs
#SBATCH --cpus-per-task=17 # Maximum recommended for CPU jobs
# GPU jobs can use 1 CPU (most common case)
#SBATCH --partition=normal
#SBATCH --cpus-per-task=1
#SBATCH --gres=gpu:1
Resource Limits
The XLence cluster operates on a trust-based model without strict enforcement of time limits or resource quotas.
Our Philosophy
For over five years, our user community has successfully self-regulated resource usage through mutual respect and consideration. As long as users continue to:
- Be mindful of resource requests
- Consider other users' needs
- Follow the local conventions (CPU limits)
- Communicate about large or long-running jobs
No hard limits will be imposed.
What This Means
- No time walls: Jobs can run as long as needed
- No strict quotas: Request resources based on actual needs
- No rigid limits: Flexibility for legitimate research requirements
- Community-based: We trust users to be reasonable
If Problems Arise
If the community experiences resource conflicts or abuse, administrators may: 1. Contact users to discuss resource usage 2. Work with research groups to coordinate large jobs 3. Only as a last resort: implement technical limits
We prefer to maintain the current collaborative environment. The success of the past five years demonstrates that our community can manage shared resources responsibly without bureaucratic overhead.
Your Responsibility
- Estimate resource needs realistically: Don't over-request "just in case"
- Monitor your jobs: Cancel jobs that fail early to free resources
- Communicate: If planning large resource usage, coordinate with administrators
- Be considerate: Remember others are waiting for resources too
Additional Resources
- Official Slurm documentation: https://slurm.schedmd.com/
- Submitting Jobs - Detailed job submission guide
- GPU Jobs - GPU-specific information
- Monitoring & Troubleshooting - Job management and debugging
Support
For help with job scheduling or Slurm issues:
- Uliano Guerrini: uliano.guerrini@unimi.it
- Omar Ben Mariem: omar.benmariem@unimi.it
Last Updated: October 2025