Skip to content

Slurm

Documentation updated: 15 September 2026. For the live partition list and time limits, run sinfo.

Access to CPU cores and GPUs is managed by the Slurm batch system.

Submitting

Specify the requested resources in the header of a batch script. For example:

#!/bin/bash
#SBATCH --job-name=myJobName  # the name of the job
#SBATCH -p main               # partition (queue)
#SBATCH --time=2:00:00        # amount of time the job takes
#SBATCH --cpus-per-task=8     # how many threads you wish to use for the given job
#SBATCH --error=%x-%j.err     # STDERR: job name and job ID
#SBATCH --output=%x-%j.out    # STDOUT: job name and job ID
env                           # print the environment variables used
date                          # print the datetime when the script starts

python myProgram.py           # run the program

Test a workload with a small input before submitting many jobs. See the job-submission checklist.

Available queues

Partition Logical nodes Slurm CPU units Time limit1 Intended use
main 147 3,798 2-00:00:00 Default partition for regular CPU jobs
short 71 470 02:00:00 Short CPU jobs
io 147 3,798 2-00:00:00 I/O-heavy jobs; at most 10 CPUs per node
long 5 320 14-00:00:00 Long CPU jobs on a limited set of nodes
gpu 5 122 8-00:00:00 GPU jobs

These are the general user partitions. Additional restricted partitions used for CMS production and infrastructure workloads are intentionally not listed. Node and CPU figures are from the 15 September 2026 configuration. Partitions overlap, so their capacities must not be added together; a Slurm CPU is an allocation unit and does not always correspond to a distinct physical core.

For a GPU job, select the gpu partition and request the required number of accelerators explicitly:

#SBATCH --partition=gpu
#SBATCH --gres=gpu:1

For more up-to-date information on the available queues and corresponding time limits, run sinfo.

Useful commands

For all available options, see the Slurm documentation.

Cancelling your job(s)

In order to cancel your jobs use scancel. For example in order to cancel all your jobs with status PENDING and with a name MyJob:

scancel -u $USER -t PENDING --name MyJob
In order to cancel a single job:
scancel <jobid>

Checking job queues

Check how many jobs are currently in the queue

squeue -h | wc -l

Check how many jobs have you submitted to the queue

squeue -u $USER -h | wc -l

Check how many jobs are currently in the running state

squeue -h -t r | wc -l

Check how many jobs are currently in the pending state

squeue -h -t pd | wc -l

Check how many jobs each user has submitted to the queue

squeue -h -o "%u" | sort | uniq -c | sort -nr -k2

Display the actual command, runtime, node and user who submitted jobs to the queue

squeue -h -o "%o %A %M %u"

in order to not type out/copy the command every time you want to check this, you can add it to your .bashrc:

alias sstatus='squeue -h -o "%u" | sort | uniq -c | sort -nr -k2'
Add that line to .bashrc if you want the alias to persist, then run sstatus.

Job info

Once your job has completed, you can get additional information that was not available during the run. This includes run time, memory used, etc. To get statistics on completed jobs by jobID:

sacct -j <jobid> --format=JobID,JobName,MaxRSS,Elapsed

To view the same information for all jobs of a user:

sacct -u $USER --format=JobID,JobName,MaxRSS,Elapsed

Alternatively, one can similarly use scontrol to gather more information about jobs, but the output is more difficult to parse:

scontrol show -od job | grep $JOB_ID


  1. Timelimit is given in days: hours-minutes-seconds ↩