Skip to content

Troubleshooting

Exercises

  1. Diagnose why a job is pending

Submit a job to the gpu partition requesting 8 GPUs (more than any single node has if nodes have 4 each) and observe the pending reason. Use squeue to identify the reason code, then determine how to fix the request.

Hint / Solution
# Submit a job requesting more GPUs than a single node has
sbatch -p gpu --gres=gpu:8 --nodes=1 --wrap="sleep 120"

# Check the pending reason
squeue -u $USER -o "%.8i %.9P %.20j %.2t %.30R"
# Expected reason: (ReqNodeNotAvail) or (Resources)

# Get more details
scontrol show job <jobid> | grep Reason

# Fix: either request fewer GPUs per node, or span multiple nodes:
#   --gres=gpu:4 --nodes=2 --ntasks-per-node=1
scancel <jobid>
Expected output: Validated 2026-04-15 on slurm-training-val (Slurm 25.11.4, ParallelCluster 3.15.0).
$ squeue -j 42 -o '%.8i %.9P %.20j %.2t %.30R'
   JOBID PARTITION                 NAME ST               NODELIST(REASON)
     42     batch              val_ts1 PD                  (JobHeldUser)

$ scontrol show job 42 | grep -E '(JobState|Reason)'
   JobState=PENDING Reason=JobHeldUser Dependency=(null)
  1. Drain a node and verify it stops accepting jobs

Drain node cpu010 with a reason string. Confirm the node shows as draining/drained in sinfo. Submit a job targeting that specific node and verify it won't run there.

Hint / Solution
# Drain the node
scontrol update NodeName=cpu010 State=DRAIN Reason="exercise: testing drain"

# Verify the state
sinfo -n cpu010
# State should show drain, drng (draining), or drained

# Check the reason
scontrol show node cpu010 | grep -i "state\|reason"

# Try to submit a job to that node -- it should pend
sbatch --nodelist=cpu010 --wrap="hostname"
squeue -u $USER
# Job will pend with reason (ReqNodeNotAvail) until the node is resumed

# Clean up
scancel -u $USER
Expected output: Validated 2026-04-15 on slurm-training-val (Slurm 25.11.4, ParallelCluster 3.15.0).
$ sudo /opt/slurm/bin/scontrol update NodeName=debug-dy-small-2 State=DRAIN Reason='val: ts2 drain test'

$ sinfo -n debug-dy-small-2
PARTITION AVAIL  TIMELIMIT  NODES  STATE NODELIST
batch*       up   infinite      0    n/a 
debug        up   infinite      1 drain~ debug-dy-small-2
gpu          up   infinite      0    n/a 
highmem      up   infinite      0    n/a 

$ scontrol show node debug-dy-small-2 | grep -E '(State|Reason)'
   State=IDLE+CLOUD+DRAIN+POWERED_DOWN+NOT_RESPONDING ThreadsPerCore=1 TmpDisk=0 Weight=1000 Owner=N/A MCS_label=N/A
   Reason=val: ts2 drain test [root@2026-04-15T02:06:38]
  1. Bring a drained node back online

After the previous exercise, resume node cpu010 and verify it returns to an idle state and can accept new jobs.

Hint / Solution
# Resume the node
scontrol update NodeName=cpu010 State=RESUME

# Verify it's back to idle
sinfo -n cpu010
# State should show idle (or idle* if no slurmd running in a lab environment)

scontrol show node cpu010 | grep -i state
# State=IDLE

# Test that it accepts jobs
srun --nodelist=cpu010 hostname
# Should return: cpu010
Expected output: Validated 2026-04-15 on slurm-training-val (Slurm 25.11.4, ParallelCluster 3.15.0).
$ sudo /opt/slurm/bin/scontrol update NodeName=debug-dy-small-2 State=RESUME

$ scontrol show node debug-dy-small-2 | grep State=
   State=IDLE+CLOUD+POWERED_DOWN+NOT_RESPONDING ThreadsPerCore=1 TmpDisk=0 Weight=1000 Owner=N/A MCS_label=N/A
  1. Find a job that was killed by OOM

Search completed jobs from the past 7 days for any that ended in the OUT_OF_MEMORY state. For each, determine how much memory was requested vs. how much was actually used at peak.

Hint / Solution
# Find OOM-killed jobs
sacct --starttime=now-7days --state=OUT_OF_MEMORY \
    --format=JobID%-12,User%-10,JobName%-20,ReqMem,MaxRSS,Elapsed,ExitCode,NodeList

# For a specific job, get the full picture
sacct -j <jobid> --format=JobID,JobName,State,ExitCode,ReqMem,MaxRSS,MaxVMSize,AllocCPUS

# ExitCode 0:9 means killed by SIGKILL (OOM killer)
# MaxRSS close to or at ReqMem confirms memory exhaustion

# Also check for jobs killed by signal 9 that might not be tagged as OOM
sacct --starttime=now-7days --format=JobID,User,State,ExitCode | grep "0:9"
Expected output: Validated 2026-04-15 on slurm-training-val (Slurm 25.11.4, ParallelCluster 3.15.0).
$ sacct --starttime=now-7days --state=OUT_OF_MEMORY --format=JobID%-12,User%-10,JobName%-20,ReqMem,MaxRSS,Elapsed,ExitCode,NodeList
JobID        User       JobName                  ReqMem     MaxRSS    Elapsed ExitCode        NodeList 
------------ ---------- -------------------- ---------- ---------- ---------- -------- --------------- 

$ sacct --starttime=now-7days --format=JobID,User,State,ExitCode | grep '0:9' | head -10

$ sacct --starttime=now-7days --format=JobID,User,State,ExitCode | grep '0:9' | head -10 || echo '(no SIGKILL exits in window)'
  1. Investigate a node in DOWN state

Using scontrol, examine a node that shows as down or down*. Determine the reason it went down, check whether slurmd is running on it, and bring it back if possible.

Hint / Solution
# Find down nodes
sinfo -N -l | grep down

# Get details on a specific down node
scontrol show node cpu005 | grep -i "state\|reason\|lastbusy"

# Check if slurmd is running on the node
ssh cpu005 systemctl status slurmd

# If slurmd is stopped, restart it
ssh cpu005 systemctl restart slurmd

# If slurmd is running but the node is still down*, clear the state
scontrol update NodeName=cpu005 State=RESUME

# Verify recovery
sinfo -n cpu005
Expected output: Validated 2026-04-15 on slurm-training-val (Slurm 25.11.4, ParallelCluster 3.15.0).
$ sinfo -N -l | head -20
Wed Apr 15 02:06:43 2026
NODELIST            NODES PARTITION       STATE CPUS    S:C:T MEMORY TMP_DISK WEIGHT AVAIL_FE REASON              
batch-dy-compute-1      1    batch*        idle 8       8:1:1  31129        0   1000 dynamic, none                
batch-dy-compute-2      1    batch*        idle 8       8:1:1  31129        0   1000 dynamic, none                
batch-dy-compute-3      1    batch*       idle~ 8       8:1:1  31129        0   1000 dynamic, none                
batch-dy-compute-4      1    batch*       idle~ 8       8:1:1  31129        0   1000 dynamic, none                
debug-dy-small-1        1     debug       idle~ 2       2:1:1   3891        0   1000 dynamic, none                
debug-dy-small-2        1     debug       idle~ 2       2:1:1   3891        0   1000 dynamic, none                
gpu-dy-t4-1             1       gpu       idle~ 4       4:1:1  15564        0   1000 dynamic, none                
highmem-dy-large-1      1   highmem       idle~ 8       8:1:1  31129        0   1000 dynamic, none                

$ sinfo -R --noheader || echo '(no nodes with REASON set)'

References