Skip to main content
SUNK Standard SUNK Self-Service Long training jobs on SUNK run across many GPUs for days, so hardware failures during a job are expected. Two layers handle them. CoreWeave automation detects failing hardware and takes the Node out of production. Slurm requeue returns a failed job to the queue, so it restarts on healthy nodes without anyone resubmitting it. This page explains what each layer does, which failures your job has to handle itself, and how to configure requeue for application errors. For the step-by-step setup of a requeueable job that resumes from a checkpoint, see Checkpoint and restart Slurm jobs after a node failure.

How CoreWeave takes failing Nodes out of production

CoreWeave’s Node lifecycle automation monitors the Nodes in your cluster during Day 2+ operations. When it detects a problem, it moves the Node through a lifecycle state transition, such as a reboot or a move to triage. The type of transition depends on how severe the problem is known to be. For the full model, see Node state transitions in CKS. In SUNK, each Slurm compute node runs as a slurmd Pod on one Kubernetes Node, so a transition on the Kubernetes Node reaches Slurm. When CoreWeave cordons the Node, the Syncer drains the matching Slurm node, so Slurm schedules no new jobs on it. The drain reason starts with k8s: cordon: while the slurmd Pod is running, and with k8s: pod terminated: cordon: after the Pod is gone. See Kubernetes drain reasons. CoreWeave uses two types of transition, immediate and pending, and each affects a running job differently. The following sections describe each type.

Immediate transitions

CoreWeave uses an immediate transition when the problem is known to be fatal or to severely degrade performance, such as a confirmed hardware fault or thermal throttling. The Node leaves production immediately, and CoreWeave Kubernetes Service (CKS) evicts its Pods, including the slurmd Pod. The job running on that node fails, and Slurm requeue brings it back.

Pending transitions

CoreWeave uses a pending transition when the severity is uncertain and stopping a running job would cost more than waiting. CKS cordons the Node, the Slurm node drains, and the job already running there keeps running. The transition waits until the Node is idle, however long the job takes, and then CoreWeave reboots the Node or moves it to triage. A Node that goes to triage leaves your cluster, and CKS delivers a replacement that joins Slurm as a new node. This design assumes that a serious fault crashes the job on its own. When it does, the job fails, the Node goes idle, and the transition proceeds.
Don’t undrain a node that carries a k8s: drain reason. Most of these nodes return to production on their own once the repair finishes. For how long a repair takes and which drain reasons you need to act on, see Drain and undrain Slurm nodes.

What happens to the job after a node failure

When the node running a job fails, Slurm marks the node DOWN and ends the job in the NODE_FAIL state. SUNK doesn’t change Slurm’s JobRequeue default. Unless a cluster administrator sets JobRequeue=0, a batch job is eligible for requeue after a node failure without any extra configuration. Slurm returns the job to the queue and restarts the batch script from the beginning under the same job ID. A restart from the beginning loses all progress unless your job saves checkpoints and resumes from the latest one. Checkpoint and restart Slurm jobs after a node failure covers --requeue, detecting a restart with SLURM_RESTART_COUNT, and where to store checkpoints.

Requeue on application exit codes

Node-failure requeue doesn’t cover a job that fails because your application exited with an error. Examples include an out-of-memory error or an NVIDIA Collective Communications Library (NCCL) timeout inside your program. Slurm treats a job that fails this way as a normal job failure and doesn’t requeue it, even with --requeue set. To requeue on application errors, set RequeueExit in the Slurm configuration. It lists the batch job exit codes that should send a job back to the queue. It takes a comma-separated list of single codes and hyphenated ranges, for example 1-9,18. RequeueExitHold takes the same format, but it holds the requeued job until someone releases it with scontrol release. The hold gives you a chance to inspect the failure first. Both are cluster-wide settings, and they apply only to jobs that Slurm can requeue, so a job submitted with --no-requeue isn’t requeued on these codes. See RequeueExit in the Slurm documentation. Replace [EXIT-CODES] with the exit codes your application uses for recoverable failures:
Add the key under slurmConfig in the slurm chart’s values:
Slurm matches the exit code of the batch script, not of each task. If srun is the last command in your script, the script exits with the srun exit code. Otherwise, save the code and end the script with exit and that code. Choose codes that separate a recoverable failure from a bug. A job that exits with a listed code every time requeues every time. To stop this loop, check SLURM_RESTART_COUNT at the start of the script and, once it passes your limit, exit with a code that isn’t in the list.

End a step when one rank fails

In a distributed job, one rank can exit, for example from an out-of-memory error, while the other ranks keep waiting on a collective operation that never completes. The job then holds its GPUs without making progress until a timeout fires. The srun flag --kill-on-bad-exit=1 ends every task in the step as soon as any task exits with a nonzero code. As a result, the job fails quickly and requeue can take over:
The flag has a cost if your job uses Pyxis containers. When it ends the step, Slurm stops the remaining tasks with SIGTERM, waits KillWait seconds, and then sends SIGKILL. That shutdown can interrupt Pyxis while it removes its container directory, and the next attempt can then fail with pyxis: ERROR File already exists. If you use the flag with Pyxis, give each attempt a unique container name or clean up in an Epilog. See Why does Pyxis report File already exists after a job is preempted?

Check the requeue history of a job

To see each run of a requeued job, including the run that failed, query the accounting database with sacct --duplicates. Replace [JOB-ID] with your job ID:
A job requeued after a node failure shows a run in the NODE_FAIL state followed by a later run. To see how many times Slurm has restarted a job, run scontrol show job [JOB-ID] and read the Restarts field. Run these commands from a login node when you need them. Don’t poll them from inside a batch script. Repeated scontrol calls add load to the Slurm controller, and repeated sacct calls add load to the accounting database.
Last modified on September 29, 2026