128 + signal number. The most common are 137 (SIGKILL: out of memory, a liveness probe kill, or a preemption), 139 (SIGSEGV: an application crash, a driver problem, or, less often, hardware), and 143 (SIGTERM: a graceful shutdown request). The same code can mean several different things, so translate the code to a signal and then read the matching section to find which cause applies.
What you see
Your job ended with a numeric exit code. In Kubernetes, you find it in the Pod status. In Slurm, you find it in the job’s exit code or in the launcher logs. Common values are 137, 139, 143, 132, and 135. Exit codes above 128 mean a signal killed the process, and the code is128 + signal number. For example, 137 is 128 + 9 (SIGKILL). Use this to translate a code to a signal, then read the matching section.
Slurm and Python report the same signals as negative numbers instead, so a job that Kubernetes records as 137 can appear in sacct output or a Python traceback as -9. The signal is the same. Subtract 128 from the positive form, or drop the minus sign from the negative form, to get the signal number in the following table.
Exit-code table
The following table maps each exit code to its signal, common meaning, and first thing to check.Exit 137 (SIGKILL) out of memory, a probe kill, or preemption
Exit 137 alone doesn’t prove an out-of-memory kill. Anything that sends SIGKILL lands here, including routine platform actions. Start with the Pod’s last terminated reason. A training Pod usually runs more than one container, so print the reason for each one rather than assuming the first is the training container:OOMKilledmeans a memory kill, but not necessarily one you caused. Read the kernel out-of-memory constraint to classify it:CONSTRAINT_MEMCGis a container limit you set too low, andCONSTRAINT_NONEis Node-level pressure that may belong to CoreWeave. For the full procedure, see How do I diagnose out-of-memory errors?. Once you’ve confirmed it’s your limit, reduce memory use (smaller batch, gradient checkpointing) or raise the Pod memory limit.Error(notOOMKilled) with liveness probe failure events means a liveness probe killed the Pod, not memory. See the liveness probe section.Errorwith no probe failure and no memory pressure often means something outside the workload ended it. A Slurm preemption is the common one on SUNK: preemption uses the partition’s or QOS’sGraceTime, which defaults to0, so SIGKILL can follow SIGTERM almost immediately and the job records 137 rather than the 143 you’d expect from a graceful stop. See Handle Slurm signals. A rolling update, a Pod deletion, or a Node drain also ends in SIGKILL.
Exit 139 (SIGSEGV) application or hardware
A segmentation fault is most often an application crash, often inside a native extension such as numpy, a custom CUDA op, or a framework C++ layer. Use the failure pattern to classify it:- Same GPU or same node every run points at hardware. Check
dmesgon that node for a GPU XID at the time of the crash. - A random rank each run, different node each time points at the application. The bug travels with the code, not the hardware. Investigate the workload.
- During checkpoint saving, followed by a collective timeout points at the GPU driver. On GB200 nodes running driver 580 this has a specific signature. See Checkpoint segfault at scale.
Exit 143 (SIGTERM) graceful termination
SIGTERM is a request to shut down cleanly. Something external asked the job to stop, and the process had time to exit before SIGKILL arrived. Check for a preemption by a higher-priority job, a Node Pool scale-down, a Slurm time limit reached, or a manual cancellation. This is usually expected behavior, not a fault. Look at the job’s events and the scheduler log to find what sent the signal. The grace window matters here. On cancellation or a time limit, SUNK gives your handler roughly theKillWait interval (30 seconds by default) to save a checkpoint, and the job records 143. Preemption uses GraceTime instead, which defaults to 0, so a preempted job often records 137 rather than 143. See Handle Slurm signals.
SIGILL and misleading errors on aarch64 (Grace CPU)
GB200, GB300, and GH200 instances use NVIDIA Grace CPUs, which are aarch64 (Arm), not x86_64. Two exit-code symptoms point at this architecture:- SIGILL (exit 132, or
-4) from x86_64 native code reaching the CPU. An image whose top-level executable is x86_64 fails earlier, at exec, with anexec format error, so SIGILL means the process started and then ran an instruction the Arm CPU doesn’t implement. The documented triggers are all x86 code carried into an otherwise-aarch64 environment: DeepSpeed or Triton JIT kernel caches shared across architectures on network storage, checkpoints saved on x86_64 nodes that embed compiled ops, and numpy or BLAS builds carrying x86 SIMD instructions. The numpy and BLAS case lands in dataloader workers, so the crash appears deep inside data loading rather than at startup, and it’s scale-dependent: it tracksnum_workersandprefetch_factor, so more concurrent workers make it more likely. Look forRuntimeError: DataLoader worker pid [N] is killed by signal: Illegal instruction. - A confusing “out of memory” report instead of a clean SIGILL. An aarch64 binary compiled with a hardcoded 4 KiB page-size assumption can segfault or report a misleading memory error on Grace, which uses 64 KiB pages. The crash looks like a memory bug but is a page-size bug. If a job that runs on x86_64 nodes fails only on Grace nodes, check the architecture and page size before you request more memory.
Liveness probe pitfall
A liveness probe that fires during slow startup can kill a healthy Pod. The signature has two parts:- Exit 137 with
lastState.terminated.reasonequal toError(notOOMKilled). - Liveness probe failure events in
kubectl describe pod.
initialDelaySeconds or timeoutSeconds to cover the real startup time. Tune only the probe configuration in your own workload.
Works at 50 nodes but fails at 100 or more with a different rank each run
A job that succeeds at small scale but fails at larger scale, with a different failing rank each run, is an application issue at scale, not hardware. Hardware faults stay tied to specific nodes. An application bug that depends on rank count or data sharding moves around. Investigate the workload’s scaling behavior (file-handle limits, per-rank memory growth, rendezvous timeout) rather than suspecting the nodes. See torch.distributed rendezvous and TCPStore troubleshooting.Bisect by scale to isolate the cause
To separate a workload problem from an infrastructure problem, bisect by GPU count and run a hardware test at the scale where your job fails. Run the same workload at increasing scale (for example, 8 GPUs, then 64, then 256) and record the smallest scale at which the failure appears. At that scale, run an NCCL test across the same set of nodes (how to run one):- If the NCCL test also fails at that scale, the problem is in the hardware, fabric, or placement. See NCCL error reference.
- If the NCCL test passes at that scale, the problem is in your workload. Investigate file-handle limits, per-rank memory growth, distributed checkpoint contention, and rendezvous timeout.
Read incomplete utilization
If the Kubernetes Training Workloads dashboard shows a Pods count lower than the number you requested, for example, 62 of 64, work through the causes in this order:- Pods that never scheduled. Check the Pod Readiness Timeline panel for rows sitting in Pending. A Pending Pod is a scheduling problem, which can be capacity or quota rather than anything in your workload.
- Pods that keep restarting. A repeated sawtooth on the Topology Churn panel means Pods are crash-looping or rescheduling across nodes. Diagnose the crash with the exit code, not the Pod count.
- A workload that doesn’t address every rank. Only once all replicas are Running and stable does a shortfall point at your job. Confirm that your job requested the expected number of replicas and that your launcher addresses every rank.
Confirm the cause with kubectl
Two commands show most of what you need:kubectl describe pod, the Last State block shows the exit code and reason. The events list shows whether a probe, the out-of-memory killer, or the scheduler acted on the Pod.
What this is not
The following points correct common misreadings of these exit codes:- Exit 137 isn’t always out of memory. A liveness probe, a preemption at the default
GraceTimeof0, a rolling update, or a Node drain all send SIGKILL too. OOMKilledisn’t always your memory limit. Read the kernel out-of-memory constraint:CONSTRAINT_NONEmeans Node-level pressure that may be CoreWeave’s to fix, not a limit for you to raise.- Exit 139 on a random rank each run isn’t a hardware fault. It’s an application bug that travels with the code.
- A SIGILL on Grace nodes isn’t a node defect. It’s x86_64 native code reaching an Arm CPU, usually through a shared JIT cache, an x86-saved checkpoint, or an x86 numpy or BLAS build.
- A Grace out-of-memory error isn’t always a memory shortage. An aarch64 binary that assumes 4 KiB pages misreports as one.
Related pages
- NCCL error reference and diagnostic decision tree
- torch.distributed rendezvous and TCPStore troubleshooting
- Run jobs on Grace CPU instances
- How do I diagnose out-of-memory errors?: classifying an
OOMKilledcontainer by kernel constraint. - Handle Slurm signals: the
KillWaitandGraceTimewindows behind 137 and 143.