> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Why does my SUNK-scheduled Kubernetes Pod fail to start?

A SUNK-scheduled Pod fails to start when Slurm and Kubernetes account for a node's resources independently, or when a scheduler annotation holds an invalid value. Run `kubectl describe pod [POD-NAME] -n [NAMESPACE]` to see the admission message, and `squeue -u root` to find the placeholder Slurm job. Common fixes: right-size requests to fit below the node's Kubernetes-allocatable capacity, set `sunk.coreweave.com/exclusive` to a valid value, and don't use `exclusive: "ok"` for GPU Pods.

This article helps you diagnose Kubernetes Pods that fail to start or behave unexpectedly under the [SUNK Pod Scheduler](/products/sunk/run_workloads/schedule-kubernetes-pods). Use it when a Pod with your SUNK `schedulerName` is stuck, gets rejected right after scheduling, or appears briefly and vanishes.

For how the scheduler works and how to configure a Pod, see [Schedule Kubernetes Pods with the SUNK Pod Scheduler](/products/sunk/run_workloads/schedule-kubernetes-pods). This article assumes that setup and focuses on failures.

## What you see

Match your symptom to the section that follows:

* The Pod is rejected with `OutOfcpu` or `OutOfmemory` right after it lands on a node: see [OutOfcpu or OutOfmemory after scheduling](#outofcpu-or-outofmemory-after-scheduling).
* The Pod appears briefly, then disappears, and events mention an invalid exclusive value: see [Invalid exclusive value](#invalid-exclusive-value).
* A GPU Pod collides with a Slurm job over GPUs: see [GPU double allocation](#gpu-double-allocation).
* Pods spread across nodes and leave GPUs stranded during scale-up: see [Fragmentation under scale-up](#fragmentation-under-scale-up).
* A multi-volume Pod (such as a LeaderWorkerSet) stays in `ContainerCreating` and reports no events: see [Duplicate volume declarations](#duplicate-volume-declarations).

## How a SUNK-scheduled Pod works

When a Pod sets your SUNK `schedulerName`, the SUNK Pod Scheduler creates a placeholder Slurm job for it. Slurm determines node placement, GPU allocation, and resource tracking. Kubernetes then runs the Pod on the chosen node. Failures usually come from the two systems accounting for resources independently.

To map a Pod back to its placeholder Slurm job, list the queue with the job name and submitting user shown:

```bash theme={"system"}
# Show queued jobs with their name and submitting user. Read-only.
squeue -o "%.18i %.9P %.30j %.10u %.8T %.10M %R"
```

The scheduler names each placeholder job after the Pod it holds, in the form `[NAMESPACE]/[POD-NAME]`, so the `NAME` column maps a job back to its Pod directly.

The SUNK Pod Scheduler submits placeholder jobs, not you. Unless the Pod sets `sunk.coreweave.com/user-id`, the job is submitted as `root`, so you can also filter with `squeue -u root`. If your Pods do set `user-id`, filter on that user instead.

To view the Pod's admission and scheduling messages, describe the Pod:

```bash theme={"system"}
# Show the Pod's events to see admission and scheduling messages. Safe; read-only.
kubectl describe pod [POD-NAME] -n [NAMESPACE]
```

## OutOfcpu or OutOfmemory after scheduling

Slurm admits the Pod and places it on a node, then the kubelet rejects it with `OutOfcpu`, `OutOfmemory`, or `UnexpectedAdmissionError`.

**What this means.** Slurm and Kubernetes count resources independently, so Slurm can pick a node that looks free to it but whose Kubernetes-allocatable capacity is already partly spent on system Pods and the `slurmd` container. The kubelet then rejects the Pod on its own admission check.

**What to do.** Right-size the Pod's requests so they fit below the node's Kubernetes-allocatable capacity, and make Slurm memory-aware.

Check what `slurmd` actually requests before you try to lower it. On current SUNK versions, enabling the Pod Scheduler applies a low-profile overlay that sets `slurmd` requests to 10 CPU and `10Gi` and leaves its limits untouched, whatever `compute.lowProfileRequests` is set to. If your cluster already shows those values, the headroom is there and the Pod's own requests are what need to change.

For the full procedure and examples, see [Reserve resources for system Pods](/products/sunk/run_workloads/manage-scheduler-resources#reserve-resources-for-system-pods).

## Invalid exclusive value

The Pod appears for a moment, then disappears, and events or scheduler logs report `SlurmJobSubmitFailed` with an invalid exclusive value.

**What this means.** On SUNK v5.7.0 and later, the `sunk.coreweave.com/exclusive` annotation accepts only a fixed set of string values (`"none"`, `"ok"`, `"user"`, `"mcs"`, `"topo"`). Any other non-empty value makes the placeholder Slurm job submission fail, so the Pod never runs. The usual culprits are a typo and the boolean `"true"` that older SUNK versions accepted. An annotation set to an empty string is ignored rather than rejected, and the Pod submits with the partition's own sharing setting.

**What to do.** Set `sunk.coreweave.com/exclusive` to one of the valid values. For the meaning of each value and when to use it, see [Exclusive annotation values](/products/sunk/run_workloads/manage-scheduler-resources#exclusive-annotation-values). Invalid values for other annotations also cause the Pod to fail scheduling.

## GPU double allocation

A GPU Pod and a Slurm job each claim the same GPUs on a node.

**What this means.** Slurm manages GPUs through GRES, which assigns specific GPU indices to each job. Setting `exclusive: "ok"` on a GPU Pod opens the node to any other job. That can collide with Slurm's GRES accounting, leaving two workloads contending for the same physical GPUs.

**What to do.** For GPU Pods, never use `exclusive: "ok"`. Use:

* `exclusive: "none"` when the Pod uses all GPUs on the node (full-node isolation).
* `exclusive: "user"` with a dedicated `user-id` when multiple Pods each use a subset of the node's GPUs.

Never schedule GPU Pods through the standard Kubernetes scheduler on Slurm nodes. That bypasses GRES tracking. See [GPU allocation](/products/sunk/run_workloads/manage-scheduler-resources#gpu-allocation).

## Fragmentation under scale-up

During autoscaling, Pods spread across partially used nodes and leave GPUs stranded.

**What this means.** The SUNK Pod Scheduler doesn't bin-pack. Slurm selects nodes by its internal ordering, which can spread Pods across several partially used nodes instead of filling one before moving to the next. With `exclusive: "user"` or `"none"`, a node holding even one Pod can block other users' jobs, so unused GPUs on it become inaccessible.

**What to do.** This is a known limitation, not a misconfiguration. To reduce its impact:

* Consider a dedicated NodeSet for inference Pods so they don't fragment training capacity.
* Plan exclusive mode with fragmentation in mind. See [Scale-up and scale-down behavior](/products/sunk/run_workloads/manage-scheduler-resources#scale-up-and-scale-down-behavior).

## Duplicate volume declarations

A multi-volume Pod, often part of a LeaderWorkerSet or a templated workload, stays in `ContainerCreating` indefinitely and reports no events.

**What this means.** Declaring the same PersistentVolumeClaim (PVC) as two separate volume entries in one Pod spec can silently break `subPath` mounts. The kubelet fails to create the mount points but emits no Kubernetes event, so `kubectl describe pod` shows nothing and the CSI driver logs a successful attach. Slurm may also show the placeholder job submitting and canceling in a loop, sometimes with `OutOfmemory` errors from the retries. That is a symptom rather than the cause.

**What to do:**

1. Declare each PVC once in `volumes`, then reference it with multiple `volumeMounts` (using `subPath` for distinct locations) rather than listing the PVC twice. Look for two entries under `spec.volumes` that have different `name` values but the same `persistentVolumeClaim.claimName`.
2. Confirm the diagnosis in the kubelet's logs for the node, because the Pod's own events stay empty. Look for a line of this form:

   ```text theme={"system"}
   Error syncing pod, skipping err=unmounted volumes=[tokenizers], unattached volumes=[], failed to process volumes=[data-vast]: context deadline exceeded
   ```

## What this is not

The following points clarify common misconceptions:

* `OutOfcpu` after scheduling isn't a Slurm bug. It's the kubelet enforcing Kubernetes accounting that Slurm doesn't share. The fix is right-sizing requests, not retrying.
* The SUNK Pod Scheduler isn't a gang scheduler. It schedules each Pod as a separate Slurm job and doesn't support multi-node PodGroups. It's best for single-node workloads such as inference.
* Kubernetes scheduling features (`podAffinity`, `podAntiAffinity`, `topologySpreadConstraints`) aren't honored. Slurm makes placement decisions. Node-level affinity such as `gpu.nvidia.com/model` is honored.

## When to file a support ticket

Open a support ticket in any of these situations:

* Pods are rejected with `OutOfcpu` or `OutOfmemory` even after you right-sized requests and made Slurm memory-aware.
* Placeholder Slurm jobs are created but the corresponding Pods never run, and the exclusive value and other annotations are valid.
* GPU contention persists despite using `exclusive: "none"` or `"user"` correctly.

Include the Pod spec (with annotations), the output of `kubectl describe pod`, and `squeue` output showing the placeholder job (`squeue -u root`, unless your Pods set `sunk.coreweave.com/user-id`).

## Related pages

* [Schedule Kubernetes Pods with the SUNK Pod Scheduler](/products/sunk/run_workloads/schedule-kubernetes-pods): scheduler setup, annotations, and known limitations.
* [Manage resources with the SUNK Pod Scheduler](/products/sunk/run_workloads/manage-scheduler-resources): exclusive values, GPU allocation, and system-Pod reservations.
* [Static CPU allocation and the SUNK Pod Scheduler](/products/sunk/run_workloads/static-cpu-allocation): Guaranteed-QoS CPU pinning that drains nodes.
* [SUNK troubleshooting](/support/sunk): all SUNK triage flows.

<Badge stroke shape="pill" color="blue" size="md">[Workload Scheduling](/support/sunk/tags/workload-scheduling)</Badge>


## Related topics

- [Workload Scheduling](/support/sunk/tags/workload-scheduling.md)
