Skip to main content
A SUNK-scheduled Pod fails to start when Slurm and Kubernetes account for a node’s resources independently, or when a scheduler annotation holds an invalid value. Run kubectl describe pod [POD-NAME] -n [NAMESPACE] to see the admission message, and squeue -u root to find the placeholder Slurm job. Common fixes: right-size requests to fit below the node’s Kubernetes-allocatable capacity, set sunk.coreweave.com/exclusive to a valid value, and don’t use exclusive: "ok" for GPU Pods. This article helps you diagnose Kubernetes Pods that fail to start or behave unexpectedly under the SUNK Pod Scheduler. Use it when a Pod with your SUNK schedulerName is stuck, gets rejected right after scheduling, or appears briefly and vanishes. For how the scheduler works and how to configure a Pod, see Schedule Kubernetes Pods with the SUNK Pod Scheduler. This article assumes that setup and focuses on failures.

What you see

Match your symptom to the section that follows:

How a SUNK-scheduled Pod works

When a Pod uses your SUNK schedulerName, the SUNK Pod Scheduler uses a placeholder Slurm job to track its resources. A standalone Pod has its own placeholder. With gang scheduling, the members of a Kubernetes Job or LeaderWorkerSet replica group share a multi-node allocation. Slurm determines placement, GPU allocation, and resource tracking. Kubernetes runs the Pods on the allocated nodes. To map a Pod back to its placeholder Slurm job, list the queue with the job name and submitting user shown:
The NAME column identifies the workload that owns the allocation: Gang members share a Slurm job ID. Use the Pod’s sunk.coreweave.com/slurm-job-id annotation to identify the allocation. See Verify the allocations. The SUNK Pod Scheduler submits placeholder jobs, not you. Unless the Pod sets sunk.coreweave.com/user-id, the job is submitted as root, so you can also filter with squeue -u root. If your Pods do set user-id, filter on that user instead. To view the Pod’s admission and scheduling messages, describe the Pod:

OutOfcpu or OutOfmemory after scheduling

Slurm admits the Pod and places it on a node, then the kubelet rejects it with OutOfcpu, OutOfmemory, or UnexpectedAdmissionError. What this means. Slurm and Kubernetes count resources independently, so Slurm can pick a node that looks free to it but whose Kubernetes-allocatable capacity is already partly spent on system Pods and the slurmd container. The kubelet then rejects the Pod on its own admission check. What to do. Right-size the Pod’s requests so they fit below the node’s Kubernetes-allocatable capacity, and make Slurm memory-aware. Check what slurmd actually requests before you try to lower it. The requests depend on your SUNK version and configuration. Low slurmd requests leave more capacity for other Pods, but you must still account for system Pods and the workload’s own requests. See Lower slurmd resource requests for the version-specific behavior and procedure. For the full procedure and examples, see Reserve resources for system Pods.

Invalid exclusive value

The Pod appears for a moment, then disappears, and events or scheduler logs report SlurmJobSubmitFailed with an invalid exclusive value. What this means. On SUNK v5.7.0 and later, the sunk.coreweave.com/exclusive annotation accepts only a fixed set of string values ("none", "ok", "user", "mcs", "topo"). Any other non-empty value makes the placeholder Slurm job submission fail, so the Pod never runs. The usual culprits are a typo and the boolean "true" that older SUNK versions accepted. An annotation set to an empty string is ignored rather than rejected, and the Pod submits with the partition’s own sharing setting. What to do. Set sunk.coreweave.com/exclusive to one of the valid values. For the meaning of each value and when to use it, see Exclusive annotation values. Invalid values for other annotations also cause the Pod to fail scheduling.

GPU double allocation

A GPU Pod and a Slurm job each claim the same GPUs on a node. What this means. Slurm manages GPUs through GRES, which assigns specific GPU indices to each job. Setting exclusive: "ok" on a GPU Pod opens the node to any other job. That can collide with Slurm’s GRES accounting, leaving two workloads contending for the same physical GPUs. What to do. For GPU Pods, never use exclusive: "ok". Use:
  • exclusive: "none" when the Pod uses all GPUs on the node (full-node isolation).
  • exclusive: "user" with a dedicated user-id when multiple Pods each use a subset of the node’s GPUs.
Never schedule GPU Pods through the standard Kubernetes scheduler on Slurm nodes. That bypasses GRES tracking. See GPU allocation.

Fragmentation under scale-up

During autoscaling, Pods spread across partially used nodes and leave GPUs stranded. What this means. The SUNK Pod Scheduler doesn’t bin-pack. Slurm selects nodes by its internal ordering, which can spread Pods across several partially used nodes instead of filling one before moving to the next. With exclusive: "user" or "none", a node holding even one Pod can block other users’ jobs, so unused GPUs on it become inaccessible. What to do. This is a known limitation, not a misconfiguration. To reduce its impact:
  • Consider a dedicated NodeSet for inference Pods so they don’t fragment training capacity.
  • Plan exclusive mode with fragmentation in mind. See Scale-up and scale-down behavior.

Duplicate volume declarations

A multi-volume Pod, often part of a LeaderWorkerSet or a templated workload, stays in ContainerCreating indefinitely and reports no events. What this means. Declaring the same PersistentVolumeClaim (PVC) as two separate volume entries in one Pod spec can silently break subPath mounts. The kubelet fails to create the mount points but emits no Kubernetes event, so kubectl describe pod shows nothing and the CSI driver logs a successful attach. Slurm may also show the placeholder job submitting and canceling in a loop, sometimes with OutOfmemory errors from the retries. That is a symptom rather than the cause. What to do:
  1. Declare each PVC once in volumes, then reference it with multiple volumeMounts (using subPath for distinct locations) rather than listing the PVC twice. Look for two entries under spec.volumes that have different name values but the same persistentVolumeClaim.claimName.
  2. Confirm the diagnosis in the kubelet’s logs for the node, because the Pod’s own events stay empty. Look for a line of this form:

What this is not

The following points clarify common misconceptions:
  • OutOfcpu after scheduling isn’t a Slurm bug. It’s the kubelet enforcing Kubernetes accounting that Slurm doesn’t share. The fix is right-sizing requests, not retrying.
  • SUNK supports fixed-size gang scheduling for Kubernetes Jobs and LeaderWorkerSet replica groups. Support starts with 7.6.0 in the 7.x series and 8.1.0 in the 8.x series. These workloads use shared Slurm allocations and don’t require PodGroup resources. Standalone Pods and unsupported controllers continue to use one placeholder per Pod. See Gang scheduling limitations and lifecycle.
  • Kubernetes scheduling features (podAffinity, podAntiAffinity, topologySpreadConstraints) aren’t honored. Slurm makes placement decisions. SUNK reads gpu.nvidia.com/class from required node affinity for GPU-type selection; it doesn’t apply other node-affinity rules.

When to file a support ticket

Open a support ticket in any of these situations:
  • Pods are rejected with OutOfcpu or OutOfmemory even after you right-sized requests and made Slurm memory-aware.
  • Placeholder Slurm jobs are created but the corresponding Pods never run, and the exclusive value and other annotations are valid.
  • GPU contention persists despite using exclusive: "none" or "user" correctly.
Include the Pod spec (with annotations), the output of kubectl describe pod, and squeue output showing the placeholder job (squeue -u root, unless your Pods set sunk.coreweave.com/user-id). Workload Scheduling
Last modified on October 7, 2026