Skip to main content
Kueue keeps a Job or JobSet suspended (spec.suspend: true) until it admits the matching Workload. The usual causes are a kueue.x-k8s.io/queue-name label that doesn’t name a LocalQueue, a ClusterQueue without enough quota for the request, or a topology-aware scheduling (TAS) fit failure. Describe the Workload, read its admission events and status conditions, and match the message to a fix in the table below.
Use the fully qualified resource name. The k9s workloads view is a synthetic summary of Pods, Deployments, and similar resources, not the Kueue resource. If the Workload is admitted but its Pods stay Pending, see Why is my Pod stuck in pending on CKS? instead.

Match the message to a fix

Admitted, but a Pod is stuck Pending on a cordoned Node

Understand automatic Node recovery

TAS records a topology assignment at admission and adds Node selectors when it ungates Pods. The selector includes kubernetes.io/hostname only when that label is in the assignment’s topology levels. Automatic Node recovery depends on your installed Kueue version and feature gates. In Kueue 0.19, a NotReady Node is not unconditionally replaced after 30 seconds with the default feature gates. Recovery also considers whether the Workload’s Pods are still running on the Node. An untolerated NoSchedule or NoExecute taint can trigger recovery, but a cordon adds a NoSchedule taint without evicting running Pods. With the default gates, Pods that keep running can delay replacement. See Kueue’s Node failure handling. When kubernetes.io/hostname is the lowest topology level, Kueue records a single failed Node in the Workload’s status.unhealthyNodes and attempts replacement. With the default TASFailedNodeReplacementFailFast gate enabled, it evicts and requeues the Workload if replacement fails.

Force fresh admission for a stuck JobSet

To force fresh admission for a stuck JobSet, deactivate and reactivate its Kueue Workload. This interrupts the workload, including any running Pods. Toggling only the JobSet’s spec.suspend can resume the existing assignment.
  1. Set the Workload’s spec.active to false.
  2. Wait for eviction to finish and the Workload’s status.admission to clear.
  3. Set the Workload’s spec.active to true to allow fresh admission. Kueue reevaluates quota, topology fit, and any admission checks.

Configure recovery for Pods that do not become ready

To recover automatically when Pods do not become ready within a configured timeout, configure waitForPodsReady through the chart’s kueue.managerConfig.controllerManagerConfigYaml value. This value replaces the whole configuration. Set recoveryTimeout for readiness failures after startup and requeuingStrategy.backoffLimitCount to limit retries; reaching the retry limit deactivates the Workload. See Set up all-or-nothing scheduling with ready Pods. Topology.spec.levels is immutable. Before you change topology levels on a production queue, contact CoreWeave Support. For chart installation and topology examples, see Kueue. To monitor admission over time, use the Kueue Scheduling Dashboard. Workload Scheduling
Last modified on October 7, 2026