About Kueue
Kueue is a Kubernetes-native system that manages jobs using quotas. Kueue makes job decisions based on resource availability, job priorities, and the quota policies defined in your cluster queues. Kueue determines when a job waits for available resources, when a job starts (Pods created), and when a job is preempted (active Pods deleted). Use Kueue on CKS to prioritize batch, AI, and ML workloads, share GPU capacity fairly across teams, and reduce idle time on expensive accelerators. CKS supports Kueue by default. To simplify getting started, CoreWeave provides a Helm chart for installing Kueue. Thecks-kueue chart also includes a kueue subchart, used to configure Kueue for deployment into your CKS cluster.
When you install Kueue through the CoreWeave Helm chart, Kueue metrics are automatically scraped and ingested into the Kueue Scheduling Dashboard in CoreWeave Grafana.
Usage
Install thecks-kueue Helm chart to deploy Kueue into your CKS cluster. The chart installs the Kueue controller and CRDs so you can begin defining queues and submitting workloads.
Add the CoreWeave Helm repo so Helm can locate the cks-kueue chart.
kueue-system namespace and the Kueue CRDs are available for use in your cluster.
Sample Kueue configuration
After you install thecks-kueue chart, use the following sample configuration to set up a basic Kueue environment for CKS. This configuration creates a single shared queue that gets you started with a proof of concept. Production deployments usually divide capacity among teams instead. For more information, see Share capacity across teams with Cohorts.
The samples use kueue.x-k8s.io/v1beta2. Before applying them, confirm that your installed Kueue CRDs serve this API version.
The configuration includes the following Kueue components:
ResourceFlavor: Defines the characteristics of compute resources (CPU, memory, GPUs) available in your cluster.ClusterQueue: Establishes resource quotas and admission policies across your entire cluster.LocalQueue: Creates namespaced queues that reference aClusterQueuefor job submission.WorkloadPriorityClass: Defines priority levels for jobs to determine scheduling order and preemption behavior.
For more information, see Using Kueue with Ray on CKS.
Share capacity across teams with Cohorts
ACohort groups ClusterQueues so that a team can borrow quota that other teams aren’t using. With the reclaimWithinCohort policy shown in the following sample, a team can reclaim its nominal quota by preempting eligible workloads in borrowing queues. Admission still depends on the workload’s resource requirements and scheduling constraints.
A single ClusterQueue is enough for a proof of concept, but most organizations divide cluster capacity among several teams. You can give each team its own ClusterQueue with a nominal quota and a LocalQueue in the team’s namespace. Nominal quota defines an admission entitlement; it does not reserve physical GPUs.
The following configuration reuses the previous sample’s ResourceFlavor and WorkloadPriorityClass resources. It splits the same resource quotas between two teams, team-a and team-b, and joins their ClusterQueues in a Cohort named org-cohort. It also creates the team namespaces and a default LocalQueue in each namespace.
Use these team queues instead of the previous sample’s shared queue. If you applied that sample, stop submitting jobs to its
default LocalQueue in the default namespace. Wait for its workloads to finish, then remove that LocalQueue and cluster-queue before applying the team configuration. Otherwise, both configurations can admit work against the same physical capacity.kueue.x-k8s.io/queue-name: default. To select a priority, also set kueue.x-k8s.io/priority-class to prod-priority or dev-priority on the job. Creating the priority classes alone does not assign them to jobs.
- When
team-bis idle,team-ahas a quota ceiling of 12 GPUs: its 8 nominal GPUs plus 4 borrowed from the Cohort. Running those jobs also requires sufficient CPU, memory, RDMA resources, and physical capacity. - When
team-bneeds its nominal quota, Kueue can preempt eligible workloads inteam-a’s borrowing queue to make room (reclaimWithinCohort: Any). Reclaim can select workloads regardless of their priority, provided the incoming workload can fit after preemption. - Within each queue, jobs with the
prod-priorityclass can preempt jobs with thedev-priorityclass when needed for admission (withinClusterQueue: LowerPriority).
Kueue’s fair sharing feature supports
withinClusterQueue and reclaimWithinCohort, but not borrowWithinCohort. If you enable fair sharing, remove borrowWithinCohort from your ClusterQueues.Plan your ClusterQueue split
The right way to divide quota among ClusterQueues depends on your organization. When you plan the split, consider the following:- Create one
ClusterQueuefor each team that needs a nominal quota, and aLocalQueuein each namespace where that team submits jobs. - Set each
nominalQuotato the team’s share. For a fixed-size cluster, keep the sum of quotas within the allocatable capacity of the Nodes that theResourceFlavormatches, after accounting for system components and workloads outside Kueue. Kueue doesn’t validate configured quotas against actual Node capacity, although topology-aware scheduling checks physical fit during workload admission. - Use
namespaceSelectorto control which namespaces can submit jobs to each queue. - Use
borrowingLimitandlendingLimitto bound how much capacity moves between teams. For more information, see ClusterQueue concepts in the Kueue documentation. - Define
WorkloadPriorityClassresources and assign them to jobs for consistent priority ordering. WithreclaimWithinCohort: Any, reclaim can preempt even higher-priority workloads in borrowing queues. - Revisit the split when you add Node capacity or when teams change. Adding Nodes does not automatically increase the nominal quotas configured in these samples. Update those quotas when you want the queues to admit more work.
Observability
This section describes how to monitor Kueue activity after the chart is installed. CoreWeave Grafana provides a Kueue Scheduling Dashboard that you can use to monitor your Kueue cluster.Topology-aware scheduling
Topology-aware scheduling (TAS) lets Kueue improve scheduling decisions by considering the physical topology of your cluster’s Nodes. This is important for HPC, AI, and ML workloads, where network latency between Nodes can be a performance bottleneck. TAS can co-locate a job’s Pods to minimize communication overhead and maximize performance. TheTopologyAwareScheduling feature in the Kueue controller is enabled by default. However, to use it, you must adjust some of the Kueue resources so that Kueue references the Node labels that describe your cluster’s topology.
After the Helm chart is installed and the Kueue CRDs exist, choose one of the following topologies based on CKS Node labels for Kueue to use:
- The
infinibandtopology is for instance types that are a part of InfiniBand fabrics. See About GPU instances to find InfiniBand connected instances. - The
multinode-nvlink-ibtopology extends theinfinibandtopology to also include instance types with rack-scale NVLink. See About GPU instances to find NVLink connected rack instances. - The
hostnametopology is for instance types without InfiniBand fabrics. It prevents Kueue from admitting workloads when total capacity is sufficient but fragmented across Nodes.
Topology CRs appear in the cluster.
The following example configuration is an adjustment of the preceding one. It demonstrates how to use the Topology resources by referencing them in ResourceFlavor resources, which are then used by ClusterQueue and LocalQueue resources.
Each ResourceFlavor must select a disjoint set of Nodes. Use nodeLabels so that no Node matches more than one flavor. Overlapping selectors cause a Node to belong to multiple flavors, which leads to ambiguous quota accounting and unpredictable scheduling.
Match each topologyName to a Node pool whose hardware actually provides that topology. For example, pair the infiniband topology with an InfiniBand connected pool like B200, the multinode-nvlink-ib topology with a rack-scale NVLink pool like GB200, and the hostname topology with pools that lack InfiniBand.
Example jobs with topology constraints
Kueue provides several annotations for expressing topology constraints on a job. Two of the most common are:kueue.x-k8s.io/podset-required-topologyrequires that all Pods in a job are scheduled within the same topology domain. The job stays pending until a domain can fit it.kueue.x-k8s.io/podset-preferred-topologytreats the topology domain as best effort. Kueue tries to fit all Pods within the same domain, but falls back to spreading across domains if needed.