SunkCluster resource into the resources that make up a cluster. The in-cluster operator runs inside your CoreWeave Kubernetes Service (CKS) cluster and turns those resources into a running Slurm cluster. This page explains what each operator does, where it runs, and how the two divide the work.
For what SUNK Self-Service provisions and the two ways to create a cluster, see About SUNK Self-Service. For the order in which a cluster comes up and how to read its status conditions, see SUNK Self-Service cluster lifecycle and reconciliation.
Two operators
The following table summarizes where each operator runs and what it manages. The sections after the table describe each operator in more detail.Out-of-band operator
The out-of-band operator reconcilesSunkCluster resources. It runs outside your CKS cluster, and CoreWeave deploys, operates, and upgrades it. You don’t install or maintain the out-of-band operator yourself.
When you apply a SunkCluster, the out-of-band operator does two kinds of work in your cluster:
- Installs SUNK. It applies the SUNK CRDs, including
NodeSetandSlurmCluster, creates thecw-sunknamespace with the permissions SUNK needs, and deploys the in-cluster operator there. - Builds your cluster’s resources. It creates one Node Pool for each entry in the
SunkClusterresource’snodeslist, and one NodeSet for each compute entry. A control-plane entry, one withcontrolPlane: true, gets a Node Pool but no NodeSet. It also creates the shared storage volumes, deploys supporting components such as GPU device scheduling, and generates aSlurmClusterresource from yourSunkCluster. It creates the NodeSets and theSlurmClusterin the same namespace as theSunkCluster, and theSunkClusterowns them.
In-cluster operator
The in-cluster operator runs inside your CKS cluster as thesunk-controller-manager Deployment in the cw-sunk namespace. It reconciles the SlurmCluster resource into the login Pods your users connect to and the Slurm control plane: slurmctld, the accounting database, and the Syncer that keeps Slurm node state aligned with Kubernetes. It runs the scheduler only when spec.scheduler.enabled is true in your SunkCluster, and that field defaults to false. It also reconciles each NodeSet into the slurmd Pods, one per Kubernetes Node.
The Node Controller and the Pod Controller are part of the same Deployment. The sunkVersion field in your SunkCluster sets which version of the in-cluster operator runs. See Upgrade a SUNK Self-Service cluster.
The in-cluster operator is the SUNK operator, the same operator that the
sunk Helm chart installs in SUNK Standard. The out-of-band operator is the part that SUNK Self-Service adds to it.The SlurmCluster is an internal interface
You configure a SUNK Self-Service cluster through the SunkCluster resource. The out-of-band operator translates it into the SlurmCluster resource that the in-cluster operator consumes, and reapplies the fields it manages every time it reconciles. A direct edit to one of those fields in the generated SlurmCluster doesn’t last, so make the change in the SunkCluster instead. For the fields you can set, see the SunkCluster reference.
The SlurmCluster is still useful for troubleshooting. Both resources report status conditions. When the SunkCluster condition SlurmClusterAvailable isn’t True, the SlurmCluster conditions show which part of the Slurm control plane is pending. See Track progress with status conditions.
Where each component runs
The CoreWeave-managed control plane holds only the out-of-band operator. Everything else runs in your CKS cluster:- In the
cw-sunknamespace: the in-cluster operator, and supporting components that the out-of-band operator deploys, such as the NVIDIA device plugin. - In your
SunkClusterresource’s namespace: the NodeSets, theSlurmCluster, and the Slurm control plane and login Pods the in-cluster operator creates from it. - On your CKS Nodes: the Slurm compute nodes, each a
slurmdPod on one Kubernetes Node. See Compute and login nodes.
Related topics
- About SUNK Self-Service: what SUNK Self-Service provisions and what stays your responsibility.
- SUNK Self-Service cluster lifecycle and reconciliation: creation order, status conditions, and ongoing reconciliation.
- Get started with SUNK Self-Service using the SunkCluster CR: create a cluster from a manifest, end to end.
SunkClusterreference: every field and status condition.- SlurmCluster: the
SlurmClusterCRD and the topology files it generates.