Skip to main content
SUNK Self-Service SUNK Self-Service is the operation mode in which CoreWeave deploys SUNK for you on CoreWeave Kubernetes Service (CKS). You declare a Slurm cluster as a single Kubernetes resource, and CoreWeave provisions and operates it for you, including compute nodes, the Slurm control plane, login Pods, shared storage, and user access. SUNK Self-Service is for teams that want SUNK without taking on the operational burden of installing the SUNK Helm charts, wiring up node pools, and keeping component versions in sync. If you’re evaluating SUNK, or are new to running Slurm on Kubernetes, SUNK Self-Service is the recommended entry point. For an introduction to the underlying technology, see About SUNK. To decide between this mode and SUNK Standard, see Choose a SUNK operation mode.

What SUNK Self-Service provisions

When you create a SUNK Self-Service cluster, CoreWeave sets up the following:
  • A CKS cluster, when you choose to provision a new one through the Cloud Console.
  • One Node Pool for each entry in your cluster’s nodes list, and one NodeSet for each compute entry, kept in sync automatically as you scale.
  • The Slurm control plane, including its MySQL database, running as Kubernetes Pods.
  • Login Pods that your users SSH into, sized per user or per group.
  • Shared storage volumes for home directories and additional mounts.
  • The supporting cluster components that Slurm needs, such as GPU device scheduling, deployed and managed for you. You don’t assemble or maintain these yourself.
  • User provisioning through SCIM and an Identity and Access Management (IAM) group for cluster access, when you use the Cloud Console path.
You configure all of this with a single SunkCluster Kubernetes resource. CoreWeave keeps the underlying components reconciled, scales the cluster when you change node counts, and reports overall progress with status conditions that you can inspect with kubectl.

What stays on your side

SUNK Self-Service provisions and maintains the cluster. The following remain your responsibility:
  • Submitting and managing Slurm jobs through sbatch, srun, and the rest of the Slurm CLI.
  • Building or selecting the container images that your Slurm jobs run through Pyxis and Enroot, separate from the slurmd, login, and control plane images that CoreWeave provides.
  • Designing your Slurm partition layout and quality-of-service policies beyond the defaults.
  • Managing the data you store in the shared volumes and the lifecycle of that data.
  • Adding users to the IAM groups that grant cluster access, and making sure each user has an SSH public key on their CoreWeave profile.
  • Monitoring and acting on job-level observability data, even though cluster-level health is reported through SunkCluster status conditions.

SUNK Self-Service compared to SUNK Standard

SUNK is CoreWeave’s technology for running Slurm on Kubernetes. A Kubernetes operator and a set of CRDs run inside the cluster to schedule Slurm jobs alongside Kubernetes workloads, but the components you deploy and manage depend on the operation mode. In SUNK Standard, you deploy and operate the sunk and slurm Helm charts yourself. SUNK Self-Service is the managed operation mode built on the same components. It adds:
  • A single, higher-level resource (SunkCluster) that captures the full cluster definition. You don’t manage the SUNK Helm charts, Node Pool configuration, or individual SUNK CRDs by hand.
  • An out-of-band operator, running in a CoreWeave-managed control plane, that provisions and maintains those underlying components.
  • Optional integration with the Cloud Console for guided cluster creation and user provisioning.
You can think of SunkCluster as the entry point and SUNK as the runtime. When you change a SunkCluster, CoreWeave reconciles the change into the SUNK components that actually run Slurm.

Architecture at a glance

SUNK Self-Service is built around an out-of-band operator that runs in a CoreWeave-managed control plane, outside your CKS cluster:
  • Out-of-band operator. It watches SunkCluster resources and reconciles them into the underlying SUNK components. CoreWeave operates and upgrades it. You don’t install or maintain it yourself. See SUNK Self-Service architecture.
  • Your CKS cluster. The Slurm control plane, login Pods, compute nodes, and SUNK CRDs all run inside your CKS cluster. This is where Slurm jobs execute and where your users connect.
When you submit a SunkCluster resource, the out-of-band operator creates and updates the Node Pools, NodeSets, the SlurmCluster resource, shared storage, and supporting components in your cluster. Your day-to-day workflow stays in kubectl against your CKS cluster. The out-of-band operator runs outside your cluster, so cluster upgrades and operator maintenance don’t require any action on your side.

Two ways to manage your cluster

You can create and manage SUNK Self-Service clusters in two ways:
  • Cloud Console: a guided form with a live YAML preview. The Console can also create a new CKS cluster, enable the SCIM API for user provisioning, and create an IAM group during cluster creation.
  • SunkCluster resource: apply YAML directly to your CKS cluster. This is the GitOps-friendly path. CKS cluster creation, SCIM setup, and IAM group creation are separate manual steps.
Both paths submit the same SunkCluster resource. The Console emits YAML that you can copy into source control, and a cluster created in the Console can be edited later with kubectl. For step-by-step instructions and a full prerequisites list, see Create a SUNK Self-Service cluster.

Next steps

Last modified on September 29, 2026