Skip to main content
SUNK Standard SUNK Self-Service SUNK logo SUNK is CoreWeave’s system for running Slurm on Kubernetes. Slurm nodes run as Kubernetes Pods on CoreWeave Kubernetes Service (CKS), so Slurm jobs and Kubernetes workloads share the same compute. The Slurm cluster also gains the Node lifecycle management, observability, and GitOps tooling that Kubernetes brings. Without SUNK, a team picks Slurm or Kubernetes, or runs both and splits its GPUs between them. With SUNK, one pool of compute serves batch training jobs through Slurm and long-running or bursty services through Kubernetes. The SUNK Pod Scheduler arbitrates between them on the same Nodes.
In SUNK, Slurm nodes run in Kubernetes Pods. These aren’t the same as Kubernetes Nodes, which are the worker machines that run the Pods. To distinguish between them, this page capitalizes Kubernetes Nodes.

Operation modes

SUNK runs in two operation modes. They run the same Slurm, and your users connect to a login node and submit jobs the same way in either. SUNK Standard. You deploy and operate the sunk and slurm Helm charts, usually through ArgoCD or another GitOps pipeline. You set every chart value and own each upgrade. SUNK Standard runs SUNK 7.x and SUNK 8.0 and later, and it’s the mode that runs outside CoreWeave, as SUNK Anywhere. SUNK Self-Service. CoreWeave provisions and operates the cluster from a single SunkCluster resource that you manage through the Cloud Console or kubectl. SUNK Self-Service requires SUNK 8.0 or later and is the recommended mode for a new cluster on CKS.

Choose a SUNK operation mode

What each mode gives you, and which one to pick for what your cluster has to do.

About SUNK Self-Service

What CoreWeave provisions and operates for you, and what stays on your side.

SUNK Standard in depth

Two Helm charts make up a SUNK Standard deployment. The sunk chart installs the SUNK operator and the cluster-wide components it depends on, such as the NVIDIA device plugin and the MOCO MySQL operator. The slurm chart describes one Slurm cluster: its control plane, login pods, compute node definitions, Slurm configuration, and accounting database. In SUNK 7.x, the slurm chart rendered every component of the cluster directly. In SUNK 8.0 and later, the chart renders a SlurmCluster resource for the control plane, login pods, and accounting, plus the compute NodeSets. The SUNK operator builds the cluster from them. Most keys keep their names, but the defaults you don’t set now come from the operator. Upgrading from 7.x to 8.0 isn’t a routine helm upgrade, so read the SUNK v8.0.0 release note first. Start with these pages:

How SUNK works

The following pieces make SUNK work:
  • Compute and login nodes are Pods. Each Slurm compute node is a Pod running slurmd, placed one-to-one on a Kubernetes Node. Login nodes are Pods with exposed IP addresses that users connect to over SSH. See Compute and login nodes.
  • NodeSets define compute. A NodeSet declares a group of Slurm nodes: image, resource requests, affinity, and replica count. Its controller creates the Pods and scales the set. See NodeSet.
  • The Syncer keeps both sides in agreement. It mirrors state between Slurm nodes and their Pods in both directions, so a Pod that isn’t ready shows as a drained Slurm node. See Syncer.
  • The SUNK Pod Scheduler shares Nodes. Pods that name it as their scheduler are placed through Slurm’s own scheduling logic, on the same Nodes as Slurm jobs and under Slurm’s priority and preemption rules. See SUNK Pod Scheduler.
  • SlurmCluster carries cluster-wide Slurm configuration and generates topology.conf from Node labels. See SlurmCluster.
  • Slurm runs from images CoreWeave publishes. The slurmd images ship per CUDA version, built on NCCL test images. Controller and login images ship per Ubuntu version. See Slurm images.

Key features

SUNK supports the following:

Next steps

Last modified on September 22, 2026