
In SUNK, Slurm nodes run in Kubernetes Pods. These aren’t the same as Kubernetes Nodes, which are the worker machines that run the Pods. To distinguish between them, this page capitalizes Kubernetes Nodes.
Operation modes
SUNK runs in two operation modes. They run the same Slurm, and your users connect to a login node and submit jobs the same way in either. SUNK Standard. You deploy and operate thesunk and slurm Helm charts, usually through ArgoCD or another GitOps pipeline. You set every chart value and own each upgrade. SUNK Standard runs SUNK 7.x and SUNK 8.0 and later, and it’s the mode that runs outside CoreWeave, as SUNK Anywhere.
SUNK Self-Service. CoreWeave provisions and operates the cluster from a single SunkCluster resource that you manage through the Cloud Console or kubectl. SUNK Self-Service requires SUNK 8.0 or later and is the recommended mode for a new cluster on CKS.
Choose a SUNK operation mode
What each mode gives you, and which one to pick for what your cluster has to do.
About SUNK Self-Service
What CoreWeave provisions and operates for you, and what stays on your side.
SUNK Standard in depth
Two Helm charts make up a SUNK Standard deployment. Thesunk chart installs the SUNK operator and the cluster-wide components it depends on, such as the NVIDIA device plugin and the MOCO MySQL operator. The slurm chart describes one Slurm cluster: its control plane, login pods, compute node definitions, Slurm configuration, and accounting database.
In SUNK 7.x, the slurm chart rendered every component of the cluster directly. In SUNK 8.0 and later, the chart renders a SlurmCluster resource for the control plane, login pods, and accounting, plus the compute NodeSets. The SUNK operator builds the cluster from them. Most keys keep their names, but the defaults you don’t set now come from the operator. Upgrading from 7.x to 8.0 isn’t a routine helm upgrade, so read the SUNK v8.0.0 release note first.
Start with these pages:
- Manage a SUNK Standard deployment with CI and GitOps for the ArgoCD app-of-apps pattern and the
NodeSetsync customization. Its examples predate SUNK 8.0, so check key names against the release note. - Configure compute nodes for SUNK Standard for node definitions, partitions, and images.
- Bind NodeSets to Node Pools for SUNK Standard to place Slurm nodes on specific CKS Node Pools.
- SUNK parameter reference and Slurm parameter reference for every chart value.
How SUNK works
The following pieces make SUNK work:- Compute and login nodes are Pods. Each Slurm compute node is a Pod running
slurmd, placed one-to-one on a Kubernetes Node. Login nodes are Pods with exposed IP addresses that users connect to over SSH. See Compute and login nodes. - NodeSets define compute. A NodeSet declares a group of Slurm nodes: image, resource requests, affinity, and replica count. Its controller creates the Pods and scales the set. See NodeSet.
- The Syncer keeps both sides in agreement. It mirrors state between Slurm nodes and their Pods in both directions, so a Pod that isn’t ready shows as a drained Slurm node. See Syncer.
- The SUNK Pod Scheduler shares Nodes. Pods that name it as their scheduler are placed through Slurm’s own scheduling logic, on the same Nodes as Slurm jobs and under Slurm’s priority and preemption rules. See SUNK Pod Scheduler.
- SlurmCluster carries cluster-wide Slurm configuration and generates
topology.conffrom Node labels. See SlurmCluster. - Slurm runs from images CoreWeave publishes. The
slurmdimages ship per CUDA version, built on NCCL test images. Controller and login images ship per Ubuntu version. See Slurm images.
Key features
SUNK supports the following:- Containerized jobs through Pyxis and enroot, including Docker workflows. See Use Docker in SUNK.
- Topology-aware scheduling with Slurm’s tree and block plugins. See Topology and block scheduling in Slurm.
- GPU straggler detection for multi-node training jobs. See Introduction to GPU straggler detection.
- Prolog and Epilog scripts that prepare a node before a job and clean up after it. See Run Prolog and Epilog scripts on SUNK.
- s6 scripts that install software and configure Slurm nodes at startup. See Run custom scripts with s6.
- Task plugins and cgroups for CPU, memory, and GPU binding. See Manage resource binding with task plugins.
- User provisioning from CoreWeave Identity and Access Management (IAM) through SCIM, so POSIX identities, SSH keys, and Slurm accounts follow your identity provider. See Provision users in SUNK.
- On SUNK Standard only: custom images, Lua plugin files, and LDAP directories. Choose a SUNK operation mode has the full list.
Next steps
- Choose a SUNK operation mode if you haven’t picked one yet.
- Create a SUNK Self-Service cluster to start a SUNK Self-Service cluster from the Cloud Console or the
SunkClusterCR. - Manage a SUNK Standard deployment with CI and GitOps to deploy SUNK Standard with ArgoCD.
- Train on SUNK to set up a cluster, submit a training job, and monitor it.