Ready, Running, or Drain, and lets operations made on either side update the state of the other. The Syncer is deployed along with each Slurm cluster.
Information flow and reconcile operations
The Syncer supports several flows of information and operations. The following sections describe each flow and the conditions that trigger it.Slurm drains from Kubernetes
When certain conditions happen on the Kubernetes side that make the respective Slurm node either inoperable or undesired for continued Slurm job scheduling, the Syncer propagates these conditions as a drain on the Slurm node. When the condition clears, the Syncer then removes the drain. A drain from Kubernetes uses thek8s: prefix to indicate within Slurm that the drain originated on the Kubernetes side.
The Syncer only removes or updates drain reasons that are prefixed with
k8s:. A non-prefixed drain is left as is.- The Kubernetes Pod associated with this Slurm node is not ready.
- The Kubernetes Pod associated with this Slurm node has been deleted.
- The Kubernetes Pod associated with this Slurm node is pending deletion. See NodeSet Controller.
- The Kubernetes Pod associated with this Slurm node is Cordoned. See Pod Controller.
Slurm downed nodes
Nodes in Slurm can be setdown by three routes:
- Automatically by the Slurm Controller.
- Manually by the user from within Slurm.
- Upon Pod deletion in Kubernetes.
The default (and recommended) configuration of the Slurm chart uses
ReturnToService=2, which automatically resumes any down node that starts communicating with the Slurm controller. To change this behavior, adjust this value. A non-default value requires the user to take action within Slurm after a Pod is updated in Kubernetes, before the node is usable in Slurm.Slurm node deletion
The Syncer updates the state within Slurm following NodeSlice changes. For example, when nodes are removed from NodeSets, these changes are reflected in the underlying NodeSlice(s). After detecting removed NodeSlice entries, the Syncer requests deletion of the corresponding Slurm nodes. Enable this functionality using the Slurm chart option .syncer.config.syncer.slurmNodeCleanUp.Slurm node status
The Syncer converts the current running, responding, and drain Slurm states into conditions and labels on the Pod. The conditionsSlurmDrain, SlurmRunning, and SlurmNotResponding mirror the state within Slurm. The Syncer propagates the reason for the drain into the Message for the SlurmDrain condition.
The labels sunk.coreweave.com/running, sunk.coreweave.com/drain, and sunk.coreweave.com/not_responding aid dashboards that use metrics from kube-state-metrics, and aren’t used for any logic within the Operator or Syncer. When a Slurm node is drained from within Slurm, that drain propagates up to the Node as well. For more information, see NodeController.
NHC drain and HPC verification
Although this feature was implemented for CoreWeave’s particular environment, it can be used for similar workflows in other environments.
verify-undrain anywhere in the node’s drain reason. When this string is present, the Syncer checks the HPCVerification condition of the Pod to see if a newer verification pass has happened since the drain.
Extra field
Nodes in Slurm have anExtra field that can be used to store user-specified information. SUNK uses this to store information that provides visibility within Slurm to conditions on the Kubernetes side. The Syncer manages updates to the Extra field to reflect the information. The information is stored as JSON in the Extra field to allow for parsing and manipulation.
Users can add more information into the
Extra field in Slurm. The Extra field must contain valid JSON, or the contents are cleared and replaced at the next synchronization. When there are conflicts, JSON fields set by the Syncer overwrite those set by the user.Hook API
The hook API provided by the Syncer lets events in Slurm directly trigger operations within Kubernetes. Some of these hooks facilitate blocking synchronization or immediate actions. The Syncer provides several hooks for node objects, described in the following sections.Pre-hook
The pre-hook endpoint ensures that other jobs running outside Slurm on the Node are removed before the Slurm jobs start. It also begins the state propagation that triggers the Pod Controller and Node Controller to perform further actions.Reboot
This endpoint is only available when the Syncer has permissions to perform operations on the Nodes. The Syncer node permissions are set with the Slurm chart option .syncer.nodePermissions.enabled.
PhaseState condition on the Node along with the associated reason production-powerreset, which then triggers other Node management tooling to reboot the Node. To modify the condition type and associated reason, use .syncer.hooksAPI.nodeRebootCondition and .syncer.hooksAPI.nodeRebootReason.
Metrics
The Syncer provides a scrapeable metrics endpoint, which exposes metrics for the nodes, jobs, and the overall Slurm cluster. The PodMonitor deployed with the SUNK chart labels all metrics with their associated Slurm cluster using theslurm_cluster label. The Syncer also exports additional metrics for the standard Go runtime and controller runtime.
Labels added by the code are shown in the following list. More labels can be added by the scrape configuration.
The Syncer applies the following labels:
- account: Slurm account name
- id: Slurm job ID
- name: Slurm job name
- node: Slurm node name
- partition: Slurm partition
- state: Slurm job current state
- user: Slurm user name
- message_type: Slurm RPC message type