Skip to main content
Gang scheduling is currently available for SUNK 7.6.0 and later in the 7.x series. Support for the 8.x series support is planned for a future release.
CoreWeave SUNK gang scheduling reserves resources for a complete group of Kubernetes Pods before binding any member to a node. Use it for fixed-size workloads that need every member to run before they can make progress. This guide shows workload owners how to submit a Kubernetes Job or LeaderWorkerSet and verify its Slurm allocation. A Kubernetes Job forms one gang. Each LeaderWorkerSet (LWS) replica group forms a separate gang. After all expected Pods exist, SUNK requests one Slurm job with one node per Pod. Members wait unbound until Slurm grants the complete allocation. Kubernetes then starts the Pods independently, so they don’t necessarily start at the same instant.

Prerequisites

Before you submit a gang, confirm the following requirements:
  • Your cluster runs SUNK 7.6.0 or later in the 7.x series. If you need to upgrade, follow the SUNK upgrade procedure. Support for the 8.x series support is planned for a future release, as noted above.
  • The SUNK Pod Scheduler is enabled and watches the namespace where you create the workload.
  • You have kubectl access to create and inspect workloads in that namespace, and access to a Slurm login node to inspect allocations.
  • You know the configured scheduler name. Replace [SCHEDULER-NAME] in each manifest and [WORKLOAD-NAMESPACE] in each command with your values.
  • The cluster has Slurm-managed nodes that can satisfy the complete gang’s resource and placement requirements.
For the LWS example, if the LWS controller and custom resource definition aren’t installed, have a cluster administrator install them. These examples use LWS v0.8.0:
The examples use small CPU and memory requests to demonstrate scheduling. Adjust the image, requests, node selection, and tolerations for your workload and nodes. Keep the members’ resource requirements homogeneous. For GPU workloads and node sharing, see Manage resources with the SUNK Pod Scheduler. Each example sets terminationGracePeriodSeconds: 20. Confirm that this is below your scheduler’s termination threshold.

Submit a Kubernetes Job

This example sets both parallelism and completions to 2, so SUNK requests one two-node Slurm allocation for the Job.
  1. Save the following manifest in a job-gang.yaml file and replace [SCHEDULER-NAME]:
    job-gang.yaml
  2. Submit the Job:
When Slurm grants the allocation, the two Pods run on distinct nodes and share one Slurm job ID. Follow Verify the allocations to check their placement.

Submit a LeaderWorkerSet

This example creates two independent replica groups with replicas: 2. Each group has two Pods because leaderWorkerTemplate.size: 2 includes both the leader and one worker. SUNK requests a separate two-node allocation for each group.
  1. Save the following manifest in a lws-multigroup-gang.yaml file and replace both occurrences of [SCHEDULER-NAME]:
    lws-multigroup-gang.yaml
    Keep the default startupPolicy: LeaderCreated. It lets LWS create the workers before the leader becomes ready, so the complete gang can form before SUNK binds the leader. Don’t set LeaderReady. With that policy, LWS waits for a Ready leader before it creates any worker, and the leader can’t become Ready until the whole gang exists, so the group never schedules. The example also uses RecreateGroupOnPodRestart and the default rollout settings. Other rolloutStrategy values, such as a maxSurge above zero, haven’t been validated with gang scheduling.
  2. Submit the LeaderWorkerSet:
Each group can receive its allocation independently. Members within a group share a Slurm job ID, while the two groups have different IDs.

Verify the allocations

Check the Pods’ node assignments and Slurm job IDs, then inspect the allocation from Slurm.

Check Pod placement

For the Job example, run:
After the Pods start, the output is similar to the following:
For the LWS example, include the group index:
After both groups start, the output is similar to the following:
Members of one gang occupy distinct nodes. Separate groups can share nodes if the node sharing policy permits it.

Check Slurm job state

On a Slurm login node, replace [JOB-ID] with one of the IDs from the JOBID column and inspect the job:
For a running two-member gang, check that JobState is RUNNING, NumNodes is 2, and NodeList contains the two allocated Slurm nodes. CPU and memory totals depend on your requests and Slurm configuration.

Optional: Configure topology segments

For GB200 and GB300 workloads that use Slurm’s Topology/Block Plugin, set sunk.coreweave.com/segment on the Pod template to request a topology segment size. For a Job, the annotation belongs under spec.template.metadata.annotations. For an LWS, apply it to the leader and worker Pod templates. For example, this Pod-template metadata excerpt requests segments of two nodes:
Choose a segment size that fits your gang and available topology blocks. For segment sizing and troubleshooting, see Topology and block scheduling in Slurm.

Limitations and lifecycle

Plan gang workloads around the following constraints:
  • Supported workload types. SUNK recognizes Kubernetes Jobs and LWS replica groups as gangs. Standalone Pods and unsupported controllers continue to use one Slurm placeholder job per Pod. SUNK doesn’t require a PodGroup resource.
  • Fixed member count. A gang’s Slurm allocation can’t grow or shrink in place. In an LWS, changing replicas adds or removes entire groups. Adding or removing workers within an existing leader-and-worker group isn’t supported.
  • Homogeneous resources. Members of a gang must have compatible CPU, memory, and GPU requirements. A CPU-only leader with GPU workers isn’t supported. The examples set the same requests for all members. To keep them identical with less to maintain, omit leaderTemplate from the LWS. LWS then uses workerTemplate for the leader Pod as well.
  • LWS resource updates. When a group rolls out updated Pod templates with incompatible resource requirements, SUNK cancels the old allocation and creates a replacement after teardown. A group held on its old revision by a rollout partition keeps its compatible allocation until that group updates.
  • LWS group recreation. With RecreateGroupOnPodRestart, a recreated group can reuse its allocation if its size and resource requirements remain compatible.
  • LWS startup policy. Gang scheduling requires the default startupPolicy: LeaderCreated. With LeaderReady, LWS creates no workers until the leader is Ready, so the gang never completes and the group stays Pending. Rollout settings other than the defaults haven’t been validated with gang scheduling.
  • Slurm scheduling policy. A complete gang can still wait for capacity, queue priority, fair-share, or reservations. Gang scheduling doesn’t bypass those policies.

Troubleshoot gang scheduling

Use workload events to distinguish a gang that’s still forming from a queued or failed Slurm allocation:
Use the following events and job states to identify the problem:
  • SunkGangWaiting reports how many gang members are present and expected. Confirm that Kubernetes created every member and that every Pod uses the configured SUNK scheduler name. If an LWS group has a leader but no workers, check that the LeaderWorkerSet doesn’t set startupPolicy: LeaderReady.
  • SunkGangValidationFailed reports incompatible LWS requirements. Check that the leader and worker templates have compatible CPU, memory, and GPU requirements and use the same scheduler name.
  • A Slurm job in PENDING can be waiting for the complete node set. Check its reason with scontrol show job [JOB-ID]. See Why is my Slurm job stuck in PENDING?.
  • If a Job’s allocation ends before the Kubernetes Job completes, SUNK deletes the Job. If an LWS group’s allocation ends, the group can form again with a new allocation. Inspect workload events and Slurm accounting to determine why the allocation ended:
For Pod admission failures or invalid annotations, see Troubleshoot a SUNK-scheduled Kubernetes Pod.

Clean up

Delete the example workload objects to release their allocations:
SUNK cancels the associated Slurm allocations. Pods can remain in Terminating while Slurm completes teardown.
Last modified on October 7, 2026