Gang scheduling is currently available for SUNK 7.6.0 and later in the 7.x series. Support for the 8.x series support is planned for a future release.
Prerequisites
Before you submit a gang, confirm the following requirements:- Your cluster runs SUNK 7.6.0 or later in the 7.x series. If you need to upgrade, follow the SUNK upgrade procedure. Support for the 8.x series support is planned for a future release, as noted above.
- The SUNK Pod Scheduler is enabled and watches the namespace where you create the workload.
- You have
kubectlaccess to create and inspect workloads in that namespace, and access to a Slurm login node to inspect allocations. - You know the configured scheduler name. Replace
[SCHEDULER-NAME]in each manifest and[WORKLOAD-NAMESPACE]in each command with your values. - The cluster has Slurm-managed nodes that can satisfy the complete gang’s resource and placement requirements.
terminationGracePeriodSeconds: 20. Confirm that this is below your scheduler’s termination threshold.
Submit a Kubernetes Job
This example sets bothparallelism and completions to 2, so SUNK requests one two-node Slurm allocation for the Job.
-
Save the following manifest in a
job-gang.yamlfile and replace[SCHEDULER-NAME]:job-gang.yaml -
Submit the Job:
Submit a LeaderWorkerSet
This example creates two independent replica groups withreplicas: 2. Each group has two Pods because leaderWorkerTemplate.size: 2 includes both the leader and one worker. SUNK requests a separate two-node allocation for each group.
-
Save the following manifest in a
lws-multigroup-gang.yamlfile and replace both occurrences of[SCHEDULER-NAME]:Keep the defaultlws-multigroup-gang.yamlstartupPolicy: LeaderCreated. It lets LWS create the workers before the leader becomes ready, so the complete gang can form before SUNK binds the leader. Don’t setLeaderReady. With that policy, LWS waits for aReadyleader before it creates any worker, and the leader can’t becomeReadyuntil the whole gang exists, so the group never schedules. The example also usesRecreateGroupOnPodRestartand the default rollout settings. OtherrolloutStrategyvalues, such as amaxSurgeabove zero, haven’t been validated with gang scheduling. -
Submit the LeaderWorkerSet:
Verify the allocations
Check the Pods’ node assignments and Slurm job IDs, then inspect the allocation from Slurm.Check Pod placement
For the Job example, run:Check Slurm job state
On a Slurm login node, replace[JOB-ID] with one of the IDs from the JOBID column and inspect the job:
JobState is RUNNING, NumNodes is 2, and NodeList contains the two allocated Slurm nodes. CPU and memory totals depend on your requests and Slurm configuration.
Optional: Configure topology segments
For GB200 and GB300 workloads that use Slurm’s Topology/Block Plugin, setsunk.coreweave.com/segment on the Pod template to request a topology segment size. For a Job, the annotation belongs under spec.template.metadata.annotations. For an LWS, apply it to the leader and worker Pod templates.
For example, this Pod-template metadata excerpt requests segments of two nodes:
Limitations and lifecycle
Plan gang workloads around the following constraints:- Supported workload types. SUNK recognizes Kubernetes Jobs and LWS replica groups as gangs. Standalone Pods and unsupported controllers continue to use one Slurm placeholder job per Pod. SUNK doesn’t require a PodGroup resource.
- Fixed member count. A gang’s Slurm allocation can’t grow or shrink in place. In an LWS, changing
replicasadds or removes entire groups. Adding or removing workers within an existing leader-and-worker group isn’t supported. - Homogeneous resources. Members of a gang must have compatible CPU, memory, and GPU requirements. A CPU-only leader with GPU workers isn’t supported. The examples set the same requests for all members. To keep them identical with less to maintain, omit
leaderTemplatefrom the LWS. LWS then usesworkerTemplatefor the leader Pod as well. - LWS resource updates. When a group rolls out updated Pod templates with incompatible resource requirements, SUNK cancels the old allocation and creates a replacement after teardown. A group held on its old revision by a rollout partition keeps its compatible allocation until that group updates.
- LWS group recreation. With
RecreateGroupOnPodRestart, a recreated group can reuse its allocation if its size and resource requirements remain compatible. - LWS startup policy. Gang scheduling requires the default
startupPolicy: LeaderCreated. WithLeaderReady, LWS creates no workers until the leader is Ready, so the gang never completes and the group stays Pending. Rollout settings other than the defaults haven’t been validated with gang scheduling. - Slurm scheduling policy. A complete gang can still wait for capacity, queue priority, fair-share, or reservations. Gang scheduling doesn’t bypass those policies.
Troubleshoot gang scheduling
Use workload events to distinguish a gang that’s still forming from a queued or failed Slurm allocation:-
SunkGangWaitingreports how many gang members are present and expected. Confirm that Kubernetes created every member and that every Pod uses the configured SUNK scheduler name. If an LWS group has a leader but no workers, check that the LeaderWorkerSet doesn’t setstartupPolicy: LeaderReady. -
SunkGangValidationFailedreports incompatible LWS requirements. Check that the leader and worker templates have compatible CPU, memory, and GPU requirements and use the same scheduler name. -
A Slurm job in
PENDINGcan be waiting for the complete node set. Check its reason withscontrol show job [JOB-ID]. See Why is my Slurm job stuck in PENDING?. -
If a Job’s allocation ends before the Kubernetes Job completes, SUNK deletes the Job. If an LWS group’s allocation ends, the group can form again with a new allocation. Inspect workload events and Slurm accounting to determine why the allocation ended:
Clean up
Delete the example workload objects to release their allocations:Terminating while Slurm completes teardown.