CoreWeave Serverless sandboxes are in public preview.
CPU-only supportGPU sandboxes are in private preview and aren’t yet supported by the SUNK Pod Scheduler integration. See GPU sandboxes for supported placement options and how to request access.CPU-only sandboxes can land on idle CPU nodes, share CPU nodes with other sandboxes or Slurm jobs, and share CPU resources on GPU nodes where other workloads are running. See Known limitations for details.
Prerequisites
- A CKS cluster with SUNK deployed and the SUNK Pod Scheduler enabled.
- A CoreWeave sandbox runner deployed on the same cluster.
- The CoreWeave Intelligent CLI (
cwic) v1.44.0 or later, installed and authenticated. See Deploy and manage a runner for setup details. - The
sandbox_adminrole to read and update the runner’s policy, and permission to update the Slurm deployment. cwsandbox >= 1.0.0, installed and authenticated, for the v1 sandbox API.kubectlandyqfor cluster inspection. The HTTP examples also usecurl,jq, and a CoreWeave API token in$TOKEN.
[RUNNER-ID] throughout this page with the ID of the runner on your SUNK cluster.
Step 1: Verify the SUNK Pod Scheduler
Validate that your SUNK deployment is configured to work with CoreWeave sandboxes.-
Verify the scheduler is running:
If no pods are returned, enable the scheduler in your Slurm Helm values by setting
scheduler.enabled: true. See Enable the scheduler for details. -
Look up the scheduler name. You set it on the runner policy in Step 2. On SUNK 8.0 and later, the scheduler reads its name from a ConfigMap:
On SUNK 7.x, the name is a
--scheduler-nameargument on the scheduler Pod instead. See Look up the scheduler configuration for the underlying behavior. -
Look up the
KillWaitvalue. You use it to check Pod termination:Look for--slurm-kill-waitin the output: -
Configure Slurm’s termination window for sandbox Pods. SUNK requires
terminationGracePeriodSecondsto be strictly less than--slurm-kill-waitminus 5 seconds. The policy doesn’t acceptbase.spec.terminationGracePeriodSeconds. For a Pod with Kubernetes’ default 30-second grace period, setKillWaitabove 35 seconds. For example, merge this into your Slurm Helm values and apply the update through your normal SUNK deployment process:This changes Slurm’s termination window for jobs across the cluster. Verify that the scheduler now reports--slurm-kill-wait=60susing the preceding command. See Slurm configuration parameters and Set the termination grace period. -
Confirm the scheduler scope covers the sandbox namespaces. The runner manages namespace selection. It isn’t a policy field. Set
scheduler.scope.type: clusterto watch all sandbox namespaces, including ones created for new users. If you usescheduler.scope.type: namespace, include each sandbox namespace inscheduler.scope.namespacesand update the list as new namespaces are created. Apply scope changes through Helm so the scheduler also receives the required RBAC permissions in those namespaces. See Enable the scheduler. -
Lower
slurmdresource requests on NodeSets you want to share with sandboxes. This is a change on the Slurm side, in your Slurm Helm values, not in any sandbox configuration. The default NodeSet resource requests consume most of the node’s allocatable capacity in Kubernetes, leaving no room for sandbox pods. Without this change, sandbox pods are rejected withOutOfMemoryorOutOfcpuerrors. See Manage resources with the SUNK Pod Scheduler for configuration details.
Step 2: Configure the runner policy
Every runner carries one policy. Setbase.spec.schedulerName in that policy to route its sandbox Pods through SUNK. You don’t create or bind a separate policy resource.
Merge the following fragment into the runner’s existing policy. Replace slurm-scheduler with the scheduler name from Step 1:
- CLI
- curl
Read the current policy, then open it in Add
$EDITOR:schedulerName under base.spec, keeping the rest of the document. Save and close the editor to apply it.If you maintain the complete policy in a file, validate and apply that file:constraints.network field and the sandbox’s declared egress govern network access. Keep the runner’s existing network constraints when adding SUNK scheduling. See Configure a sandbox policy.
Policy changes can take time to propagate to sandbox creation. Before relying on a new scheduler or pinned annotation, create a sandbox and verify its Pod fields as described in Pods remain pending. A policy read can show the new value before new sandboxes use it.
For the complete schema, see Policy reference. For runner setup, see Deploy and manage a runner.
Step 3: Create sandboxes
Setplacement_mode="cks" and runner_ids to select the runner whose policy you configured. The runner policy supplies the scheduler automatically.
SandboxDefaults:
Resource requests and Slurm accounting
SUNK reads the pod’s resource requests (not limits) and converts them to Slurm job parameters:
Slurm uses these values for scheduling decisions and
sacct accounting. SUNK does not require any particular Quality of Service class. Guaranteed (requests equal limits) and Burstable (requests lower than limits) both work.
If you set requests lower than limits with ResourceOptions, the pod can burst up to the limits when capacity is free, but Slurm only sees the requests. For example, a sandbox configured with:
sacct as a 500m CPU, 512Mi memory job, even though the sandbox can use up to 2 CPUs and 2Gi when the node has room. See Resources for the full ResourceOptions reference.
Size the requests to match what your sandbox workloads need, leaving enough room on the target nodes for the slurmd requests you lowered in Step 1. For the underlying rules, see Set resource requests and Manage resources with the SUNK Pod Scheduler.
Step 4: Control placement with Slurm annotations
To control sandbox placement, set SUNK annotations on the sandbox at launch time. The following example pins the sandbox to thehpc-prod partition:
Common annotations
All annotations share thesunk.coreweave.com/ prefix. The annotations commonly used to control sandbox placement are:
Passing a username instead of a numeric UID to the
user-id annotation causes a blocking error that prevents Slurm from scheduling the sandbox. To find the numeric UID from a Slurm login node, replace [USERNAME] with the Linux username to look up:
Enforce annotations in the policy
To set a Slurm job parameter for every sandbox on the runner, add its annotation tobase.metadata.annotations. For example, merge this fragment into the existing policy to use the sandboxes partition:
base.metadata.annotations and ensure constraints.metadata.denied_annotation_prefixes permits it. For example, leave sunk.coreweave.com/user-id unpinned when clients need to match the user of a training job.
Match the Slurm user from a training job
Training jobs often run with--exclusive=user to claim entire nodes for a single user. This prevents other users’ jobs from competing for resources on those nodes while still allowing the same user to run additional jobs there, such as sandboxes that use spare CPU alongside GPU training.
By default, SUNK placeholder jobs run as root (UID 0). Because root is a different user than the one who submitted the training job, Slurm does not place the sandbox placeholder on the exclusive node.
When training code uses the cwsandbox Python client to create sandboxes from within a running Slurm job, it can read the job’s Slurm user ID from the environment and pass it as an annotation. This ensures the sandbox placeholder jobs are submitted under the same Slurm user as the training job, allowing Slurm to place them on the same exclusive nodes:
user-id annotation must be a numeric Linux UID, not a username. When set, SUNK also defaults the group-id to the same value. Set sunk.coreweave.com/group-id separately if the group ID differs.
Troubleshooting
Use the following sections to diagnose common issues when running sandboxes through the SUNK Pod Scheduler.Pods remain pending
Inspect the sandbox Pod to confirm the runner applied the scheduler name and to check its termination grace period:SCHEDULER matches the SUNK scheduler, its scope includes NAMESPACE, and GRACE is strictly less than the scheduler’s kill wait minus 5 seconds. A 30-second grace period fails with the default 30-second kill wait. If the policy pins a partition, confirm that PARTITION matches it.
If the scheduler name or pinned annotations are missing or incorrect, read back the runner policy and confirm that the client selected the intended runner. If newly created Pods still use previous values, allow more time for policy propagation and retry with another sandbox.
Placeholder jobs temporarily in completing state
When a sandbox stops, its Slurm placeholder job spends 30 to 60 seconds in theCG (completing) state while cleanup scripts run and the node is released back to the pool. This is normal Slurm behavior and does not affect other jobs running on the same node. See Slurm job states for the full state reference.
Sandboxes not landing where expected
Slurm determines sandbox placement based on the annotations the policy sets or the client passes. If sandboxes aren’t landing on the expected nodes, verify the Slurm job parameters. SUNK creates a placeholder Slurm job for each sandbox pod with the name<namespace>/<pod-name>. The pod name includes the sandbox ID, which is available from the client as sb.sandbox_id. Find the placeholder job and inspect its parameters by searching for the sandbox ID:
[SANDBOX-ID] with the value of sb.sandbox_id from the Python client.