Skip to main content
This guide shows how to create a sandbox with one or more GPUs, confirm the GPU is visible from inside the sandbox, and set the CPU and memory requests that go with it. GPU sandboxes run in either placement mode: on serverless capacity that CoreWeave operates, or on a CoreWeave Kubernetes Service (CKS) cluster you own. The two modes differ in how the sandbox is isolated, which GPUs you can get, and who sets the limits, so this page covers them in separate sections. For CPU-only sandboxes, see Get started. To compare serverless and CKS capacity, see Sandbox placement and spillover.
GPU sandboxes are in private preview and require your organization to be allowlisted, even if it already uses CPU-only sandboxes. To request access, contact your account team or email forge-support@coreweave.com. Until your organization is allowlisted, requests to create GPU sandboxes fail with CWSANDBOX_GPU_NOT_ALLOWED.
For concurrent sandbox quotas and resource limits, see Limits and quotas.

Isolation depends on the placement mode

GPU virtualization is supported only on serverless capacity. On a CKS cluster you own, GPU sandboxes aren’t virtualized.
GPU sandboxes on your own CKS cluster run under the default runtime class (runc), with no virtualization. They have the same isolation as any other container on the Node and can’t be considered a secure container runtime. Don’t use CKS GPU sandboxes to run untrusted code, such as code generated by a model or submitted by users. To run untrusted code on a GPU, use serverless capacity.

Before you begin

You need the following:
  • GPU sandboxes enabled for your organization.
  • A CoreWeave API access token with the SANDBOX_USER Identity and Access Management (IAM) action. See Choose and get credentials.
  • The Python client, cwsandbox 1.14.2 or later:
The TypeScript client doesn’t expose GPU resources yet. Use the Python client for GPU sandboxes.

How GPU requests work

The same rules apply in both placement modes:
  • A GPU request reserves GPUs and nothing else. Set CPU and memory explicitly, sized for the work you’ll run alongside the GPU. Omitted CPU and memory values can use the policy’s defaultCpu and defaultMemory, but the resolved allocation must still satisfy the policy’s resource requirements.
  • Sandboxes get whole GPUs. GPUs aren’t shared, partitioned, or time-sliced between sandboxes.
  • A sandbox runs on a single Node. On serverless capacity, a sandbox gets exactly 1 GPU. On your CKS cluster, the largest sandbox is one full Node, and how many GPUs that is depends on the GPU type. On a Node with 8 GPUs, a CKS sandbox can request anywhere from 1 GPU up to 8. Requests for several GPUs need that many free GPUs, plus the CPU and memory you asked for, on one Node at the same time.
  • The GPU type is a filter, not a menu. The optional type key must match exactly, including case, one of the GPU types the runner advertises. Omit it to accept any GPU the runner has.

Run a GPU sandbox on serverless capacity

Serverless placement needs no runner and no policy of your own. CoreWeave owns the policy and the hardware. Serverless is currently the only placement mode where GPU sandboxes run in a virtual machine, so it’s the mode to use for untrusted code.

Available GPUs

Serverless capacity runs one GPU model, the NVIDIA RTX PRO 6000 Blackwell Server Edition. Only 1-GPU sandboxes are supported on serverless capacity. To run a sandbox with more than 1 GPU, use your CKS cluster. Two instance types carry it: High Memory and Standard Memory. They differ in host RAM, not in the GPU, and both present the same GPU type to a sandbox, so which one a sandbox lands on isn’t something you select. Leave the type key out of the GPU request so the platform assigns a GPU from CoreWeave-managed compute. The field is still accepted here, but it filters rather than selects: a type that matches no runner fails with CWSANDBOX_RUNNER_UNAVAILABLE. A runner that receives an unsupported type rejects it with CWSANDBOX_PLACEMENT_CONSTRAINT_UNSATISFIED. Use resources to choose CPU and memory alongside the GPU count, as shown in the following example. Disk is requested separately from CPU, memory, and GPU rather than alongside them: ResourceOptions has no disk field. The container’s root filesystem is Node-local ephemeral storage that the sandbox doesn’t reserve a share of, so df inside the sandbox reports the Node’s filesystem rather than a per-sandbox quota. For a dedicated writable path, declare a scratch volume:
A disk-backed volume, the default, draws on the same Node-local storage as the root filesystem. Set medium="memory" for a tmpfs instead: a memory-backed volume must declare a size, and the memory-backed volumes on one container can’t total more than 80% of its memory request. Leave runtime_class unset too. A GPU request selects the GPU virtual machine runtime class on its own, and a runtime class you pin is used exactly as given, so pinning the CPU class alongside a GPU request produces a sandbox that can’t reach the GPU.

Create the sandbox

Set your credential as described in Choose and get credentials, then run the following example. It creates a sandbox with 1 GPU, 2 CPUs, and 8 GiB of memory, then prints the GPU that nvidia-smi reports.
This example passes AuthStrategy.COREWEAVE_API_KEY, which reads CWSANDBOX_API_KEY.
The flat resources dict sets requests and limits to the same values. "gpu": 1 is shorthand for "gpu": {"count": 1}. Serverless supports only 1 GPU per sandbox, so keep the count at 1.The resource_gpu property returns the GPU allocation the platform confirmed, such as {'count': 1}.
Sample output:
GPU sandboxes take longer to start than CPU-only sandboxes because the platform attaches the GPUs to the sandbox’s virtual machine. Allow several minutes if you set a request timeout.

Run a GPU sandbox on your CKS cluster

On CKS placement, your administrators decide which GPUs sandboxes can use and how many. The runner offers the GPU types present on the cluster’s Nodes, and the policy it carries bounds what a sandbox may request.
GPU virtualization isn’t supported on your own CKS cluster. GPU sandboxes on CKS by default run under the cluster’s default runtime class, which uses runc. The sandbox is an ordinary container that shares the Node’s kernel and has direct access to the GPU devices assigned to it. This mode isn’t a secure container runtime: don’t use it to run untrusted code. Use it only for workloads you’d already be comfortable running as a regular Pod on the cluster. For untrusted code, use serverless capacity.

Prerequisites

Complete the following before you request a GPU on CKS:
  • A CKS cluster with GPU Nodes and a runner in the Ready state. See Use your own compute.
  • A policy on that runner that permits GPUs. maxGpuCount under resources caps how many GPUs one sandbox may request: leave it out for no cap, or set it to 0 to reject GPU requests. Leave defaultGpuRuntimeClass unset so GPU sandboxes run under the cluster’s default runtime class. Kata-based GPU runtime classes, such as cw-kata-nvidia-gpu, aren’t supported on customer CKS clusters. For every field, see Configure a sandbox policy.
  • The SANDBOX_USER IAM action and a CoreWeave API access token, set as CWSANDBOX_API_KEY.

Find the GPU types on your runner

Each runner advertises the GPU types its Nodes carry. List them before you choose a type:
The values are the exact strings to pass as type. Matching is case-sensitive, so B200 and b200 aren’t the same type. To filter runners by GPU type instead, pass gpu_type to list_runners(). For more, see Discover available runners.

Create the sandbox

The following example places a sandbox on CKS with 2 GPUs of a specific type, 8 CPUs, and 32 GiB of memory. Replace [GPU-TYPE] with one of the types your runner advertises, or drop the type key to accept any GPU on the runner. Leave runtime_class unset so the sandbox uses the cluster’s default runtime class.
To target one runner rather than any CKS runner in your organization, add runner_ids=["[RUNNER-ID]"]. The confirmed allocation in resource_gpu reports the count only, such as {'count': 2}, even when you requested a type. The GPU count is capped by the policy’s maxGpuCount and by the number of free GPUs on a single Node. A request above the policy cap fails with CWSANDBOX_RESOURCE_CEILING_EXCEEDED. A type the runner doesn’t offer, or one with a different case, finds no eligible runner and fails with CWSANDBOX_RUNNER_UNAVAILABLE.

Container images

The platform provides the NVIDIA driver and the nvidia-smi tool inside a GPU sandbox, so the default image can already see the GPU. To run CUDA applications, use an image that ships the CUDA runtime and libraries your code needs, such as a pytorch/pytorch or nvidia/cuda image. Match the image to the GPU. The RTX PRO 6000 Blackwell Server Edition is compute capability 12.0 (sm_120), which needs CUDA 12.8 or later, and PyTorch 2.7 was the first release built for it. An older image still reports the GPU’s name correctly, because that reads device metadata through the driver, then fails at the first kernel launch with CUDA error: no kernel image is available for execution on the device. Check that your framework lists sm_120 rather than trusting the device name:
Sample output:
A large framework image takes longer to pull than the default image, so allow a few minutes for the sandbox to become ready.

Common errors

Next steps

Last modified on October 6, 2026