View pricing for instance types on CoreWeave’s pricing page.
Availability
Node Pools are available in all General Access Regions.Sizing your Node Pools
Because you can run multiple Node Pools in a single cluster, you decide how to divide your Nodes among them. There is no single correct layout. The right structure depends on how many distinct workloads you run, how you want to scale them, and how much management overhead you can absorb. Use the trade-offs and recommendations in this section to plan a layout, then create and autoscale your Node Pools to match. A Node Pool maps to Nodes through its target. For most instance types,spec.targetNodes sets the number of Nodes CKS provisions and keeps running. For rack-based instance types such as GB200 and GB300, spec.targetRacks sets the number of racks, and each rack contains 18 Nodes. Every Node in a Node Pool shares that pool’s instance type, labels, taints, and annotations, so a Node Pool is the unit where you set Node configuration and capacity together.
Recommendations by workload type
Use these recommendations as a starting point, then adjust based on how your workloads schedule and scale.-
A single large training job: Use one Node Pool sized to the job. Training jobs run across many Nodes of the same instance type, so a single pool matches the workload and keeps management simple. Set the scaling strategy to
IdleOnlyso that loweringtargetNodesdoesn’t remove Nodes that are still running the job. The Cluster Autoscaler ignores the scaling strategy, so if the pool autoscales, also prevent Node removal with an annotation. Consider Node Pool prefill to avoid a capacity gap if a Node is replaced mid-run. -
Multiple independent inference endpoints: Use a separate Node Pool per endpoint, or per group of endpoints with the same instance type and scaling needs. Separate pools let you scale each endpoint independently in response to its own demand and isolate one endpoint’s Nodes from another’s. The
PreferIdlestrategy scales these pools down faster when demand drops. See Autoscale Node Pools for how the autoscaler selects a pool based on a Pod’snodeSelectorandaffinity. - Mixed instance types: Use one Node Pool per instance type. Because every Node in a pool shares one instance type, you need a distinct pool for each GPU or CPU type your workloads require.
- Bursty or intermittent workloads: Put these workloads in their own Node Pool and configure scale-to-zero so the pool drops to zero Nodes when there’s no demand.
Node Pool sizing considerations
The following table summarizes the trade-offs between consolidating Nodes into fewer, larger Node Pools and splitting them across more, smaller Node Pools.
In short, fewer large Node Pools reduce management overhead at the cost of increased impact scope and less flexibility, while more smaller Node Pools give you workload isolation and finer scaling control at the cost of more pools to manage.
Node cordoning
CKS sometimes cordons Nodes to ensure that workloads are only scheduled to healthy Nodes. In most cases, CKS eventually removes the cordon, which makes the Node schedulable again. CoreWeave manages this kind of Node cordoning entirely, though you can also manually cordon Nodes for your own reasons.CoreWeave-managed Node conditions, labels, states, and events can change without notice. Don’t use them for your own automation or change them except through documented procedures, such as cordoning a Node yourself or requesting a Node reboot. You can still add your own custom Node conditions and labels.
- Maintenance: If a Node requires maintenance, updates, or hardware fixes, CKS cordons it to ensure no new workloads are placed on it during that time. This lets CoreWeave’s Node lifecycle automation make necessary changes without disrupting running tasks.
- Node draining for removal: If a Node must be removed from the cluster, CKS cordons the Node before draining it. CKS automatically reschedules workloads onto healthy Nodes, and no new workloads are scheduled to the cordoned Node.
- InfiniBand or Ethernet link flaps: Link flaps are intermittent, unpredictable up-down transitions in a network connection, which can result in networking or communication failures. If an InfiniBand or Ethernet link is flapping, a Node can experience inconsistent or unreliable connectivity. In this case, CKS cordons the Node to ensure no workloads are scheduled to a Node with an unreliable network connection.
- Temporary health check failures: Kubernetes uses health checks to assess the state of a user’s system. A temporary check failure might indicate transient issues that could degrade Node performance. CKS cordons the Node until the issue is resolved.
Node Pool prefill
Node Pool prefill helps maintain capacity while individual Nodes are replaced. For Nodes selected for prefill, CKS provisions replacement capacity before draining and removing the existing Node. Replacement capacity depends on availability. You aren’t billed for both Nodes during the prefill overlap. Prefill is enabled by default for supported new Node Pools.Node Pools created before October 8, 2026, retain their existing prefill settings. See the prefill changelog entry.
spec.prefill field, CKS uses prefill default values.
When to use prefill
Use prefill for Node Pools whose workloads are sensitive to capacity loss during Node replacement, such as long-running training jobs or latency-sensitive inference.- Without prefill enabled: CKS replaces a Node only after the old Node is removed, so the Node Pool is below its target Node count until the replacement is delivered and ready. This leaves a capacity gap that can last 20 to 40 minutes.
- With prefill enabled: CKS provisions a replacement Node before draining and removing the old Node, reducing the risk of a capacity gap. Workloads can continue on the existing Node while CKS provisions its replacement.
Prefill availability
Prefill provisions replacement Nodes from on-demand capacity, so a replacement Node is not guaranteed. If on-demand capacity is unavailable, prefill keeps retrying, and the Node stays in the prefill flow until a replacement can be provisioned. Account for this when you rely on prefill for time-sensitive replacements. Prefill requirements
If prefill is enabled on an unsupported Node Pool, CKS sets
PrefillEnabled to False and doesn’t mark new Nodes for prefill. This also applies when prefill is enabled by default. See Prefill status.
How prefill works
A Node is marked unschedulable as soon as it enters triage, so no new workloads are scheduled onto it. If a replacement Node can’t be provisioned, CKS keeps retrying until one is available or the Node is no longer marked for prefill. After the replacement joins the Node Pool, CKS waits up to the idle timeout (spec.prefill.timeout) for the old Node to become idle, then removes it. If the Node hasn’t become idle when the timeout expires, CKS starts draining it. CKS removes the Node once it becomes idle.
Prefill Node conditions
When prefill is enabled, CKS sets aPrefill condition on Nodes that are in the prefill flow. The condition’s reason field indicates the current state. You can use LastTransitionTime on the condition to see when the reason last changed.
Prefill flow
The following diagram shows how a Node moves through the prefill flow, including the two paths CKS can take depending on whether the Node becomes idle within the default 24 hour idle timeout (spec.prefill.timeout) period.
For details on configuring or disabling prefill, see Configure Node Pool prefill. For the prefill reference specification, see Prefill reference.
Scaling strategies
A Node Pool has two scale-down strategies, which control how CKS chooses Nodes to remove when you lowertargetNodes. The Cluster Autoscaler doesn’t use them. To protect active Nodes from the autoscaler, see Prevent Pod eviction or Node removal. Set spec.lifecycle.scaleDownStrategy to one of the following values:
Idle Nodes
A Node is idle when it has no Pods in theRunning or Pending phase. When determining a Node’s idle status, CKS ignores Pods with any of the following conditions:
You can determine a Node’s idle status by inspecting the
CWActive condition. CWActive = False means the Node is idle.
Node idle status and maintenance
CKS defers certain maintenance operations, such as moving a Node out of production after detecting elevated InfiniBand link flaps, until the Node becomes idle. A long-running Pod without an interruptible label keeps the Node active indefinitely, preventing those deferred operations from completing. For Pods that run continuously but can tolerate interruption, apply theqos.coreweave.cloud/interruptable label so that deferred maintenance doesn’t wait for them. For stateful Pods that need cleanup time, use qos.coreweave.com/graceful-interruptible, which makes deferred maintenance wait only for the Pod’s termination grace period. Gracefully interruptible Pods still count toward CWActive until they finish terminating, as the preceding table shows. See Pod interruption and eviction policies. To learn how deferred maintenance operations behave, see Node state transitions in CKS.
Once CKS selects a Node for removal, CKS removes the Node by first cordoning it to prevent any new Pods from being scheduled onto it. Then, CKS drains the Node to perform a graceful cleanup. Once the Node is fully drained, CKS removes it from the Node Pool.
DeletionGracePeriodSeconds defaults to 30 seconds unless the Pod’s spec.terminationGracePeriodSeconds is set to a different value.Scale down a Node Pool
To scale down a Node Pool:- Set
spec.targetNodesto the desired number of Nodes, orspec.targetRacks, depending on the target type set on the Node Pool. For instructions on doing this with Cloud Console or the Kubernetes CLI, see Manage Node Pools. - Set
spec.lifecycle.scaleDownStrategyto the preferred scaling strategy.
How CKS selects Nodes to remove
When a Node Pool scales down, CKS picks which Nodes to remove based on the configured scale-down strategy and Node idleness. There is no Node Pool field or Node label that lets you single out a specific Node for removal during scale-down. CKS chooses among eligible Nodes automatically. If you need a particular Node to be the one removed during the next scale-down, do one of the following before loweringtargetNodes:
- Drain the Node yourself. Use standard Kubernetes tooling (
kubectl cordonandkubectl drain) to evict workloads from the target Node. With the defaultIdleOnlystrategy, an empty Node becomes eligible for removal first. - Reschedule workloads off the Node. Delete or move the Pods running on the target Node so the Node becomes idle. CKS ignores Pods that are part of a DaemonSet, control-plane Pods, and Pods labeled as interruptible when determining idleness (see Idle Nodes).
No nodes eligible for removal right now, see Current is above target for how to find the Pods blocking it.
Node Pool types
CKS uses thedefault Node Pool type to manage Reserved and On-Demand instances. The following section describes when to use the default type and how to configure it.
Default Node Pools
CKS uses the default Node Pool type (spec.computeClass: default) for Reserved and On-Demand instances. Billing for default Node Pools depends on Reservation and utilization.
You don’t need to specify default Node Pools.If you don’t specify a
computeClass, the Node Pool defaults to the default type.Example default Node Pool manifest
Manage Node Pools with Terraform
The CoreWeave Terraform provider manages clusters, VPCs, and Object Storage, but it doesn’t include a Node Pool resource. Deploy Node Pools through the Cloud Console or the Kubernetes API instead. A Terraform-only workflow typically combines the CoreWeave provider for the cluster and VPC with a separate mechanism, such as the Kubernetes provider, to apply Node Pool manifests. See How do I use Terraform to manage clusters and Node Pools?.Image pull best practices
CoreWeave operates a region-level registry proxy that accelerates container image pulls and reduces exposure to public registry rate limits for your cluster’s Nodes. To ensure predictable and fast rollouts when scaling Node Pools, use immutable tags, or pin by digest, for production workloads. This ensures every Node pulls the same artifact and avoids stale results from proxy metadata caching. Avoid mutable tags like:latest. With metadata caching enabled, the proxy can continue serving a cached manifest until the cache expires, which can lead to inconsistent versions across Nodes. For more information, see the region-level image proxy documentation.
Reboot methods
You can manually reboot Nodes in the following two ways:Node conditions for reboots (deprecated)
CoreWeave’s Node conditions are visible when you manage reboots. You previously set these Node conditions with the Conditioner Kubectl plugin. The CoreWeave Intelligent CLI doesn’t set them. It requests reboots through thePhaseState and PendingPhaseState lifecycle conditions instead. See the Reboot Nodes and Apply Node Pool updates guides for more information.
Control Plane Node Pool (deprecated)
CoreWeave provisioned clusters created before July 7, 2025 with acpu-control-plane Node Pool for CKS-managed components. Clusters created after this date don’t have this Node Pool. The CKS Control Plane now manages the components out-of-band. See the Control Plane Node Pool release notes for more information.