- Silent data corruption: Detects memory errors that do not trigger standard system alerts.
- Numerical performance regression: Benchmarks common deep-learning kernels (FP8, FP16, BF16) to detect slowdowns.
- Thermal deficiencies: Runs all streaming multiprocessors at 100% load for approximately 20 minutes to uncover thermal issues.
.spec.nodeName, and work started outside Slurm on SUNK clusters are exceptions. See Test schedule, Why is my Pod stuck in pending on CKS?, and Preemption on SUNK clusters. A transient hpc-verification Pod in your cluster is a test that ran on an idle Node or was preempted to make way for your job.
HPC Verification is separate from Node Health Check (NHC).
SUNK runs Node Health Check (NHC) as one of the scripts in the Epilog chain after a job, and only when no other job is still running on the node. A failed check drains the node with a reason that begins with NHC:.
A passing HPC Verification test is what returns a node whose drain reason carries sunk:verify-undrain to service. See Node Health Check (NHC) drains.
Test schedule
At the top of each hour, if a Node is idle, meaning no customer Pod holds any of its GPUs, CoreWeave runs a 20- to 30-minute verification test. This rule is narrower than the idle status CKS uses for maintenance, which also counts Pods that don’t hold GPUs. This test uses all GPUs and available InfiniBand resources to exercise the hardware uniformly without interrupting your jobs. Thehpc-verification Pods run at a low Kubernetes priority class to allow preemption and avoid blocking customer workloads. If the Node is not idle, CoreWeave skips the test and retries at the next hour mark.
Third-party schedulers don’t preempt verification Pods by default, because those Pods are placed by the default-scheduler. Respecting Kubernetes PriorityClasses isn’t sufficient on its own. If you run a third-party scheduler such as Volcano or the KAI Scheduler, see Custom schedulers and third-party integrations, and contact CoreWeave before you run it on a production cluster. CoreWeave’s cks-kueue chart integrates with HPC Verification natively, so it doesn’t need this coordination. Contact CoreWeave Support or email support@coreweave.com if you have questions about HPC Verification and how it interacts with your scheduler.
Taints
UseNoSchedule taints to keep Pods without matching tolerations off GPU Nodes while allowing automated HPC Verification to run. Verification Pods tolerate arbitrary NoSchedule and PreferNoSchedule taints, but they don’t tolerate arbitrary NoExecute taints.
A NoExecute taint can prevent a verification Pod from scheduling or evict it after it starts. This affects verification in two ways:
- Idle Nodes: Automated verification can’t run, so the Node remains unverified even when no customer workloads are running.
- Rebooted Nodes: The Node returns to
productionwithout waiting for verification, so a blocked test leaves it unverified rather than holding it back. A Node can still remain inproduction-powerresetorproduction-reboot. CoreWeave might delete and replace a Node that remains in this state too long.
If you apply a
NoExecute taint, change its effect to NoSchedule in the Node or Node Pool configuration that manages the taint. Check that your workload tolerations still match the intended taint effect. Leave CoreWeave-managed taints unchanged. See User taints.
A rebooted Node returns to production as soon as the reboot completes, and the verification test runs after that. Let the test finish before scheduling workloads onto the Node: the test needs the Node’s GPUs free, so a workload that lands first prevents it from starting, and one that lands mid-test stops it. For a Node that remains in a reboot state, see Reboot troubleshooting.
Verify a Node on demand
Use the CoreWeave Intelligent CLI to request verification of a specific GPU Node without waiting for the hourly test. The Node must be idle, with no customer Pod holding any of its GPUs. Before you begin, install and authenticate the CoreWeave Intelligent CLI and configure access to the Node’s cluster. CheckNoExecute taints and follow the guidance in Taints.
Replace [NODE-NAME] with the name of the Node to verify:
production lifecycle state before resuming workloads. See Monitor the reboot.
If verification doesn’t complete or the Node remains unavailable, contact CoreWeave Support. Include the cluster name, Node name, taints, and approximate time of the verification request.
Customer impact
HPC Verification jobs are not billable and don’t appear in usage reports. They run only when Nodes are idle, so a workload that Kubernetes or Slurm schedules never shares GPU resources with a test. If you launch a workload that needs the GPUs a test holds, the test stops, and CoreWeave releases resources before your Pod starts running. Your job and the test never run at the same time. Three cases work differently. A workload placed by a third-party scheduler doesn’t stop the test, so the workload waits inPending until the test ends. A Pod pinned to a Node with .spec.nodeName bypasses the scheduler, so it can’t preempt the test and can fail to start on a Node that looks free. See Why is my Pod stuck in pending on CKS?. On SUNK clusters, work that you start outside Slurm, such as a process launched over SSH, can run on the same GPUs as a test. See Preemption on SUNK clusters.
Preemption on SUNK clusters
On a SUNK cluster, Slurm stops a verification test when it allocates a node for a job. When Slurm allocates a node for a job submitted withsrun, salloc, or sbatch, SUNK applies the sunk.coreweave.com/lock taint during the job’s Prolog and evicts every Pod that doesn’t tolerate it, including verification Pods. See The SUNK /lock taint.
Work that you start on a compute node outside Slurm, such as a process launched from an SSH session on a node where you hold no job, doesn’t run the Prolog. The Node still looks idle to Kubernetes and Slurm, so a verification test can start on it while your process runs. This has two effects:
- If you run
nvidia-smion a compute node that you reached over SSH without an allocation, you might see GPU utilization and GPU memory usage from a verification test, even though no Slurm job is running. - Your process and the test compete for the same GPUs, which can cause the Node to fail verification.
srun, salloc, or sbatch, including interactive sessions. SSH to a node where you already hold an allocation doesn’t cause this conflict, because that job’s Prolog has already applied the lock. SUNK removes the lock once no Slurm job is running on the node, so stop any process that you started over SSH before your job ends. See Access Slurm compute nodes.
Identify the test
Use the following signals to confirm that anhpc-verification Pod or metric spike on your Node comes from HPC Verification rather than a customer workload.
You might briefly see a Pod whose name starts with hpc-verification- (single-Node tests) or hpcv-mpi- (multi-Node group tests) in kubectl get pods -A while the test runs. The test Pods run in the cw-hpc-verification namespace and use a CoreWeave PriorityClass with a negative value: cw-hpc-verification at -1, or cw-hpc-validation at -50 on clusters where CoreWeave has rolled out the new class. Either way, they run at a lower priority than customer workloads that use the default priority.
While the test is running, you might notice short spikes in your Grafana dashboards for:
- GPU SM Utilization: Measures how heavily the GPU cores are used.
- CPU Utilization: Shows CPU usage during the test.
- GPU Memory Utilization: Tracks GPU memory usage.
Key points
- HPC Verification tests can’t be disabled. They are a core safeguard that supports the CoreWeave SLA.
- You might briefly see an
hpc-verification-*Pod inkubectl get pods -A. This means the test is running on an idle Node. Kubernetes preempts it as soon as your workload starts, unless a third-party scheduler places your workload or the Pod is pinned to the Node with.spec.nodeName. On SUNK clusters, work started outside Slurm doesn’t preempt it. See Preemption on SUNK clusters. - These tests aren’t billed and don’t appear in usage reports.
support@coreweave.com. Include the cluster name, Node name (from kubectl describe node), and the approximate time. The CoreWeave team checks and, if needed, uncordons or replaces the Node to restore capacity.