all_reduce test from CoreWeave’s nccl-tests repository on a small, known-good group of Nodes and record the bus bandwidth (busbw) it reports. That’s your healthy baseline for the same message sizes: a job running correctly over InfiniBand lands close to it, while a job that has fallen back to TCP usually reports busbw an order of magnitude lower. If Straggler Detection is enabled on your SUNK cluster (the default on v8.0.0 and later) and the NCCL your job loads meets its prerequisites, the Slurm Job Metrics dashboard reports bus bandwidth (BusBW) per collective continuously from live jobs. You can then compare against the baseline without stopping the job to run nccl-tests.
If the dashboard panels are empty, rule out the expected cases first. Jobs shorter than about 2 minutes and process-group sizes your job doesn’t use leave panels empty by design (see NCCL Metrics).
If the Straggler Detection plugin is enabled and the NCCL BusBW panels stay empty for a longer job, run it with NCCL_DEBUG=INFO. Check the NCCL version the job reports and look for the Successfully loaded external profiler plugin line. An NCCL version older than the prerequisites require can come from the image itself or from a Python virtual environment. For example, pip install torch bundles its own NCCL, which takes precedence over the image’s. If the version meets the prerequisites but the line is missing, the plugin didn’t load. See Enable the plugin.
When InfiniBand is the cause, the slowness is usually one of these: InfiniBand is not actually being used, the InfiniBand interfaces are down, or NCCL is missing the environment variables that point it at the IB HCA.
Quick diagnostic. From inside your training Pod, check InfiniBand port status. If a port intended for the job’s InfiniBand traffic shows Down or Disabled, contact support with the Node names and ibstat output.
Common causes:
- Pod spec is missing
rdma/ib: 1. Check the RDMA resource request against Configure the Pods. - NCCL environment variables are missing or incorrect. Use the NCCL configuration reference for your fabric’s settings. If your stack needs UCX on the RDMA fabric, follow the reference’s UCX exception.
- Node Pool does not have InfiniBand. See InfiniBand and RoCE labels for the labels to look for.
NCCL_DEBUG=INFO and look for NET/IB (good) compared to NET/Socket (TCP fallback).
NVLink domain placement on rack-scale systems
On rack-scale, NVLink-connected instances such as NVL72, GPUs communicate over NVLink within an NVLink domain and over the scaleout fabric (InfiniBand or RoCE) beyond it. NCCL is fastest when collectives stay inside the NVLink domain, so a job whose Pods are spread across domains can route traffic that should run over NVLink onto the slower scaleout fabric instead. When this happens, throughput drops even though InfiniBand is healthy and every cluster-wide check passes. Align scheduling with physical connectivity so related Pods land in the same NVLink domain. Use thenvidia.com/gpu.clique label as a Pod affinity topologyKey, and see IMEX overview for how NVLink domains, partitions, and placement work on CoreWeave.
Isolating a bad Node on large jobs
If you are training across hundreds or thousands of GPUs and the cluster-wide checks above pass, the cause may be a small group of Nodes with degraded HCAs rather than a configuration issue. Split your nodelist in half, run thenccl-tests jobs on each half, and keep bisecting the half that performs worse until you identify the rack or Nodes contributing to the slowdown. Then contact support with the Node names.
Administrator