What you see
NCCL prints errors to your training log that look like one of these exact strings:NCCL error ... internal error(NCCL Error 3)NCCL error ... unhandled system error(NCCL Error 2)NCCL error ... unhandled cuda error(NCCL Error 1)NCCL error ... remote process exited or there was a network error(NCCL Error 6)[Rank N] Watchdog caught collective operation timeoutorWorkNCCL ... timed outtransport/net_ib.cc ... vendor err 249orvendor err 129ibv_reg_mr_iova2 failed: Invalid argument
NCCL error codes
NCCL Error 3 (internal error)
Despite the name, NCCL Error 3 is almost never an NCCL bug. The “internal error” wording is misleading: it covers a range of transport and GPU failures. Work through the following order:- Job layout that breaks rank-to-device mapping. If the error is deterministic and reproducible rather than intermittent, check the job’s GPU request before anything else. On a cluster with
ConstrainDevicesenabled incgroups.conf,--gpus-per-task=1with--ntasks-per-node=8gives every task a single GPU renumbered as device 0, which breaks NCCL’s expectation that rank N on a node maps to GPU N. The same job works with--gpus-per-node=8. Drop--gpus-per-task=1and keep--ntasks-per-node=8so each task sees all eight GPUs with their real device ordinals. The same misconfiguration can surface instead asunhandled cuda errorwithCuda failure 101 'invalid device ordinal'. - Connectivity between nodes. NCCL reports Error 3 when it can’t reach a peer over the transport it selected. Run an NCCL test on the same set of nodes (how to run one). If the test always fails, the path is broken; if it fails intermittently, re-run with
NCCL_DEBUG=INFOand keep the logs; if it never fails, the problem is in the workload. - GPU hardware fault. An underlying XID error can surface as Error 3. Check the node’s kernel log for an XID at the timestamp of the NCCL failure.
- Application memory corruption. A bad pointer or buffer overflow in CUDA code can corrupt NCCL’s internal state. Suspect this when XID 31 (MMU fault) appears on many nodes in a single job: XID 31 across many nodes in one job is application-side, not hardware.
- Application crash on another rank. A different rank may have crashed and left this rank unable to complete the collective. Find the first rank that died. See Training job exit-code triage.
NCCL Error 2 (unhandled system error)
NCCL Error 2 means a system call that NCCL depends on failed. It points at the network, the OS, or the driver rather than at NCCL itself. Check, in this order:- A network device that went down mid-run. On InfiniBand clusters, from a Pod on an affected node, check InfiniBand port status. For RoCE clusters, check backend link health.
- A stalled storage mount. An NFS or VAST mount that stops responding can fail the system calls underneath NCCL.
- A peer rank that exited and dropped its connections.
Call to ibv_reg_mr_iova2 failed with error Invalid argument, the cause is a container library mismatch, not the fabric. See ibv_reg_mr_iova2 failed with Invalid argument.
NCCL Error 1 (unhandled cuda error)
unhandled cuda error is NCCL Error 1: a CUDA call failed underneath NCCL. This is a GPU-side error, so start with the GPU rather than with storage or the fabric. Check the node’s kernel log for a GPU XID event at the time of the failure, then check the driver version and GPU health on that node.
If the error carries Cuda failure 101 'invalid device ordinal', it’s the job-layout problem described in NCCL Error 3 step 1, not a GPU fault.
NCCL Error 6 (remote process exited)
NCCL Error 6 means either a peer rank exited while a collective was in flight, or the network path to that peer failed. The full message,remote process exited or there was a network error, covers both. The rank that prints Error 6 is almost always reporting someone else’s failure, so start by finding the rank that died first:
- Sort all rank logs by timestamp.
- Find the earliest rank that printed a real failure (out of memory, segfault, XID, exit code). That rank is the cause.
- Diagnose that rank’s failure, not the Error 6 messages downstream of it.
WorkNCCL timeout diagnostic order
AWatchdog caught collective operation timeout or WorkNCCL ... timed out means a collective didn’t complete within the timeout. This is one of the most common and most misdiagnosed NCCL errors. Work through this exact order before suspecting the fabric:
- Dataloader or storage starvation. Check whether a rank was waiting on data. Look for GPUs at 0 percent utilization at the time of the timeout and a drop in storage read throughput. A rank starved of data can’t reach the collective, so every other rank times out waiting for it. This is the single most common cause.
- Silent rank crash before the timeout. Check whether a rank crashed (out of memory, segfault) just before the timeout fired. The crash is the cause. The timeout is the symptom on the surviving ranks.
- NCCL test on the same set of nodes. Run an NCCL test on the same nodes the job used (how to run one). If it passes, the fabric is healthy and the cause is in the workload or storage.
- Only then suspect the fabric. If the NCCL test fails on the same set of nodes, isolate the bad node by bisecting the nodelist, then contact support with the node names.
vendor err 249 compared to vendor err 129
These appear in InfiniBand transport errors such astransport/net_ib.cc ... vendor err NNN. They imply different fabric states:
vendor err 249indicates an actively flapping link. The connection goes up and down during the run. This points at a physical or transient link problem on a specific node.vendor err 129indicates transport retries were exhausted. CheckNCCL_IB_RETRY_CNTbefore the fabric: the value is stored in a 3-bit hardware register, so 7 is the maximum, and anything higher is silently truncated. A common setting of 16 truncates to zero retries, which makes the job die on the first transport error and makes a healthy fabric look broken. See NCCL_IB_RETRY_CNT has a hard maximum of 7. If the value is 7 or lower and the error persists, isolate the node.
rdma-core or UCX mismatch. If the onset points at the fabric, isolate the affected node by bisecting the nodelist, then contact support with the node names and the exact error string.
ibv_reg_mr_iova2 failed with Invalid argument
This error means the libibverbs or OFED userspace inside your container doesn’t match the kernel RDMA modules on the node.ibv_reg_mr_iova2 registers memory regions for RDMA; a newer userspace passes registration parameters that an older kernel module rejects with EINVAL. It reads like a fabric problem because the message mentions IB, but it’s a local software version mismatch present on every node. Don’t investigate switch health for this error.
Two signals make it close to definitive:
- Single-node jobs succeed, multi-node jobs fail. Single-node traffic goes over NVLink, which doesn’t use libibverbs. Only the IB path calls
ibv_reg_mr_iova2. - CoreWeave’s
nccl-testsimage passes on the same nodes while your image fails. The CoreWeave image is built against an OFED version known to match the node, so this narrows the problem to your container.
- Upgrade the node image (preferred when the nodes are old). If the nodes are running an old ncore or driver, upgrading is the right call regardless of this error. Set the driver version on the Node Pool, as described in Update the GPU driver, then reboot the GPU nodes. Expect roughly an hour per node for the reboot and health-check cycle. Ask your CoreWeave representative to coordinate if you can’t schedule the reboots yourself.
- Match the container to the node (faster, no reboot). If you can’t take the downtime, use a CUDA base image whose OFED version matches the node’s, install a matching libibverbs in the container, or start from CoreWeave’s
nccl-testsimage.
Stuck in NCCL communicator init after a previous crash
If the job stops responding during NCCL communicator initialization specifically after a previous run crashed, the cause is usually a corrupted on-disk compile cache, not NCCL or the fabric. A previous run that died without cleaning up can leave both the PyTorch Inductor and Triton caches in a bad state. The strong signal is adestroy_process_group ... was not called warning from the previous run, which means the prior process exited without tearing down cleanly.
The two caches default to different storage, which determines how far the corruption spreads:
- Inductor defaults to a
torchinductor_[USERNAME]directory under the system temporary directory, normally/tmp/torchinductor_$USER. That’s node-local, so corruption stays on the nodes the crashed run used. - Triton defaults to
$HOME/.triton/cache. Because home directories are on shared storage, corruption there follows your jobs to any node.
TORCHINDUCTOR_CACHE_DIR to a path on shared storage is the obvious one, and it makes the Inductor cache behave like the Triton case. The subtler one is TMPDIR: Inductor resolves its default through Python’s temporary directory, which honors TMPDIR, so wherever TMPDIR points is where the cache lands. Check TMPDIR before you assume the cache is on the node.
Under Pyxis or Enroot, /tmp inside the container is tmpfs (RAM) rather than the node’s NVMe, so an Inductor cache left at the default doesn’t survive the container at all and can’t be the stale cache you’re chasing. It also consumes memory while the job runs. Jobs that set TMPDIR to an NVMe-backed path, as Node-local storage and /tmp on Slurm nodes recommends, do keep a cache across runs on that node. If your job runs in a container and doesn’t set TMPDIR, suspect the Triton cache in your home directory instead.
Fix it by clearing the caches or disabling them for the next run:
Checkpoint segfault at scale, followed by a collective timeout
A segfault during checkpoint saving followed by an NCCL collective timeout is a driver-level problem, not a fabric fault. The signature on GB200 nodes running driver 580 is specific:- The segfault lands in
crc32_16bytes, the checksum routine that verifies checkpoint data. - The segfault appears in the log before the NCCL timeout. The timeout is the surviving ranks waiting on the ranks the segfault killed.
- It’s scale-dependent, more likely at eight racks or more, and it can follow the workload to another cloud provider after the same driver upgrade.
- Call
torch.cuda.synchronize()before checkpointing, so all in-flight GPU work completes before the save begins. - Move checkpoint writes to object storage, which changes the data path.
- Roll the Node Pool back to driver 570. See Update the GPU driver.
Disambiguate a training stall from an NCCL stall
When a job stops responding and you can’t tell whether NCCL is stuck or the stall is in your code upstream of NCCL, turn on both debug channels:all_to_all and mixture-of-experts sensitivity
A passingall_reduce doesn’t certify the fabric for an all_to_all workload, and the reason is the traffic pattern. Mixture-of-experts models use all_to_all_single for expert parallelism, which needs point-to-point communication between every pair of participating nodes. all_reduce uses a ring or tree topology, so each rank only talks to its neighbors. One unreachable node breaks an all_to_all outright while an all_reduce routes around it.
So if your workload uses all_to_all and fails, test with all_to_all. A clean all_reduce benchmark on the same set of nodes can hide a connectivity problem that only the workload’s actual collective triggers.
Run an NCCL test on the same set of nodes
Several sections above tell you to run an NCCL test on the nodes the job used. This is how you get a comparable result:- Get the test jobs. CoreWeave publishes sample NCCL test jobs for Slurm and for the MPI Operator. See the
coreweave/nccl-testsrepository for the jobs and instructions. CoreWeave’sslurmdimages are built on these images, so the binaries are already present in a standard SUNK container. See Slurm images. - Match the collective to your workload. Run
all_reducefor a data-parallel job andall_to_allfor a mixture-of-experts job, for the reason in the preceding section. - Use the same nodes. Pass the job’s nodelist so the test exercises the same paths. A test on different nodes proves nothing about the ones that failed.
- Compare against a baseline, not a feeling. Record the bus bandwidth (
busbw) the test reports on a small known-good group of nodes first, then compare. For how to baseline and how to bisect a nodelist down to the bad node, see Why is my multi-node NCCL training slow?.
nccl-tests pass on the same nodes, the gap is in the harness, not the fabric. Check message sizes, NCCL environment variables, and whether the timing loop calls torch.cuda.synchronize().
When to open a support ticket
Self-diagnose first using the preceding decision trees. Open a ticket in any of these situations:- An NCCL test fails on a specific set of nodes and you have isolated the node by bisection.
- You see
vendor err 249(actively flapping link) arriving on several nodes at once. - On an InfiniBand cluster,
ibstatshows a port intended for the job’s InfiniBand traffic asDownorDisabled, and the port doesn’t recover.
ibstat output.
What this is not
The following points correct the most common misdiagnoses:- NCCL Error 3 isn’t an NCCL bug in almost all cases, and it isn’t the storage-first case either. Check the job’s GPU request and node-to-node connectivity first.
- A WorkNCCL timeout isn’t a fabric problem by default. It’s most often a starved dataloader. This is the error the storage-first order applies to, not Error 3.
- A passing
all_reducetest doesn’t certify the fabric for anall_to_allworkload. unhandled cuda errorisn’t a system or fabric error. It’s NCCL Error 1, and it points at the GPU or the driver.ibv_reg_mr_iova2 failedisn’t an IB fabric fault. It’s a container library mismatch on every node.
Related pages
- Training job exit-code triage
- torch.distributed rendezvous and TCPStore troubleshooting: failures during startup, before training begins.
- Measuring MFU and job performance: data loading and step-time diagnosis.
- Why is my multi-node NCCL training slow?: baselining with nccl-tests and bisecting a nodelist.
- NCCL configuration reference: the NCCL environment variables CoreWeave documents.