> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# What do NCCL errors mean and how do I diagnose them?

Most NCCL errors aren't NCCL bugs. The message you see usually names the layer where NCCL gave up, not the layer that failed, so the wording is often different from the real cause. Match the exact error string against the following sections. Which cause to check first depends on the error: a collective **timeout** is most often storage or a starved dataloader, while an **internal error** is most often connectivity, a GPU fault, or a job-layout mistake.

## What you see

NCCL prints errors to your training log that look like one of these exact strings:

* `NCCL error ... internal error` (NCCL Error 3)
* `NCCL error ... unhandled system error` (NCCL Error 2)
* `NCCL error ... unhandled cuda error` (NCCL Error 1)
* `NCCL error ... remote process exited or there was a network error` (NCCL Error 6)
* `[Rank N] Watchdog caught collective operation timeout` or `WorkNCCL ... timed out`
* `transport/net_ib.cc ... vendor err 249` or `vendor err 129`
* `ibv_reg_mr_iova2 failed: Invalid argument`

Match your error against the following sections.

## NCCL error codes

| Code | Name | Message | What it usually means |
| - | - | - | - |
| 1 | `ncclUnhandledCudaError` | unhandled cuda error | A CUDA call failed underneath NCCL. Check GPU XIDs, driver state, and GPU health. |
| 2 | `ncclSystemError` | unhandled system error | A system call failed: network, OS, or driver. Check the IB fabric and the storage mounts. |
| 3 | `ncclInternalError` | internal error | Misleading wording. Almost never an NCCL bug. Usually a transport failure, a GPU fault, application memory corruption, or a job layout that breaks NCCL's rank-to-device mapping. |
| 4 | `ncclInvalidArgument` | invalid argument | A bad parameter reached NCCL. Usually an application bug. |
| 5 | `ncclInvalidUsage` | invalid usage | NCCL API misuse. An application bug. |
| 6 | `ncclRemoteError` | remote process exited or there was a network error | A peer rank crashed, or the network path to it failed. |

## NCCL Error 3 (internal error)

Despite the name, NCCL Error 3 is almost never an NCCL bug. The "internal error" wording is misleading: it covers a range of transport and GPU failures. Work through the following order:

1. **Job layout that breaks rank-to-device mapping.** If the error is deterministic and reproducible rather than intermittent, check the job's GPU request before anything else. On a cluster with `ConstrainDevices` enabled in `cgroups.conf`, `--gpus-per-task=1` with `--ntasks-per-node=8` gives every task a single GPU renumbered as device 0, which breaks NCCL's expectation that rank N on a node maps to GPU N. The same job works with `--gpus-per-node=8`. Drop `--gpus-per-task=1` and keep `--ntasks-per-node=8` so each task sees all eight GPUs with their real device ordinals. The same misconfiguration can surface instead as `unhandled cuda error` with `Cuda failure 101 'invalid device ordinal'`.
2. **Connectivity between nodes.** NCCL reports Error 3 when it can't reach a peer over the transport it selected. Run an NCCL test on the same set of nodes ([how to run one](#run-an-nccl-test-on-the-same-set-of-nodes)). If the test always fails, the path is broken; if it fails intermittently, re-run with `NCCL_DEBUG=INFO` and keep the logs; if it never fails, the problem is in the workload.
3. **GPU hardware fault.** An underlying XID error can surface as Error 3. Check the node's kernel log for an XID at the timestamp of the NCCL failure.
4. **Application memory corruption.** A bad pointer or buffer overflow in CUDA code can corrupt NCCL's internal state. Suspect this when XID 31 (MMU fault) appears on many nodes in a single job: XID 31 across many nodes in one job is application-side, not hardware.
5. **Application crash on another rank.** A different rank may have crashed and left this rank unable to complete the collective. Find the first rank that died. See [Training job exit-code triage](/support/sunk/articles/what-do-training-job-exit-codes-mean).

If the job stops responding during NCCL communicator init right after a previous crash, the cause is usually not any of the above. See [Stuck in NCCL communicator init after a previous crash](#stuck-in-nccl-communicator-init-after-a-previous-crash).

## NCCL Error 2 (unhandled system error)

NCCL Error 2 means a system call that NCCL depends on failed. It points at the network, the OS, or the driver rather than at NCCL itself. Check, in this order:

1. A network device that went down mid-run. On InfiniBand clusters, from a Pod on an affected node, [check InfiniBand port status](/products/networking/hpc-interconnect/use-gpudirect-rdma#check-infiniband-port-status). For RoCE clusters, check [backend link health](/products/networking/hpc-interconnect/infiniband-roce-labels#when-speedcurrent-does-not-equal-speedexpected).
2. A stalled storage mount. An NFS or VAST mount that stops responding can fail the system calls underneath NCCL.
3. A peer rank that exited and dropped its connections.

If the error text ends in `Call to ibv_reg_mr_iova2 failed with error Invalid argument`, the cause is a container library mismatch, not the fabric. See [ibv\_reg\_mr\_iova2 failed with Invalid argument](#ibv_reg_mr_iova2-failed-with-invalid-argument).

## NCCL Error 1 (unhandled cuda error)

`unhandled cuda error` is NCCL Error 1: a CUDA call failed underneath NCCL. This is a GPU-side error, so start with the GPU rather than with storage or the fabric. Check the node's kernel log for a GPU XID event at the time of the failure, then check the driver version and GPU health on that node.

If the error carries `Cuda failure 101 'invalid device ordinal'`, it's the job-layout problem described in [NCCL Error 3](#nccl-error-3-internal-error) step 1, not a GPU fault.

## NCCL Error 6 (remote process exited)

NCCL Error 6 means either a peer rank exited while a collective was in flight, or the network path to that peer failed. The full message, `remote process exited or there was a network error`, covers both. The rank that prints Error 6 is almost always reporting someone else's failure, so start by finding the rank that died first:

1. Sort all rank logs by timestamp.
2. Find the earliest rank that printed a real failure (out of memory, segfault, XID, exit code). That rank is the cause.
3. Diagnose that rank's failure, not the Error 6 messages downstream of it.

If no rank printed a real failure first, treat it as the network half of the message: run an NCCL test on the same set of nodes ([how to run one](#run-an-nccl-test-on-the-same-set-of-nodes)).

## WorkNCCL timeout diagnostic order

A `Watchdog caught collective operation timeout` or `WorkNCCL ... timed out` means a collective didn't complete within the timeout. This is one of the most common and most misdiagnosed NCCL errors. Work through this exact order before suspecting the fabric:

1. **Dataloader or storage starvation.** Check whether a rank was waiting on data. Look for GPUs at 0 percent utilization at the time of the timeout and a drop in storage read throughput. A rank starved of data can't reach the collective, so every other rank times out waiting for it. This is the single most common cause.
2. **Silent rank crash before the timeout.** Check whether a rank crashed (out of memory, segfault) just before the timeout fired. The crash is the cause. The timeout is the symptom on the surviving ranks.
3. **NCCL test on the same set of nodes.** Run an NCCL test on the same nodes the job used ([how to run one](#run-an-nccl-test-on-the-same-set-of-nodes)). If it passes, the fabric is healthy and the cause is in the workload or storage.
4. **Only then suspect the fabric.** If the NCCL test fails on the same set of nodes, isolate the bad node by bisecting the nodelist, then [contact support](/support/contact) with the node names.

```mermaid theme={"system"}
---
title: WorkNCCL timeout diagnostic order
---
graph TD
start["WorkNCCL timeout"] --> data["1. Dataloader / storage starvation?"]
data -->|"GPUs at 0%, storage throughput dropped"| fixdata["Fix data pipeline"]
data -->|"No"| crash["2. Did a rank crash first?"]
crash -->|"OOM / segfault before timeout"| fixcrash["Diagnose the crashed rank"]
crash -->|"No"| test["3. Run NCCL test on same nodes"]
test -->|"Passes"| workload["Workload or storage issue"]
test -->|"Fails"| fabric["4. Bisect nodelist, contact support"]
```

## vendor err 249 compared to vendor err 129

These appear in InfiniBand transport errors such as `transport/net_ib.cc ... vendor err NNN`. They imply different fabric states:

* **`vendor err 249`** indicates an actively flapping link. The connection goes up and down during the run. This points at a physical or transient link problem on a specific node.
* **`vendor err 129`** indicates transport retries were exhausted. Check `NCCL_IB_RETRY_CNT` before the fabric: the value is stored in a 3-bit hardware register, so 7 is the maximum, and anything higher is silently truncated. A common setting of 16 truncates to zero retries, which makes the job die on the first transport error and makes a healthy fabric look broken. See [NCCL\_IB\_RETRY\_CNT has a hard maximum of 7](/products/networking/hpc-interconnect/nccl-configuration-reference#nccl_ib_retry_cnt-has-a-hard-maximum-of-7). If the value is 7 or lower and the error persists, isolate the node.

For both, use the onset pattern before you file. A sudden onset across several nodes at the same timestamp points at a fabric-level event, which is CoreWeave's to fix. A failure confined to one Pod points at that container's configuration, such as a `rdma-core` or UCX mismatch. If the onset points at the fabric, isolate the affected node by bisecting the nodelist, then [contact support](/support/contact) with the node names and the exact error string.

## ibv\_reg\_mr\_iova2 failed with Invalid argument

This error means the libibverbs or OFED userspace inside your container doesn't match the kernel RDMA modules on the node. `ibv_reg_mr_iova2` registers memory regions for RDMA; a newer userspace passes registration parameters that an older kernel module rejects with `EINVAL`. It reads like a fabric problem because the message mentions IB, but it's a local software version mismatch present on every node. Don't investigate switch health for this error.

Two signals make it close to definitive:

* **Single-node jobs succeed, multi-node jobs fail.** Single-node traffic goes over NVLink, which doesn't use libibverbs. Only the IB path calls `ibv_reg_mr_iova2`.
* **CoreWeave's `nccl-tests` image passes on the same nodes while your image fails.** The CoreWeave image is built against an OFED version known to match the node, so this narrows the problem to your container.

Ask what changed first. This almost always follows a CUDA base image bump, which pulls in newer OFED userspace.

**What to do:** Two paths, in order of preference.

1. **Upgrade the node image (preferred when the nodes are old).** If the nodes are running an old ncore or driver, upgrading is the right call regardless of this error. Set the driver version on the Node Pool, as described in [Update the GPU driver](/products/cks/nodes/gpu-driver-management/update-gpu-driver), then reboot the GPU nodes. Expect roughly an hour per node for the reboot and health-check cycle. Ask your CoreWeave representative to coordinate if you can't schedule the reboots yourself.
2. **Match the container to the node (faster, no reboot).** If you can't take the downtime, use a CUDA base image whose OFED version matches the node's, install a matching libibverbs in the container, or start from CoreWeave's `nccl-tests` image.

If the error persists after both, [contact support](/support/contact) with the image reference, the node's driver and ncore versions, and the exact error string.

## Stuck in NCCL communicator init after a previous crash

If the job stops responding during NCCL communicator initialization specifically after a previous run crashed, the cause is usually a corrupted on-disk compile cache, not NCCL or the fabric. A previous run that died without cleaning up can leave both the PyTorch Inductor and Triton caches in a bad state. The strong signal is a `destroy_process_group ... was not called` warning from the previous run, which means the prior process exited without tearing down cleanly.

The two caches default to different storage, which determines how far the corruption spreads:

* **Inductor** defaults to a `torchinductor_[USERNAME]` directory under the system temporary directory, normally `/tmp/torchinductor_$USER`. That's node-local, so corruption stays on the nodes the crashed run used.
* **Triton** defaults to `$HOME/.triton/cache`. Because home directories are on shared storage, corruption there follows your jobs to any node.

Two things move the Inductor cache, and both change how the corruption behaves. Setting `TORCHINDUCTOR_CACHE_DIR` to a path on shared storage is the obvious one, and it makes the Inductor cache behave like the Triton case. The subtler one is `TMPDIR`: Inductor resolves its default through Python's temporary directory, which honors `TMPDIR`, so wherever `TMPDIR` points is where the cache lands. Check `TMPDIR` before you assume the cache is on the node.

Under Pyxis or Enroot, `/tmp` inside the container is `tmpfs` (RAM) rather than the node's NVMe, so an Inductor cache left at the default doesn't survive the container at all and can't be the stale cache you're chasing. It also consumes memory while the job runs. Jobs that set `TMPDIR` to an NVMe-backed path, as [Node-local storage and `/tmp` on Slurm nodes](/products/sunk/manage_sunk/node-local-storage-and-tmp) recommends, do keep a cache across runs on that node. If your job runs in a container and doesn't set `TMPDIR`, suspect the Triton cache in your home directory instead.

Fix it by clearing the caches or disabling them for the next run:

```bash theme={"system"}
# Clear the on-disk compile caches left by a crashed run.
# Safe to delete; they are regenerated on next launch.
# Resolves both the explicit overrides and a relocated TMPDIR.
rm -rf "${TORCHINDUCTOR_CACHE_DIR:-${TMPDIR:-/tmp}/torchinductor_$USER}" "${TRITON_CACHE_DIR:-$HOME/.triton/cache}"
```

```bash theme={"system"}
# Or disable the Inductor caches so a corrupt cache cannot be reused.
export TORCHINDUCTOR_FORCE_DISABLE_CACHES=1
```

Don't rely on a node reboot to fix this. Replacing or rebooting a node clears a node-local Inductor cache, but it leaves the Triton cache in your home directory untouched, so a corrupted Triton cache follows the job to its next node. Clear the caches explicitly instead. To prevent cross-job corruption, point the cache paths at a per-job-id directory so a crash in one job can't corrupt the cache another job reuses.

## Checkpoint segfault at scale, followed by a collective timeout

A segfault during checkpoint saving followed by an NCCL collective timeout is a driver-level problem, not a fabric fault. The signature on GB200 nodes running driver 580 is specific:

* The segfault lands in `crc32_16bytes`, the checksum routine that verifies checkpoint data.
* The segfault appears in the log **before** the NCCL timeout. The timeout is the surviving ranks waiting on the ranks the segfault killed.
* It's scale-dependent, more likely at eight racks or more, and it can follow the workload to another cloud provider after the same driver upgrade.

The cause is CDMM (Confidential Data Memory Management), which is enabled by default on driver 580 for GB200 and changes the timing of GPU-to-CPU transfers. The checkpoint code reads host memory before the transfer finishes. Driver 570 doesn't enable CDMM by default and isn't affected.

If you see a segfault before an NCCL timeout in a checkpointing job, don't investigate the IB fabric, NVLink, or NCCL configuration. They'll all read healthy. Work these options instead, lightest first:

1. Call `torch.cuda.synchronize()` before checkpointing, so all in-flight GPU work completes before the save begins.
2. Move checkpoint writes to object storage, which changes the data path.
3. Roll the Node Pool back to driver 570. See [Update the GPU driver](/products/cks/nodes/gpu-driver-management/update-gpu-driver).

## Disambiguate a training stall from an NCCL stall

When a job stops responding and you can't tell whether NCCL is stuck or the stall is in your code upstream of NCCL, turn on both debug channels:

```bash theme={"system"}
# Surface both compile-side and NCCL-side activity to locate the hang.
# Read-only diagnostics; verbose. Enable on a reproducing run.
export TORCH_COMPILE_DEBUG=1
export NCCL_DEBUG=INFO
```

If NCCL is silent in the logs at the time of the stall, the stall is upstream of NCCL, in your code or in compilation. If NCCL is actively logging and then stops at a collective, the stall is at that collective.

## all\_to\_all and mixture-of-experts sensitivity

A passing `all_reduce` doesn't certify the fabric for an `all_to_all` workload, and the reason is the traffic pattern. Mixture-of-experts models use `all_to_all_single` for expert parallelism, which needs point-to-point communication between every pair of participating nodes. `all_reduce` uses a ring or tree topology, so each rank only talks to its neighbors. One unreachable node breaks an `all_to_all` outright while an `all_reduce` routes around it.

So if your workload uses `all_to_all` and fails, test with `all_to_all`. A clean `all_reduce` benchmark on the same set of nodes can hide a connectivity problem that only the workload's actual collective triggers.

## Run an NCCL test on the same set of nodes

Several sections above tell you to run an NCCL test on the nodes the job used. This is how you get a comparable result:

* **Get the test jobs.** CoreWeave publishes sample NCCL test jobs for Slurm and for the MPI Operator. See the [`coreweave/nccl-tests` repository](https://github.com/coreweave/nccl-tests/blob/master/README.md#running-nccl-tests) for the jobs and instructions. CoreWeave's `slurmd` images are built on these images, so the binaries are already present in a standard SUNK container. See [Slurm images](/products/sunk/reference/slurm-images).
* **Match the collective to your workload.** Run `all_reduce` for a data-parallel job and `all_to_all` for a mixture-of-experts job, for the reason in the preceding section.
* **Use the same nodes.** Pass the job's nodelist so the test exercises the same paths. A test on different nodes proves nothing about the ones that failed.
* **Compare against a baseline, not a feeling.** Record the bus bandwidth (`busbw`) the test reports on a small known-good group of nodes first, then compare. For how to baseline and how to bisect a nodelist down to the bad node, see [Why is my multi-node NCCL training slow?](/support/platform/articles/why-is-my-multi-node-nccl-training-slow).

If your own harness reports a problem but CoreWeave's `nccl-tests` pass on the same nodes, the gap is in the harness, not the fabric. Check message sizes, NCCL environment variables, and whether the timing loop calls `torch.cuda.synchronize()`.

## When to open a support ticket

Self-diagnose first using the preceding decision trees. Open a ticket in any of these situations:

* An NCCL test fails on a specific set of nodes and you have isolated the node by bisection.
* You see `vendor err 249` (actively flapping link) arriving on several nodes at once.
* On an InfiniBand cluster, `ibstat` shows a port intended for the job's InfiniBand traffic as `Down` or `Disabled`, and the port doesn't recover.

Alongside the details every ticket needs, listed in [Support ticket templates](/support/ticket-templates), include the exact error string, the node names, the NCCL test output, and `ibstat` output.

## What this is not

The following points correct the most common misdiagnoses:

* NCCL Error 3 isn't an NCCL bug in almost all cases, and it isn't the storage-first case either. Check the job's GPU request and node-to-node connectivity first.
* A WorkNCCL timeout isn't a fabric problem by default. It's most often a starved dataloader. This is the error the storage-first order applies to, not Error 3.
* A passing `all_reduce` test doesn't certify the fabric for an `all_to_all` workload.
* `unhandled cuda error` isn't a system or fabric error. It's NCCL Error 1, and it points at the GPU or the driver.
* `ibv_reg_mr_iova2 failed` isn't an IB fabric fault. It's a container library mismatch on every node.

## Related pages

* [Training job exit-code triage](/support/sunk/articles/what-do-training-job-exit-codes-mean)
* [torch.distributed rendezvous and TCPStore troubleshooting](/support/sunk/articles/how-do-i-fix-torch-distributed-rendezvous-failures): failures during startup, before training begins.
* [Measuring MFU and job performance](/products/sunk/optimize_workloads/measuring-mfu-and-job-performance): data loading and step-time diagnosis.
* [Why is my multi-node NCCL training slow?](/support/platform/articles/why-is-my-multi-node-nccl-training-slow): baselining with nccl-tests and bisecting a nodelist.
* [NCCL configuration reference](/products/networking/hpc-interconnect/nccl-configuration-reference): the NCCL environment variables CoreWeave documents.

<Badge stroke shape="pill" color="blue" size="md">[Server Errors](/support/sunk/tags/server-errors)</Badge><Badge stroke shape="pill" color="blue" size="md">[Workload Scheduling](/support/sunk/tags/workload-scheduling)</Badge>
