NCCL error, NCCL WARN | A collective communication operation failed. The numeric code and message identify the cause, and the message often names a different layer than the one that failed. | What do NCCL errors mean?. |
NCCL error 3, internal error, ncclInternalError | Almost never an NCCL bug, despite the wording. Usually a job layout that breaks NCCL’s rank-to-device mapping, a broken transport path, a GPU fault, or application memory corruption. | NCCL Error 3. |
NCCL error 2, unhandled system error, ncclSystemError | A system call underneath NCCL failed: the network, the OS, or the driver. | NCCL Error 2. |
unhandled cuda error, ncclUnhandledCudaError | NCCL Error 1: a CUDA call failed underneath NCCL. A GPU-side error, so start with GPU XIDs and the driver. | NCCL Error 1. |
remote process exited or there was a network error, ncclRemoteError | NCCL Error 6: a peer rank crashed, or the network path to it failed. Find the rank that died first. | NCCL Error 6. |
WorkNCCL ... timed out, Watchdog caught collective operation timeout | A torch.distributed collective didn’t complete within the timeout. Most often a starved dataloader or slow storage, not the fabric. | WorkNCCL timeout diagnostic order. |
vendor err 249, vendor err 129 | An InfiniBand transport error. 249 is an actively flapping link; 129 is transport retries exhausted, often from an NCCL_IB_RETRY_CNT above 7. | vendor err 249 compared to vendor err 129. |
ibv_reg_mr_iova2 failed, Invalid argument (RDMA registration) | The container’s libibverbs or OFED userspace doesn’t match the node’s kernel RDMA modules. Not an IB fabric fault. | ibv_reg_mr_iova2 failed with Invalid argument. |
destroy_process_group ... was not called | A previous run exited without tearing down, which can leave a corrupted torch.compile or Triton cache that hangs the next run during init. | Stuck in NCCL communicator init. |
Exit code 137, 139, 143, 132, 135 (or -9, -11, -15, -4, -7) | A signal killed the training process, where the exit code is 128 + signal number. The same code has several possible causes. | What do training job exit codes mean?. |
Illegal instruction, DataLoader worker pid ... killed by signal: Illegal instruction | SIGILL. On Grace (aarch64) instances, x86_64 native code reaching an Arm CPU, often through a shared JIT cache or an x86 numpy or BLAS build. | SIGILL on aarch64 and Run jobs on Grace CPU instances. |
c10d sendBytes failed ... Broken pipe, waitForInput ... timed out, DistStoreError, TCP client failed to connect/validate to host | A torch.distributed rendezvous failure. These strings look alike and have different causes, from a PyTorch version bug to a hostname that didn’t resolve. | Fix torch.distributed rendezvous and TCPStore failures. |
XID (for example XID 31, XID 79, XID 119) | An NVIDIA driver-reported GPU error. The XID number classifies the fault, from an application bug to a hardware failure. | What do GPU XID error codes mean?. |
CUDA error: out of memory | The process requested more GPU memory than is free. Reduce batch size or model footprint. | Recommended resource requests and limits. |
uncorrectable ECC error, Xid ... row remap | A GPU memory hardware fault. CoreWeave detects these and replaces affected hardware. | Interpret GPU health events and Node cordoning. |