> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# How do I fix torch.distributed rendezvous and TCPStore failures?

Your `torch.distributed` job fails during startup or rendezvous with a c10d or TCPStore error such as `sendBytes failed ... Broken pipe`, `waitForInput ... timed out`, or `DistStoreError`. These strings look alike and have different causes, so match your exact error before you change anything. A `Broken pipe` on a random rank is worth ruling out as a known PyTorch bug first, while a `TCP client failed to connect/validate to host` is a name-resolution problem that a PyTorch upgrade won't touch. For the standard launch flow, see [Submit a training job](/products/sunk/tutorials/train-on-sunk/3-submit-a-training-job).

## What you see

Your job fails during startup or rendezvous with a c10d or TCPStore error such as one of the following:

* `[c10d] sendBytes failed on SocketImpl ... Broken pipe`
* `[c10d] sendBytes failed on SocketImpl ... Connection timed out`
* `[c10d] waitForInput ... timed out ... 60000 ms`
* `torch.distributed.DistStoreError`
* `TCP client failed to connect/validate to host`
* `failed to recv, got 0 bytes`

These come from the c10d store, the coordination service the ranks use to find each other. They don't share a cause, so match your exact string first.

## Match the error string

| Error string | Usual cause | First move |
| - | - | - |
| `sendBytes failed ... Broken pipe` | A TCPStore bug in older PyTorch releases | Check your PyTorch version |
| `sendBytes failed ... Connection timed out` | The destination rank was blocked and couldn't accept the connection, often mid-checkpoint | Find what that rank was doing at the timestamp |
| `waitForInput ... timed out ... 60000 ms` | Either the same TCPStore bug, or a hostname that didn't resolve | Read the lines above it for a name-resolution error |
| `failed to recv, got 0 bytes` | A transient TCP blip during communicator bootstrap | Relaunch the job. Investigate only if it reproduces |
| `TCP client failed to connect/validate to host` | The rendezvous host's name didn't resolve | Check hostname resolution from inside the job |
| `DistStoreError` during rendezvous | Same family as the broken-pipe bug when the discriminators below match | Check your PyTorch version |

Note the difference between `Broken pipe` and `Connection timed out`. They look alike and have different causes, so don't treat them interchangeably.

## Check the PyTorch version first

For a `Broken pipe` or `DistStoreError` failure, a version check is the cheapest first step: these have been traced to a c10d TCPStore bug often enough that CoreWeave support asks for the PyTorch version before investigating anything else.

CoreWeave support has seen this affect PyTorch 2.7.x and earlier (reproduced on 2.7.1+cu128) and has confirmed it resolved on 2.9.1. Upstream PyTorch changed TCPStore connection handling in 2.8.0, but no PyTorch release note attributes a `sendBytes` broken-pipe fix to a specific version, so treat 2.8.0 as the version to move past rather than as a documented fix boundary. If you're on 2.7.x or earlier, upgrade to a current release and retest before debugging your cluster or fabric.

The signature to match is a failure that appears on a random rank, depends on which node the rank lands on, and doesn't reproduce by rerunning on the same allocation.

## Check hostname resolution

If the error is `TCP client failed to connect/validate to host`, or a `waitForInput` timeout preceded by a name-resolution warning such as `The IPv6 network addresses of ([HOST], [PORT]) cannot be retrieved (gai error: -2 - Name or service not known)`, the ranks can't resolve the rendezvous host. That's a cluster configuration problem, not a PyTorch bug, and upgrading PyTorch won't help.

Check resolution from inside the job:

```bash theme={"system"}
# Confirm the rendezvous host resolves from a compute node. Safe; read-only.
getent hosts $RDZV_HOST
```

If it returns nothing while connecting by IP address works, [contact support](/support/contact) with the cluster name and the failing hostname. On SUNK self-service clusters this has been caused by `slurmd` Pods missing the subdomain they need for name resolution, which is a platform-side fix.

## Discriminators between a version bug and infrastructure

Use these signals to decide whether you're hitting a PyTorch bug or a real infrastructure or configuration problem:

| Signal | Points at a PyTorch version bug | Points at infrastructure or config |
| - | - | - |
| Which rank fails | Random, node-dependent, different each run | Always the same rank or node |
| Timing | Fails within seconds of launch | Fails after minutes, or partway through |
| Timeout value | An internal default you didn't set (for example, 60000 ms) | The timeout you configured |
| When it started | After a container or framework upgrade | After a cluster or network change |

A 60000 ms timeout you didn't configure isn't one of your settings, but it doesn't identify the version bug on its own: the same 60-second internal timeout fires when a hostname fails to resolve. Read the log lines above the timeout before you conclude which it is.

Adjusting the timeouts you *can* configure won't fix either case, because none of them control the 60-second internal wait. `init_process_group(timeout=...)`, `--rdzv_timeout`, `TORCHELASTIC_TIMEOUT`, `TORCH_DISTRIBUTED_CONNECTION_TIMEOUT`, and `DataLoader(timeout=...)` are all different timeouts. Changing them is a dead end for this failure.

If most signals point at a version bug, upgrade PyTorch. If they point at infrastructure or configuration, work through the following configuration checks.

```mermaid theme={"system"}
---
title: Is this a PyTorch bug or a configuration problem?
---
graph TD
start["c10d / TCPStore failure"] --> rand["Random rank, fails in seconds, fixed internal timeout?"]
rand -->|"Yes"| ver["Likely a PyTorch version bug. Upgrade PyTorch."]
rand -->|"No"| cfg["Same rank, your timeout, after a cluster change?"]
cfg -->|"Yes"| conf["Check the rendezvous configuration"]
cfg -->|"No"| both["Capture logs from all ranks and re-evaluate"]
```

## Required rendezvous parameters

`torchrun` requires a consistent rendezvous configuration across all ranks. The required parameters for the c10d backend are:

```bash theme={"system"}
# Launch torchrun with the c10d rendezvous backend.
# Intent: every rank joins the same rendezvous group using the same id and endpoint.
torchrun \
  --nnodes=$SLURM_NNODES \
  --nproc-per-node=8 \
  --rdzv-id=$SLURM_JOB_ID \
  --rdzv-backend=c10d \
  --rdzv-endpoint="$RDZV_HOST:$RDZV_PORT" \
  train.py
```

* `--rdzv-id` must be unique per job and identical across all of that job's ranks. Using the Slurm job id is a reliable way to get a unique, shared id.
* `--rdzv-backend=c10d` selects the built-in c10d store.
* `--rdzv-endpoint` is the host and port where the rendezvous server runs. It must resolve and be reachable from every rank.

For static rendezvous (no elastic membership), set `MASTER_ADDR` and `MASTER_PORT` instead, and they must be identical across ranks.

## Common misconfigurations

The following misconfigurations commonly cause rendezvous failures:

* **Duplicate `rdzv-id` across concurrent jobs.** If two jobs running at the same time use the same rendezvous id, their ranks attempt to join the same group and collide. Derive the id from the job id so concurrent jobs never share one. See the `$SLURM_JOB_ID` pattern in [Submit a training job](/products/sunk/tutorials/train-on-sunk/3-submit-a-training-job).
* **Rendezvous host resolved before it's ready.** Ranks attempt to reach the rendezvous endpoint before the host is running, so connections are refused or reset. Make sure the rendezvous host is running before the other ranks attempt to connect.
* **Agent or store port conflicts.** The rendezvous port must be free on the host. Deriving a port from the job id, as the tutorial does, reduces the chance of two jobs picking the same port on the same node. The tutorial's scheme uses the last four characters of the job id, so it narrows the odds rather than eliminating them. A port clash is still worth checking.

## Capture collective stack traces

When a job stops responding during or after rendezvous and you can't tell which rank is stuck, enable the flight recorder to capture per-rank collective stack traces:

```bash theme={"system"}
# Enable the NCCL flight recorder so collective stack traces are captured on hang.
# The value is the number of recorded entries, not bytes. 2000 is the documented default.
# Read-only effect on training; adds diagnostic capture. Safe to enable.
export TORCH_NCCL_TRACE_BUFFER_SIZE=2000
```

When the job stops responding or times out, the flight recorder shows what collective each rank was waiting on, which usually points directly at the stuck rank.

## When to open a support ticket

Open a ticket only after you've ruled out a PyTorch version bug and checked the rendezvous configuration and hostname resolution. Alongside the details every ticket needs, listed in [Support ticket templates](/support/ticket-templates), include the exact error string, your PyTorch version, the launcher command (with `--rdzv-*` flags), which rank failed, and whether it's reproducible on the same allocation.

## What this is not

The following errors are commonly misdiagnosed:

* A `sendBytes ... Broken pipe` on a random rank isn't necessarily a network fault. Rule out a TCPStore version bug before you treat it as one. A broken pipe can also mean the store or a peer process genuinely died, so check whether another rank crashed first.
* A 60000 ms timeout you didn't configure isn't your rendezvous setting. It's an internal c10d wait that none of your configurable timeouts control, so it points at either the version bug or a hostname that didn't resolve, not at your launcher flags.
* A rendezvous collision between two jobs isn't an infrastructure problem. It's a shared `rdzv-id` or port.

## Related pages

* [Submit a training job](/products/sunk/tutorials/train-on-sunk/3-submit-a-training-job): the working torchrun launch pattern.
* [Training job exit-code triage](/support/sunk/articles/what-do-training-job-exit-codes-mean): when rendezvous succeeds but the job exits with a code.
* [NCCL communicator init and checkpoint failures](/support/sunk/articles/what-do-nccl-errors-mean#stuck-in-nccl-communicator-init-after-a-previous-crash): when the job stops responding because of a barrier or a cache corruption, not rendezvous.
* [NCCL error reference and diagnostic decision tree](/support/sunk/articles/what-do-nccl-errors-mean)

<Badge stroke shape="pill" color="blue" size="md">[Server Errors](/support/sunk/tags/server-errors)</Badge><Badge stroke shape="pill" color="blue" size="md">[Workload Scheduling](/support/sunk/tags/workload-scheduling)</Badge>
