torch.distributed job fails during startup or rendezvous with a c10d or TCPStore error such as sendBytes failed ... Broken pipe, waitForInput ... timed out, or DistStoreError. These strings look alike and have different causes, so match your exact error before you change anything. A Broken pipe on a random rank is worth ruling out as a known PyTorch bug first, while a TCP client failed to connect/validate to host is a name-resolution problem that a PyTorch upgrade won’t touch. For the standard launch flow, see Submit a training job.
What you see
Your job fails during startup or rendezvous with a c10d or TCPStore error such as one of the following:[c10d] sendBytes failed on SocketImpl ... Broken pipe[c10d] sendBytes failed on SocketImpl ... Connection timed out[c10d] waitForInput ... timed out ... 60000 mstorch.distributed.DistStoreErrorTCP client failed to connect/validate to hostfailed to recv, got 0 bytes
Match the error string
Note the difference between
Broken pipe and Connection timed out. They look alike and have different causes, so don’t treat them interchangeably.
Check the PyTorch version first
For aBroken pipe or DistStoreError failure, a version check is the cheapest first step: these have been traced to a c10d TCPStore bug often enough that CoreWeave support asks for the PyTorch version before investigating anything else.
CoreWeave support has seen this affect PyTorch 2.7.x and earlier (reproduced on 2.7.1+cu128) and has confirmed it resolved on 2.9.1. Upstream PyTorch changed TCPStore connection handling in 2.8.0, but no PyTorch release note attributes a sendBytes broken-pipe fix to a specific version, so treat 2.8.0 as the version to move past rather than as a documented fix boundary. If you’re on 2.7.x or earlier, upgrade to a current release and retest before debugging your cluster or fabric.
The signature to match is a failure that appears on a random rank, depends on which node the rank lands on, and doesn’t reproduce by rerunning on the same allocation.
Check hostname resolution
If the error isTCP client failed to connect/validate to host, or a waitForInput timeout preceded by a name-resolution warning such as The IPv6 network addresses of ([HOST], [PORT]) cannot be retrieved (gai error: -2 - Name or service not known), the ranks can’t resolve the rendezvous host. That’s a cluster configuration problem, not a PyTorch bug, and upgrading PyTorch won’t help.
Check resolution from inside the job:
slurmd Pods missing the subdomain they need for name resolution, which is a platform-side fix.
Discriminators between a version bug and infrastructure
Use these signals to decide whether you’re hitting a PyTorch bug or a real infrastructure or configuration problem:
A 60000 ms timeout you didn’t configure isn’t one of your settings, but it doesn’t identify the version bug on its own: the same 60-second internal timeout fires when a hostname fails to resolve. Read the log lines above the timeout before you conclude which it is.
Adjusting the timeouts you can configure won’t fix either case, because none of them control the 60-second internal wait.
init_process_group(timeout=...), --rdzv_timeout, TORCHELASTIC_TIMEOUT, TORCH_DISTRIBUTED_CONNECTION_TIMEOUT, and DataLoader(timeout=...) are all different timeouts. Changing them is a dead end for this failure.
If most signals point at a version bug, upgrade PyTorch. If they point at infrastructure or configuration, work through the following configuration checks.
Required rendezvous parameters
torchrun requires a consistent rendezvous configuration across all ranks. The required parameters for the c10d backend are:
--rdzv-idmust be unique per job and identical across all of that job’s ranks. Using the Slurm job id is a reliable way to get a unique, shared id.--rdzv-backend=c10dselects the built-in c10d store.--rdzv-endpointis the host and port where the rendezvous server runs. It must resolve and be reachable from every rank.
MASTER_ADDR and MASTER_PORT instead, and they must be identical across ranks.
Common misconfigurations
The following misconfigurations commonly cause rendezvous failures:- Duplicate
rdzv-idacross concurrent jobs. If two jobs running at the same time use the same rendezvous id, their ranks attempt to join the same group and collide. Derive the id from the job id so concurrent jobs never share one. See the$SLURM_JOB_IDpattern in Submit a training job. - Rendezvous host resolved before it’s ready. Ranks attempt to reach the rendezvous endpoint before the host is running, so connections are refused or reset. Make sure the rendezvous host is running before the other ranks attempt to connect.
- Agent or store port conflicts. The rendezvous port must be free on the host. Deriving a port from the job id, as the tutorial does, reduces the chance of two jobs picking the same port on the same node. The tutorial’s scheme uses the last four characters of the job id, so it narrows the odds rather than eliminating them. A port clash is still worth checking.
Capture collective stack traces
When a job stops responding during or after rendezvous and you can’t tell which rank is stuck, enable the flight recorder to capture per-rank collective stack traces:When to open a support ticket
Open a ticket only after you’ve ruled out a PyTorch version bug and checked the rendezvous configuration and hostname resolution. Alongside the details every ticket needs, listed in Support ticket templates, include the exact error string, your PyTorch version, the launcher command (with--rdzv-* flags), which rank failed, and whether it’s reproducible on the same allocation.
What this is not
The following errors are commonly misdiagnosed:- A
sendBytes ... Broken pipeon a random rank isn’t necessarily a network fault. Rule out a TCPStore version bug before you treat it as one. A broken pipe can also mean the store or a peer process genuinely died, so check whether another rank crashed first. - A 60000 ms timeout you didn’t configure isn’t your rendezvous setting. It’s an internal c10d wait that none of your configurable timeouts control, so it points at either the version bug or a hostname that didn’t resolve, not at your launcher flags.
- A rendezvous collision between two jobs isn’t an infrastructure problem. It’s a shared
rdzv-idor port.
Related pages
- Submit a training job: the working torchrun launch pattern.
- Training job exit-code triage: when rendezvous succeeds but the job exits with a code.
- NCCL communicator init and checkpoint failures: when the job stops responding because of a barrier or a cache corruption, not rendezvous.
- NCCL error reference and diagnostic decision tree