Skip to main content
Your torch.distributed job fails during startup or rendezvous with a c10d or TCPStore error such as sendBytes failed ... Broken pipe, waitForInput ... timed out, or DistStoreError. These strings look alike and have different causes, so match your exact error before you change anything. A Broken pipe on a random rank is worth ruling out as a known PyTorch bug first, while a TCP client failed to connect/validate to host is a name-resolution problem that a PyTorch upgrade won’t touch. For the standard launch flow, see Submit a training job.

What you see

Your job fails during startup or rendezvous with a c10d or TCPStore error such as one of the following:
  • [c10d] sendBytes failed on SocketImpl ... Broken pipe
  • [c10d] sendBytes failed on SocketImpl ... Connection timed out
  • [c10d] waitForInput ... timed out ... 60000 ms
  • torch.distributed.DistStoreError
  • TCP client failed to connect/validate to host
  • failed to recv, got 0 bytes
These come from the c10d store, the coordination service the ranks use to find each other. They don’t share a cause, so match your exact string first.

Match the error string

Note the difference between Broken pipe and Connection timed out. They look alike and have different causes, so don’t treat them interchangeably.

Check the PyTorch version first

For a Broken pipe or DistStoreError failure, a version check is the cheapest first step: these have been traced to a c10d TCPStore bug often enough that CoreWeave support asks for the PyTorch version before investigating anything else. CoreWeave support has seen this affect PyTorch 2.7.x and earlier (reproduced on 2.7.1+cu128) and has confirmed it resolved on 2.9.1. Upstream PyTorch changed TCPStore connection handling in 2.8.0, but no PyTorch release note attributes a sendBytes broken-pipe fix to a specific version, so treat 2.8.0 as the version to move past rather than as a documented fix boundary. If you’re on 2.7.x or earlier, upgrade to a current release and retest before debugging your cluster or fabric. The signature to match is a failure that appears on a random rank, depends on which node the rank lands on, and doesn’t reproduce by rerunning on the same allocation.

Check hostname resolution

If the error is TCP client failed to connect/validate to host, or a waitForInput timeout preceded by a name-resolution warning such as The IPv6 network addresses of ([HOST], [PORT]) cannot be retrieved (gai error: -2 - Name or service not known), the ranks can’t resolve the rendezvous host. That’s a cluster configuration problem, not a PyTorch bug, and upgrading PyTorch won’t help. Check resolution from inside the job:
If it returns nothing while connecting by IP address works, contact support with the cluster name and the failing hostname. On SUNK self-service clusters this has been caused by slurmd Pods missing the subdomain they need for name resolution, which is a platform-side fix.

Discriminators between a version bug and infrastructure

Use these signals to decide whether you’re hitting a PyTorch bug or a real infrastructure or configuration problem: A 60000 ms timeout you didn’t configure isn’t one of your settings, but it doesn’t identify the version bug on its own: the same 60-second internal timeout fires when a hostname fails to resolve. Read the log lines above the timeout before you conclude which it is. Adjusting the timeouts you can configure won’t fix either case, because none of them control the 60-second internal wait. init_process_group(timeout=...), --rdzv_timeout, TORCHELASTIC_TIMEOUT, TORCH_DISTRIBUTED_CONNECTION_TIMEOUT, and DataLoader(timeout=...) are all different timeouts. Changing them is a dead end for this failure. If most signals point at a version bug, upgrade PyTorch. If they point at infrastructure or configuration, work through the following configuration checks.

Required rendezvous parameters

torchrun requires a consistent rendezvous configuration across all ranks. The required parameters for the c10d backend are:
  • --rdzv-id must be unique per job and identical across all of that job’s ranks. Using the Slurm job id is a reliable way to get a unique, shared id.
  • --rdzv-backend=c10d selects the built-in c10d store.
  • --rdzv-endpoint is the host and port where the rendezvous server runs. It must resolve and be reachable from every rank.
For static rendezvous (no elastic membership), set MASTER_ADDR and MASTER_PORT instead, and they must be identical across ranks.

Common misconfigurations

The following misconfigurations commonly cause rendezvous failures:
  • Duplicate rdzv-id across concurrent jobs. If two jobs running at the same time use the same rendezvous id, their ranks attempt to join the same group and collide. Derive the id from the job id so concurrent jobs never share one. See the $SLURM_JOB_ID pattern in Submit a training job.
  • Rendezvous host resolved before it’s ready. Ranks attempt to reach the rendezvous endpoint before the host is running, so connections are refused or reset. Make sure the rendezvous host is running before the other ranks attempt to connect.
  • Agent or store port conflicts. The rendezvous port must be free on the host. Deriving a port from the job id, as the tutorial does, reduces the chance of two jobs picking the same port on the same node. The tutorial’s scheme uses the last four characters of the job id, so it narrows the odds rather than eliminating them. A port clash is still worth checking.

Capture collective stack traces

When a job stops responding during or after rendezvous and you can’t tell which rank is stuck, enable the flight recorder to capture per-rank collective stack traces:
When the job stops responding or times out, the flight recorder shows what collective each rank was waiting on, which usually points directly at the stuck rank.

When to open a support ticket

Open a ticket only after you’ve ruled out a PyTorch version bug and checked the rendezvous configuration and hostname resolution. Alongside the details every ticket needs, listed in Support ticket templates, include the exact error string, your PyTorch version, the launcher command (with --rdzv-* flags), which rank failed, and whether it’s reproducible on the same allocation.

What this is not

The following errors are commonly misdiagnosed:
  • A sendBytes ... Broken pipe on a random rank isn’t necessarily a network fault. Rule out a TCPStore version bug before you treat it as one. A broken pipe can also mean the store or a peer process genuinely died, so check whether another rank crashed first.
  • A 60000 ms timeout you didn’t configure isn’t your rendezvous setting. It’s an internal c10d wait that none of your configurable timeouts control, so it points at either the version bug or a hostname that didn’t resolve, not at your launcher flags.
  • A rendezvous collision between two jobs isn’t an infrastructure problem. It’s a shared rdzv-id or port.
Server Errors Workload Scheduling
Last modified on September 29, 2026