> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Use GPUDirect RDMA with InfiniBand

> Configure GPUDirect RDMA over InfiniBand on CoreWeave, including Node Pool selection and NCCL settings

In this guide, learn how to use [GPUDirect Remote Direct Memory Access (RDMA) with InfiniBand](/products/networking/hpc-interconnect/about-hpc-interconnect#about-gpudirect-rdma) at CoreWeave, and how to test it with the NVIDIA Collective Communications Library (NCCL). GPUDirect RDMA enables direct memory access between GPUs across Nodes over InfiniBand, which reduces latency and frees CPU cycles for distributed training and other multi-Node GPU workloads. This guide is for users running multi-Node GPU workloads on CKS or Slurm who need high-bandwidth, low-latency communication between GPUs.

## Prerequisites

CoreWeave supports GPUDirect RDMA over InfiniBand for some [GPU instance types](/platform/instances/gpu-instances).

To use this feature, you must:

* Select a Node Pool with InfiniBand support.
* Install NCCL and the OpenFabrics Enterprise Distribution (OFED) driver in the Pod image.
* Configure the Pods to use GPUDirect RDMA.

## Select a Node Pool with InfiniBand support

To use GPUDirect RDMA, make sure the [Node Pool](/products/cks/nodes/nodes-and-node-pools) has Nodes with InfiniBand, as shown in our [list of GPU instance types](/platform/instances/gpu-instances).

All Nodes with InfiniBand have the required kernel drivers pre-installed. CKS manages all the required driver and operator dependencies. To avoid Node instability, you should not install other driver management tools.

Once you have a Node Pool that supports InfiniBand, the next section covers the Pod-level configuration that opts workloads into GPUDirect RDMA.

## Configure the Pods

Configure your Pods to use GPUDirect RDMA over InfiniBand by following these steps:

1. Set the value of `spec.containers.resources.requests.rdma/ib` to `1`.

   This value doesn't indicate the number of InfiniBand devices requested. It works as a boolean to schedule Pods onto servers with InfiniBand support.

   Kubernetes schedules resources through `requests` and `limits`. When you specify only `limits`, Kubernetes sets `requests` to the same amount as the limit. For more information, see [Resource management for Pods and containers](https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/) in the Kubernetes documentation.

   For a full YAML example showing how to set the `rdma/ib` value in the Pod spec for both `requests` and `limits`, see [Kubernetes example](#kubernetes-example).

2. Configure the Pods to use GPUDirect RDMA by setting these environment variables:

   * `NCCL_SOCKET_IFNAME`: The network interface name to use for NCCL communication. Set this to the InfiniBand interface name.
   * `NCCL_IB_HCA`: The InfiniBand host channel adapter (HCA) to use for NCCL communication.
   * `UCX_NET_DEVICES`: The network devices to use for Unified Communication X (UCX) communication. Set this to the InfiniBand interface name.

For Kubernetes and Slurm examples, see [Kubernetes example](#kubernetes-example) and [Slurm example](#slurm-example).

1. Optional: Enable extended logging with the `NCCL_DEBUG` environment variable.

   To increase the verbosity of NCCL's logging, set the `NCCL_DEBUG` environment variable to `INFO` for extra debug information. This helps diagnose issues with RDMA support, but it increases the log file size, so disable it when testing is complete. See `NCCL_DEBUG` in the [NCCL documentation](https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/env.html?#nccl-debug) for more logging options.

### Kubernetes example

When you deploy a Kubernetes Pod in the cluster, use the highlighted lines in this example to set the `rdma/ib` value in the Pod spec for both `requests` and `limits`, and to set the required environment variables.

```yaml title="Kubernetes example with debug logging" highlight={9,14,16-24} theme={"system"}
# [...]
spec:
  containers:
  - name: example
    resources:
      requests:
        cpu: 10
        memory: 10Gi
        rdma/ib: 1
        nvidia.com/gpu: 8
      limits:
        cpu: 10
        memory: 10Gi
        rdma/ib: 1
        nvidia.com/gpu: 8
  env:
    - name: NCCL_SOCKET_IFNAME
      value: eth0
    - name: NCCL_IB_HCA
      value: ibp
    - name: UCX_TLS
      value: tcp
    - name: UCX_NET_DEVICES
      value: eth0
    - name: NCCL_DEBUG
      value: INFO
# [...]
```

Setting `NCCL_DEBUG` to `INFO` enables extended logging. Remove it if you don't need extended logging.

### Slurm example

When you deploy a Slurm job, use the highlighted lines in this example to set the required environment variables. Remove `NCCL_DEBUG` unless you need extended logging.

```bash title="Example Slurm sbatch script" highlight={9-12} theme={"system"}
#!/bin/bash

#SBATCH --partition h100
#SBATCH --nodes 16
#SBATCH --ntasks-per-node 8
#SBATCH --gpus-per-node 8
# [...] other SBATCH options, as needed

export NCCL_SOCKET_IFNAME=eth0
export NCCL_IB_HCA=ibp
export UCX_TLS=tcp
export UCX_NET_DEVICES=eth0
export NCCL_DEBUG=INFO # Remove NCCL_DEBUG unless debug logging
```

## Test with NCCL

After you configure your Pods or Slurm jobs, verify that GPUDirect RDMA works as expected by running NCCL tests across multiple Nodes.

CoreWeave provides several sample NCCL test jobs designed for use with Message Passing Interface (MPI) Operator or Slurm. To test GPUDirect RDMA support with InfiniBand, see the [`nccl-tests` repository](https://github.com/coreweave/nccl-tests/blob/master/README.md#running-nccl-tests) for the test jobs and instructions for running them.

## Troubleshoot NCCL over InfiniBand

If a multi-Node job runs but throughput is far below expectations, NCCL has usually fallen back from InfiniBand to TCP. To confirm which transport NCCL chose, set `NCCL_DEBUG=INFO` and look for `NET/IB` (InfiniBand, good) rather than `NET/Socket` (TCP fallback). From inside the Pod, run `ibstat` and confirm each port shows `State: Active` and `Physical state: LinkUp`.

The most common causes are a Pod spec missing `rdma/ib: 1`, missing or incorrect NCCL environment variables, or a Node Pool without InfiniBand. For a full diagnostic walkthrough, including baselining with nccl-tests and isolating a degraded Node, see [Why is my multi-node NCCL training slow?](/support/platform/articles/why-is-my-multi-node-nccl-training-slow).