NVIDIA driver on your system is too old, NCCL initialization failures, crashes that appear only at scale, and throughput changes after an image upgrade.
Symptoms this page covers
Image guidance
Every recommended image builds onnccl-tests, which carries the networking libraries tuned for CoreWeave’s InfiniBand fabric, including HPC-X. Pick the image that adds only what your workload needs on top of that foundation. Use the following flowchart to help decide what image to use:
torch and torch-extras also publish a base variant. Use it only when image size is a hard constraint and you have time to test the result. It has the following limitations:
- It isn’t built on
nccl-tests, so it lacks networking libraries such as HPC-X. - Although it keeps
nvcc, it omits the CUDA development libraries, which can make some workloads fail to build.
base image doesn’t degrade performance or compatibility.
The following table summarizes the choices:
The
torch, torch-extras, and nightly images are published to ghcr.io/coreweave/ml-containers. For what each one contains, see ML container images. nccl-tests is built from a separate repository and published to ghcr.io/coreweave/nccl-tests, so it doesn’t appear in that inventory and it uses a different tag format.
Choose an image tag
A tag records the component versions that most often decide compatibility and is usually enough for choosing an image. However, the tag doesn’t list everything in the image. Extension versions such as FlashAttention and DeepSpeed, and the HPC-X version inherited fromnccl-tests, aren’t in the tag. The leading commit identifies the exact build, so the full component list is in the coreweave/ml-containers repository at that commit.
The ghcr.io/coreweave/ml-containers images use the following tag format, shown here for a production build:
- The leading commit,
8a60b2d, indicates whether the tag is a production build or a branch build, which the next section covers. - The CUDA version,
cuda12.9.1, sets the minimum GPU driver the image requires. See Match CUDA to the Node driver.
Use unprefixed tags in production
CoreWeave publishes two kinds of tags to the same repository:- Production builds start directly with the commit hash, as in
8a60b2d-nccl-cuda12.9.1-.... These are built from the default branch. - Branch builds start with a prefix before the rest of the tag, as in
es-sm-10x-...orrenovate-.... The build pipeline adds the prefix automatically for any branch other than the default one, so the prefix is the source branch name, whether that branch belongs to an engineer or to an automated dependency bot.
Pick a nccl-tests tag
The nccl-tests images use a different format that leads with the CUDA version and ends with the commit, as in 12.9.2-devel-ubuntu22.04-nccl2.31.2-1-ee0d4d1. They don’t carry a branch prefix.
To pick a nccl-tests image, use the table of current stable images in the repository README, which CoreWeave keeps up to date. Those are the recommended tags, so you don’t need to read the packages list for nccl-tests.
Match CUDA to the Node driver
The container provides the CUDA toolkit. The Node provides the NVIDIA driver. NVIDIA defines two levels of compatibility between them, and most confusion about CUDA versions comes from mixing them up:- The minimum driver is set by the CUDA major version. Any CUDA 12.x toolkit loads on driver
525.60.13or later, and any CUDA 13.x toolkit loads on driver580or later. At this level the CUDA runtime starts, but a given program may or may not run. See Partial compatibility. - The driver for full feature support is the driver version or branch NVIDIA identifies for that CUDA toolkit release. Use it or a newer driver to support the release’s new features. For CUDA 13.x, NVIDIA lists corresponding driver branches, as shown in the following table.
12080 means the driver supports CUDA 12.8.
Drivers on CoreWeave Nodes
CoreWeave ships driver580 and 595, with 595 as the default, and has deprecated 535. Driver 580 fully supports CUDA toolkits through 13.0, and driver 595 fully supports CUDA toolkits through 13.2. Both run any CUDA 12.x image. For the per-instance driver table, see About GPU driver management in CKS.
Recommended driver per CUDA version
For an image built with another CUDA toolkit, use the driver version or branch listed for full feature support, or a newer driver:
Source: NVIDIA CUDA Toolkit release notes. CUDA 13.x entries use NVIDIA’s corresponding driver branches; update releases within a CUDA minor version use the same branch. CUDA 12.x entries list Linux driver versions for the toolkit’s GA release; update releases can require newer versions. Drivers are backward compatible with applications built using older CUDA toolkits.
Partial compatibility
NVIDIA defines driver525.60.13 as compatible with any CUDA 12.x release, and driver branch 580 as compatible with any CUDA 13.x release. Under that definition, carefully written software in a CUDA 13.3 container can run on driver 580 or 590, even though CUDA 13.3 requires driver branch R610 or later for full feature support.
However, many programs don’t work like this in practice. Use the drivers listed for full feature support in the preceding table, or newer drivers. If you need a newer CUDA toolkit than your Node’s driver fully supports, a given program may or may not crash, depending on how it was written and compiled, and you can’t predict the outcome from the image and driver versions alone.
Each CUDA minor version adds driver API calls and raises the PTX ISA version the toolkit emits. That produces two distinct runtime failures on a driver that meets the major-version minimum but lacks support for the toolkit’s new features:
- Precompiled code that calls a newer API. Code compiled to SASS with the CUDA 13.1 toolkit may call a function that first appeared in driver
590. On driver580, that call fails withcudaErrorCallRequiresNewerDriver. Only the code paths that use the newer function are affected, so the program may run for a while before it hits one. - PTX that the driver can’t compile. Code compiled to PTX with the CUDA 13.1 toolkit always requires PTX ISA 9.1 or later. Driver
580can JIT-compile PTX only up to ISA 9.0, so loading that code fails withCUDA_ERROR_UNSUPPORTED_PTX_VERSION.
Check the driver on your Nodes
To see which driver each Node runs, list the driver version label:580.105.08-0ubuntu1.
If the image’s CUDA toolkit needs a newer driver than your Nodes run, whether for the major-version minimum or for full support of its minor version, you have two options:
- Select an image tag built against an older CUDA toolkit.
- Pin a newer driver on the Node Pool. See Select GPU driver versions in CKS Node Pools.
Instances with Grace CPUs need aarch64 images
Instances with Grace CPUs, including GH200, GB200, and GB300, run theaarch64 architecture rather than x86_64. Two distinct image problems surface here. An x86_64-only image crashes with SIGILL, reported as exit code 132 or -4. Separately, an aarch64 binary compiled to assume 4 KB memory pages can fail with a misleading out-of-memory error from its own allocator, because Nodes with Grace CPUs use 64 KB pages.
CoreWeave’s nccl-tests and ml-containers images, including torch, torch-extras, and their nightly channels, publish a manifest list covering both linux/amd64 and linux/arm64, so a single tag runs on x86 instances and instances with Grace CPUs alike.
On instances with Grace CPUs, take the following precautions:
- If you use an image from outside CoreWeave, confirm it publishes an
aarch64manifest by runningdocker manifest inspect [IMAGE]before you deploy. - Rebuild custom binaries and compiled extensions for
aarch64. Don’t reuse JIT caches or checkpoints that contain compiled ops from an x86 cluster.
What not to override
The maintained images and the CoreWeave platform are tuned to work together, so replacing a component that the platform also provides is a recurring cause of failures.System NCCL
You rarely need to replace the NCCL build a maintained image ships. CoreWeave tracks the latest NCCL release in these images, so when you need a newer NCCL, upgrade to a newer base image instead of swapping the library in place. If you do replace NCCL in thenccl-tests image or its derivatives, it generally works, because most NCCL components have a stable application binary interface (ABI). To change NCCL behavior rather than its version, use environment variables. See the NCCL configuration reference.
libibverbs and OFED userspace
Thelibibverbs and OFED userspace in the container must match the kernel modules on the Node. A mismatch surfaces as ibv_reg_mr_iova2 failed: Invalid argument and similar errors. Don’t layer a different OFED userspace onto a maintained image. Rebuild against a supported CUDA and OFED base instead.
This isn’t a NCCL bug. NCCL reaches InfiniBand through libibverbs, so a userspace mismatch surfaces first as a NCCL failure.
GDRCopy kernel components
CoreWeave provides the GDRCopy kernel module on the Node. Don’t install a GDRCopy kernel driver inside the container. For the supported path, see NVSHMEM and GDRCopy support.Driver features that look missing
NVIDIA OptiX, EGL rendering, and hardware video encode and decode come from the host driver. The NVIDIA Container Toolkit mounts them into the container at runtime based onNVIDIA_DRIVER_CAPABILITIES. They aren’t built into the image, so a failure here is a capability setting, not a missing package.
CoreWeave’s torch-extras images ship with NVIDIA_DRIVER_CAPABILITIES=compute,utility by default, inherited from the NVIDIA CUDA base image. That value covers CUDA compute and nvidia-smi but excludes the following:
Check the current value inside the container:
OptiX denoising weights are not mounted
Even withNVIDIA_DRIVER_CAPABILITIES=all, the NVIDIA Container Toolkit can fail to mount /usr/share/nvidia/nvoptix.bin, the OptiX denoising weights, while mounting libnvoptix.so correctly. The upstream report, nvidia-container-toolkit issue 127, is closed, but CoreWeave Support has seen the behavior as recently as driver 580.
To work around it, install the matching userspace GL package in your image, replacing [DRIVER-MAJOR-VERSION] with the major version your Nodes run, such as 580:
Throughput changed after an image upgrade
An image upgrade is rarely a single-variable change. One new tag can move the CUDA, NCCL, PyTorch, torchvision, torchaudio, ABI, and Ubuntu versions at once, along with components the tag doesn’t record, and a change of base OS also moves the system Python and glibc versions.CoreWeave doesn’t publish per-tag performance baselines. A large throughput swing across a base image change, in either direction, more often reflects workload or configuration differences than the image itself. Establish your own baseline before you attribute a change to the image.
- List the delta. Compare the two tags field by field, using Read an image tag to identify each one, and record every component that changed. Then diff the two commits in the
coreweave/ml-containersrepository for components the tag doesn’t show, such as FlashAttention, DeepSpeed, and HPC-X. - Reproduce on both tags. Confirm the difference is consistent before you bisect.
- Change one field at a time. Hold the base OS fixed while you isolate a framework or NCCL change, since a base OS change also moves Python and glibc.
- Bisect through intermediate tags. The
ml-containersrepository publishes many builds between releases, which gives you intermediate points to test. - Run at production scale and duration. Gradual effects, such as step times that erode over hours, don’t appear in short runs.
Collect per-job GPU metrics
Image choice also affects how you collect GPU telemetry because some images ship their own DCGM packages. CoreWeave collects DCGM telemetry through a manageddcgm-exporter DaemonSet, so you don’t need to run dcgmi inside your containers. If your tooling requires DCGM in the container for per-job attribution, package it in your image and scrape its exporter.
If dcgmi diag inside a container reports Detected unsupported Cuda version, the container has the legacy DCGM 3.x package, which doesn’t support CUDA 13 or driver 580 and later. Check which package is installed:
datacenter-gpu-manager. The current packages are named datacenter-gpu-manager-4-*. The nccl-tests image installs DCGM through apt, and builds from before the switch to DCGM 4.x installed the legacy package regardless of their CUDA version. Any image derived from one of those older nccl-tests tags carries DCGM 3.x. To resolve it, rebuild on a current nccl-tests, torch, or torch-extras tag.
Before you open a support ticket
Confirm the following, since each one accounts for a large share of image-related failures:- The image’s CUDA toolkit version is supported by the driver your Nodes run.
- The image architecture matches the instance, which means an
aarch64manifest for instances with Grace CPUs. - You haven’t replaced the OFED userspace or GDRCopy inside a maintained image.
gpu.coreweave.cloud/driver-version label from an affected Node. For a throughput change, include both tags and your measurements at production scale.
Additional resources
For more information, see the following resources:- ML container images: the published images, variants, and what each contains.
- Packages list: every published image and tag.
coreweave/ml-containers: Dockerfiles and build source.- About GPU driver management in CKS: supported driver versions per instance type.
- NCCL configuration reference: NCCL settings for distributed training.