Skip to main content
Powered by four NVIDIA Rubin GPUs paired with NVIDIA Vera Arm CPUs, these instances target the most demanding AI training and inference workloads. Each Rubin GPU provides up to 288 GB of next-generation HBM4 memory and connects over sixth-generation NVLink (NVLink 6) for a rack-scale, high-bandwidth memory fabric. These instances form part of a larger NVL72 rack architecture with 20.7 TB of total GPU memory. Compared to GB200, Vera Rubin delivers up to 5x inference and 3.5x training performance. For clustering, these instances use next-generation NVIDIA Spectrum-X RoCE (RDMA over Converged Ethernet) with BlueField-4 and ConnectX-9 SuperNICs for large-scale AI in Ethernet-based cloud environments.

Specifications

1 Usable storage is less than the raw capacity when configured as RAID. See RAID layout and throughput to learn about RAID configurations and their performance characteristics.

Primary use cases

Training next-generation foundation models in the trillion-parameter class, massive-scale and high-fidelity inference on the most complex AI models, and memory-intensive scientific simulations that require maximum throughput. Next-generation frontier models (multi-trillion parameters), state-of-the-art multimodal systems, and large-scale scientific discovery models.
Last modified on September 30, 2026