For instructions about accessing CoreWeave Grafana dashboards, see Access and use CoreWeave Grafana dashboards.
Kubernetes Training Workloads
Grafana dashboard for monitoring training workload hardware, placement topology, and Node health
To view the dashboard, go to the Kubernetes Training Workloads dashboard.
The Kubernetes Training Workloads dashboard helps you monitor a training workload end to end. It combines hardware metrics for GPU utilization, network bandwidth (such as InfiniBand), and storage I/O (both local and NFS) with a topology view of where the workload’s Pods are placed across Nodes, NVLink domains, leaf groups, SuperPods, and cabinets. Use it to diagnose performance bottlenecks, spot unhealthy Nodes, and understand how placement affects your workload.
Select a workload with the dashboard variables, such as Org, Cluster, Namespace, Kind, Workload, and Pod. The dashboard resolves parent workloads, such as a JobSet, MPIJob, CronJob, or Deployment, down to their child Pods.
The following sections describe the panels available on the dashboard.
Last modified on July 23, 2026