Skip to main content
To view the dashboard, go to the Kueue Scheduling Dashboard.
For instructions about accessing CoreWeave Grafana dashboards, see Access and use CoreWeave Grafana dashboards.
The Kueue Scheduling Dashboard provides an overview of Kueue scheduling activity in your cluster, including workload admissions, ClusterQueue and cohort status, preemptions, evictions, and admission wait times. It also shows the GPU topology segments available for scheduling, so you can see how much contiguous capacity is available for workloads of different sizes. Use this dashboard to monitor queue health, diagnose admission delays, and understand how workloads move through your cluster’s queues. Filter the panels with the dashboard variables, such as accelerator, org, cluster, Cluster Queue, and resource. Most panels are built on Kueue’s Prometheus metrics. For the full list of metrics and their meanings, see the Kueue metrics reference. The following sections describe the panels available in the Kueue Scheduling Dashboard.

Cluster topology

The following panels show GPU capacity as topology segments. A segment means free GPUs in the same topology domain that a workload can be scheduled into. Segment sizes of 8 or fewer count free GPUs per Node, and larger segment sizes count free GPUs per NVLink domain. For more about segments and topology domains, see Topology and block scheduling in Slurm.

Kueue health

The following panels show the health of the Kueue controller manager and its admission activity:

ClusterQueue

The following panels show workload activity, quotas, and wait times for the selected ClusterQueues:

Cohort

The following panels aggregate workload and quota metrics across each cohort subtree:

Workloads pending ready

The following panels track how long workloads take to have ready Pods, measured from workload creation and from admission. They appear in the dashboard’s Workloads Pending Ready row, which requires the waitForPodsReady setting in your Kueue configuration.
Last modified on July 23, 2026