| Summary | High-level “at-a-glance” cards for host identity, CPU and GPU utilization, network traffic, active alerts, and HPC verification status. |
| Logs | A rolling log view that merges Kubernetes events, kernel messages, Node Problem Detector findings, and recent alert annotations. Ideal for pinpointing the root cause of spikes or failures. |
| Network | End-to-end connectivity metrics, including ICMP loss from the Ping Exporter, conntrack table usage, interface throughput, packet error rates, and more. |
| Resources | Capacity and utilization for CPU, GPU, memory, NFS mounts, and local disk I/O, plus historical usage charts to identify resource pressure. |
| GPUs | Everything GPU-related, including ECC errors, power draw, clock speeds, memory usage, thermal headroom, fan RPM, and NVSwitch and NVLink health. |
| Temperatures | Real-time thermal data for GPU cores, HBM memory, motherboard sensors, and per-Pod thermal impact. Helpful for identifying cooling hot spots. |
| InfiniBand | Port state, link speed, retransmit counts, and congestion indicators for Nodes equipped with InfiniBand adapters. |
| Slurm Info | Node state within your Slurm cluster (idle, alloc, drain), running jobs, and allocation timelines. Useful for mixed Kubernetes and Slurm environments. |
| Kubernetes | Pod allocation breakdowns, taints, Reservation status, CPU-hours per job, memory by namespace, and Calico network policy statistics. |
| Verification | Results from the latest HPC verification suite, including GPU compute benchmarks, stress tests, and smoke tests. |
| Hardware | Low-level chassis data, including IPMI power readings, fan speeds, serial numbers, PCIe link status, and NVLink bandwidth graphs. |
| Unsorted | Additional metrics that do not yet belong to a dedicated group. Check here for new or experimental panels. |