Overview
This is the second step in the Deploy vLLM inference tutorial series. Before starting, complete step 1: set up infrastructure dependencies. Monitoring and observability are important for production inference workloads. This step sets up Prometheus for metrics collection and Grafana for visualization, giving you insights into your vLLM deployment’s performance. You also prepare the storage and authentication resources that the vLLM deployment in the next step depends on. The monitoring stack provides metrics for the following areas:- Request throughput and latency
- GPU utilization and memory usage
- KV cache performance
- Queue depth and autoscaling metrics
Resource allocationThe monitoring stack requires additional cluster resources. Ensure your CKS cluster has at least one CPU Node for Prometheus and Grafana to deploy to.
Step 1: Install Prometheus and Grafana
Clone the reference architecture repository:hack folder:
values.yaml file by replacing the orgID, clusterName, and hosts sections with your information.
orgID: You can get yourorgIDon the CoreWeave Console settings page.clusterName: You can get your cluster name on the CoreWeave Console Cluster page.
orgID is cw99 and your cluster name is my-inference-cluster, the values.yaml would look like the following:
values.yaml looks like the following:
observability/basic directory and deploy the monitoring stack from the observability/basic directory:
monitoring namespace:
Step 2: Verify monitoring deployment
List the Pods in themonitoring namespace:
Step 3: Get Grafana credentials
Retrieve the auto-generated Grafana admin password:Step 4: Create model cache storage
In this step, you create a namespace for the inference workload and a persistent volume claim that caches downloaded model weights so later Pod restarts don’t re-download large model files. Navigate to theinference/basic directory:
Step 5: Optional: Set up Hugging Face authentication
For models that require authentication, like the Llama 3.1 8B Instruct, create a secret with your Hugging Face token. Replace[HUGGINGFACE-TOKEN] with your actual Hugging Face access token.