Skip to main content
A memory kill on CKS is either one of the following:
  • Container OOM: Your container exceeded its own cgroup memory limit.
  • Node memory pressure: The Node ran out of physical memory.
To tell them apart, check the container’s last termination state. Replace [POD-NAME] with the Pod’s name and [NAMESPACE] with its namespace:
Reason: OOMKilled confirms a memory kill. To classify it, read the kernel OOM constraint from the Node’s kernel logs, if you have Node log access:
  • CONSTRAINT_MEMCG means a container limit you set too low
  • CONSTRAINT_NONE means Node-level pressure that may belong to CoreWeave.

OOM symptoms

  • A container restarts and kubectl describe pod shows Last State: Terminated, Reason: OOMKilled, Exit Code: 137.
  • A Pod shows Evicted with a message about the Node being low on memory.
  • Kernel logs contain a line such as oom-kill:constraint=CONSTRAINT_MEMCG or Memory cgroup out of memory: Killed process ....

Container OOM versus Node memory pressure

These are different failures with different fixes. Use the tables to help classify which failure it is.

Container OOM

Container OOM is the most common case CoreWeave support sees and is almost always a workload configuration issue.

Node memory pressure

Diagnose the OOM

Follow these steps to identify the cause of the OOM.
1

Read the kernel constraint

The kernel OOM log line names the constraint that triggered the kill. This is the most reliable signal.
Kernel OOM log lines
The Memory cgroup out of memory line reports the killed process:
Example cgroup OOM line
anon-rss is the resident memory the process used. If it’s close to your container’s memory limit, the container hit its own limit. total-vm is only virtual address space, not real memory use.The oom_memcg field in the oom-kill:constraint= line contains the Pod UID, with underscores instead of hyphens. Convert it before matching it to a Pod:
Matching a cgroup path to a Pod UID
Find the Pod with a given UID. Replace [POD-UID] with the UID from the previous step:
This command is safe to run anytime and it’s read-only.
2

Check the termination reason

Exit code 137 is 128 + 9 (SIGKILL). Rolling updates, Pod deletions, and Node drains also end in SIGKILL, so exit code 137 alone doesn’t prove an OOM.
  • Termination reason OOMKilled: a real OOM.
  • Restart count increments: a real OOM. It stays at 0 when the Pod is recreated by a rolling update or deletion.
  • A graceful Stopping container event sequence: normal termination.
Alert on the OOMKilled reason, not on exit code 137, or every rolling update produces false OOM alerts.
3

Find the source of the OOM

The namespace where you notice symptoms is often not where the OOM started.
  1. Find Pods with the highest restart counts cluster-wide. Safe; read-only.
  2. For any suspect Pod, read its last termination state. Safe; read-only.
Backlog-processing addons, such as policy report controllers, are common culprits: the backlog grows during CrashLoopBackOff, so each restart needs more memory than the last. Raise the limit well above the current value, let the backlog clear, then right-size it.A subprocess inside a multi-process Pod can be OOM-killed without the kubelet recording an OOMKilled event. If a Pod misbehaves with no OOMKilled event, check kernel logs if you have Node log access.

How to fix OOM problems

For container OOM (CONSTRAINT_MEMCG, MemoryPressure=False): For Node memory pressure (CONSTRAINT_NONE, MemoryPressure=True):
  • Reduce the total memory your Pods use on the Node, or spread the workload across more Nodes.
  • Set requests to match actual usage so the scheduler doesn’t overpack the Node.
  • If there’s no apparent workload cause, contact support with the Node name and the OOM events.

When a CoreWeave-managed Pod is the one killed

If the killed process is node_exporter, the CoreWeave-managed monitoring agent in the cw-exporters namespace, the cause is usually still a workload that creates an unusually large number of processes, sockets, mounts, or file descriptors, which node_exporter enumerates on every scrape. This case looks like:
  • The killed process is node_exporter, not one of your processes.
  • Multiple Nodes are affected at about the same time, matching where the workload runs.
  • The Nodes show no MemoryPressure condition.
List cgroup OOM kill events cluster-wide:
Look for what changed recently, such as a new monitoring agent or a workload with many processes per Node, and reduce that count. For other CoreWeave-managed Pods killed with no workload cause, contact support with the Node name and the OOM events.

When the memory graphs mislead you

  • A Pod that OOMs within seconds of starting shows near-zero memory in dashboards, because it crashes before the metrics scrape records peak usage. When a Pod is in CrashLoopBackOff, trust the termination reason, the restart count, and the kernel log, not the memory graph.
  • In Grafana memory graphs, “used” memory includes page cache, which is not memory pressure: the kernel reclaims cache freely. Judge Node-level pressure by available memory, not used memory: a Node with hundreds of GB available isn’t memory-constrained even if the used graph looks full.
  • Available memory counts /dev/shm and other tmpfs memory as reclaimable cache, but the kernel can’t reclaim it. Workloads that use /dev/shm heavily can OOM a Node whose available memory looks adequate. To check, compare node_memory_Shmem_bytes to node_memory_Cached_bytes in Grafana: if shmem tracks cached, most of the “cache” can’t be reclaimed. Reduce shared-memory usage or leave real headroom.

Common confusions

  • Exit code 137 isn’t always an OOM. Read the termination reason.
  • “Connection refused” between your Pods can be an OOM symptom, not a network problem: an OOM-killed server process is no longer listening on its port. Check kubectl get events --field-selector involvedObject.name=[NODE-NAME] for SystemOOM events before investigating networking.
  • FailedKillPod events with DeadlineExceeded are a different failure: the container runtime couldn’t stop the container. If a Pod stays stuck in Terminating, contact support with the Pod and Node names.
  • A single Node that consistently OOMs on workloads that succeed on identical Nodes may have a hardware state issue. Contact support with the Node name.
Server Errors
Last modified on September 2, 2026