- Container OOM: Your container exceeded its own cgroup memory limit.
- Node memory pressure: The Node ran out of physical memory.
[POD-NAME] with the Pod’s name and [NAMESPACE] with its namespace:
Reason: OOMKilled confirms a memory kill. To classify it, read the kernel OOM
constraint from the Node’s kernel logs, if you have Node log access:
CONSTRAINT_MEMCGmeans a container limit you set too lowCONSTRAINT_NONEmeans Node-level pressure that may belong to CoreWeave.
OOM symptoms
- A container restarts and
kubectl describe podshowsLast State: Terminated, Reason: OOMKilled, Exit Code: 137. - A Pod shows
Evictedwith a message about the Node being low on memory. - Kernel logs contain a line such as
oom-kill:constraint=CONSTRAINT_MEMCGorMemory cgroup out of memory: Killed process ....
Container OOM versus Node memory pressure
These are different failures with different fixes. Use the tables to help classify which failure it is.Container OOM
Container OOM is the most common case CoreWeave support sees and is almost
always a workload configuration issue.
Node memory pressure
Diagnose the OOM
Follow these steps to identify the cause of the OOM.1
Read the kernel constraint
The kernel OOM log line names the constraint that triggered the kill. This is the
most reliable signal.The Find the Pod with a given UID. Replace This command is safe to run anytime and it’s read-only.
Kernel OOM log lines
Memory cgroup out of memory line reports the killed process:Example cgroup OOM line
anon-rss is the resident memory the process used. If it’s close to your
container’s memory limit, the container hit its own limit. total-vm is only
virtual address space, not real memory use.The oom_memcg field in the oom-kill:constraint= line contains the Pod UID,
with underscores instead of hyphens. Convert it before matching it to a Pod:Matching a cgroup path to a Pod UID
[POD-UID] with the UID from the previous step:2
Check the termination reason
Exit code 137 is
128 + 9 (SIGKILL). Rolling updates, Pod deletions, and Node
drains also end in SIGKILL, so exit code 137 alone doesn’t prove an OOM.- Termination reason
OOMKilled: a real OOM. - Restart count increments: a real OOM. It stays at 0 when the Pod is recreated by a rolling update or deletion.
- A graceful
Stopping containerevent sequence: normal termination.
OOMKilled reason, not on exit code 137, or every rolling update
produces false OOM alerts.3
Find the source of the OOM
The namespace where you notice symptoms is often not where the OOM started.
-
Find Pods with the highest restart counts cluster-wide. Safe; read-only.
-
For any suspect Pod, read its last termination state. Safe; read-only.
CrashLoopBackOff, so each restart needs
more memory than the last. Raise the limit well above the current value, let
the backlog clear, then right-size it.A subprocess inside a multi-process Pod can be OOM-killed without the kubelet
recording an OOMKilled event. If a Pod misbehaves with no OOMKilled event,
check kernel logs if you have Node log access.How to fix OOM problems
For container OOM (CONSTRAINT_MEMCG, MemoryPressure=False):
- Raise the container’s memory limit, or reduce what the workload allocates. See What are the recommended resource requests and limits for GPU Pods?.
- Watch for a missing unit suffix:
memory: 16means 16 bytes,memory: 16Gmeans 16 gigabytes, andmemory: 16Gimeans about 17.2 gigabytes.
CONSTRAINT_NONE, MemoryPressure=True):
- Reduce the total memory your Pods use on the Node, or spread the workload across more Nodes.
- Set requests to match actual usage so the scheduler doesn’t overpack the Node.
- If there’s no apparent workload cause, contact support with the Node name and the OOM events.
When a CoreWeave-managed Pod is the one killed
If the killed process isnode_exporter, the CoreWeave-managed monitoring agent
in the cw-exporters namespace, the cause is usually still a workload that creates an
unusually large number of processes, sockets, mounts, or file descriptors, which
node_exporter enumerates on every scrape. This case looks like:
- The killed process is
node_exporter, not one of your processes. - Multiple Nodes are affected at about the same time, matching where the workload runs.
- The Nodes show no
MemoryPressurecondition.
When the memory graphs mislead you
- A Pod that OOMs within seconds of starting shows near-zero memory in
dashboards, because it crashes before the metrics scrape records peak usage.
When a Pod is in
CrashLoopBackOff, trust the termination reason, the restart count, and the kernel log, not the memory graph. - In Grafana memory graphs, “used” memory includes page cache, which is not memory pressure: the kernel reclaims cache freely. Judge Node-level pressure by available memory, not used memory: a Node with hundreds of GB available isn’t memory-constrained even if the used graph looks full.
- Available memory counts
/dev/shmand other tmpfs memory as reclaimable cache, but the kernel can’t reclaim it. Workloads that use/dev/shmheavily can OOM a Node whose available memory looks adequate. To check, comparenode_memory_Shmem_bytestonode_memory_Cached_bytesin Grafana: if shmem tracks cached, most of the “cache” can’t be reclaimed. Reduce shared-memory usage or leave real headroom.
Common confusions
- Exit code 137 isn’t always an OOM. Read the termination reason.
- “Connection refused” between your Pods can be an OOM symptom, not a network
problem: an OOM-killed server process is no longer listening on its port.
Check
kubectl get events --field-selector involvedObject.name=[NODE-NAME]forSystemOOMevents before investigating networking. FailedKillPodevents withDeadlineExceededare a different failure: the container runtime couldn’t stop the container. If a Pod stays stuck inTerminating, contact support with the Pod and Node names.- A single Node that consistently OOMs on workloads that succeed on identical Nodes may have a hardware state issue. Contact support with the Node name.
Related pages
- Pod Inspector dashboard: OOM history and memory panels.
- Node cordoning: why CoreWeave cordons Nodes and how to check the cordon reason.