Kubernetes OOMKilled (Exit Code 137): Causes and Fixes

kubernetes oomkilled container with exit code 137 compared with node OOM kill and eviction

Kubernetes OOMKilled with exit code 137 means the Linux kernel killed your container’s process for using more memory than it was allowed. Sometimes that’s the container’s own memory limit. Sometimes the whole node ran out of memory and your container lost the lottery. The status looks the same, but the fix is completely different, so the first job is working out which one you hit.

This guide shows how to confirm an OOMKilled container, tell a limit kill from a node kill, and fix the common causes, including the JVM default that trips up most Java workloads. If you need a refresher on how requests and limits work, see our Kubernetes resource requests and limits guide, or start from the complete Kubernetes guide.

What Does OOMKilled Exit Code 137 Mean in Kubernetes?

Exit code 137 is 128 + 9. Per the Bash manual, when a process dies from a fatal signal N, the exit status is 128+N. Signal 9 is SIGKILL, which a process can’t catch or clean up after.

OOMKilled is the reason Kubernetes records when that SIGKILL came from the kernel’s out-of-memory killer. The Kubernetes memory docs show exactly what you’ll see in the pod status:

lastState:
  terminated:
    exitCode: 137
    reason: OOMKilled

⚠️ Note: Exit code 137 on its own doesn’t prove an OOM kill. Anything that sends SIGKILL produces 137, including a pod that ignored SIGTERM for longer than its grace period (see our Kubernetes graceful shutdown guide). Only trust it as memory-related when the reason says OOMKilled.

How to Confirm a Kubernetes OOMKilled Container

Start with the pod. A container that keeps getting OOMKilled usually shows a climbing restart count and eventually CrashLoopBackOff:

kubectl get pod my-app-7d9f8 -n prod
kubectl describe pod my-app-7d9f8 -n prod

In the describe output, look under the container’s Last State. You want Terminated, Reason: OOMKilled, Exit Code: 137. To pull just that field across all containers in the pod:

kubectl get pod my-app-7d9f8 -n prod \
  -o jsonpath='{range .status.containerStatuses[*]}{.name}{"\t"}{.lastState.terminated.reason}{"\t"}{.lastState.terminated.exitCode}{"\n"}{end}'

Then check how much memory the container actually uses against its limit. This needs metrics-server installed in the cluster:

kubectl top pod my-app-7d9f8 -n prod --containers
kubectl get pod my-app-7d9f8 -n prod -o jsonpath='{.spec.containers[*].resources}'

If usage sits close to the limit just before each restart, you’ve found your answer.

Container Limit Kill vs Node Out-of-Memory Kill

There are two different ways a container ends up OOMKilled, plus one lookalike that isn’t an OOM kill at all.

What happened What triggers it What you see
Container limit kill Container uses more than its own limits.memory OOMKilled, 137, container restarts
Node OOM kill The whole node runs out of memory before the kubelet can react OOMKilled, 137, often a container that was under its limit
Eviction (not an OOM kill) Node memory drops below the kubelet’s eviction threshold Pod status Evicted, pod rescheduled elsewhere

Container limit kills are the common case. The kernel enforces limits.memory through the container’s cgroup. The resource management docs point out a subtlety: memory limits are enforced reactively. A container can briefly exceed its limit, and the kill happens when the kernel detects memory pressure.

Node OOM kills happen when the node itself runs out of memory before the kubelet can evict anything. The kernel’s OOM killer then picks a victim across the whole node, and the kubelet biases that choice by Quality of Service class through oom_score_adj:

QoS class oom_score_adj Killed
Guaranteed -997 Last
Burstable Between 2 and 999, based on the memory request In between
BestEffort 1000 First

This is why a pod with no memory requests or limits (BestEffort) gets killed first on a busy node, even if it’s small. To check the node side, look for kernel OOM events. The Kubernetes docs example shows them as OOMKilling warnings on the node, though whether they appear as events depends on your cluster’s node monitoring:

kubectl describe node <node-name> | grep -i -A2 oom
kubectl get events -A --field-selector reason=OOMKilling

Evictions look different. The pod status says Evicted, not OOMKilled, and the pod is replaced rather than restarted in place. If that’s what you see, the fix is node capacity and requests, not container limits.

Why Java Containers Get OOMKilled (and Why They Sometimes Don’t Use Their Memory)

Java is the most common OOMKilled search for a reason. A modern JVM sizes its heap from the container’s memory limit, but by default it only takes a quarter of it. The Java documentation lists the -XX:MaxRAMPercentage default as 25 percent.

That causes two opposite problems:

  • Heap too small. With a 2 GiB limit, the default max heap is about 512 MiB. The app throws java.lang.OutOfMemoryError while the container still has plenty of memory. That’s a JVM error, not an OOMKilled.
  • Heap too big. Someone sets -Xmx or MaxRAMPercentage=100 to “use all the memory.” But the JVM also needs memory outside the heap: metaspace, thread stacks, the JIT code cache, direct buffers. Heap plus everything else exceeds the limit, and the kernel kills the container with 137.

A common middle ground is to leave headroom for non-heap memory:

env:
  - name: JAVA_TOOL_OPTIONS
    value: "-XX:MaxRAMPercentage=75"
resources:
  requests:
    memory: "2Gi"
  limits:
    memory: "2Gi"

Tune the percentage from real measurements. Apps with many threads or large direct buffers need more non-heap headroom than 25 percent.

Other Causes of Kubernetes OOMKilled Errors

Memory-backed emptyDir volumes. If you mount an emptyDir with medium: Memory, files written there count against the container’s memory limit. The docs note the kubelet tracks tmpfs emptyDir volumes as container memory use, not local storage. A job that writes a large temp file to a RAM-backed volume gets OOMKilled even though the app itself is small.

A real memory leak. If memory climbs steadily until each kill, raising the limit only delays the next restart. Graph the container’s memory over a few restarts. A sawtooth that always climbs to the limit is a leak, and the fix is in the code.

A limit copied from another service. Limits often get set once from a template and never measured. Check actual peak usage under load before assuming the app is misbehaving.

Sidecars. Each container has its own limit. An Envoy or logging sidecar can be the one getting OOMKilled while your app container is fine. The jsonpath command above shows which container it was.

How to Fix a Kubernetes OOMKilled Container

Match the fix to the cause:

  1. Limit too low for real usage: raise limits.memory in the Deployment, based on measured peak usage plus headroom.
  2. Node OOM kills: set memory requests that reflect real usage, so the scheduler doesn’t overpack nodes. For critical workloads, set requests equal to limits so the pod is Guaranteed and killed last.
  3. Java: set MaxRAMPercentage deliberately instead of relying on the 25 percent default or -Xmx guesses.
  4. Leak: fix the code. Raising the limit buys time, not a solution.

Raising Memory Without Recreating the Pod

In-place pod resize has been stable since Kubernetes v1.35, according to the resize documentation. You can raise a running container’s memory through the pod’s resize subresource. It requires kubectl v1.32 or newer.

This changes one running pod only. It is useful for buying time during an incident:

kubectl patch pod my-app-7d9f8 -n prod --subresource resize --patch \
  '{"spec":{"containers":[{"name":"app","resources":{"requests":{"memory":"1Gi"},"limits":{"memory":"1Gi"}}}]}}'

⚠️ Important: For pods managed by a Deployment, this patch is lost the next time the pod is replaced, because new pods come from the Deployment template. Update the Deployment too, or the OOMKilled restarts come back on the next rollout.

Keep a few resize rules in mind. The pod’s QoS class can’t change through a resize, so a Guaranteed pod must keep requests equal to limits. Lowering a memory limit is best-effort: if current usage is above the new limit, the kubelet skips the resize and leaves it “In Progress.” Windows pods don’t support in-place resize.

Right-sizing memory also affects your bill. Our Kubernetes cost optimization guide covers finding over-provisioned workloads without starving them.

Frequently Asked Questions

Q: What does exit code 137 mean in Kubernetes?
A: Exit code 137 means the process was killed by signal 9 (SIGKILL), because the exit status for a fatal signal is 128 plus the signal number. When the reason is OOMKilled, the kernel’s out-of-memory killer sent it. Without that reason, 137 can also come from a pod that ignored SIGTERM past its grace period.

Q: How do I check why a Kubernetes pod was OOMKilled?
A: Run kubectl describe pod <name> and look at the container’s Last State for Reason: OOMKilled and Exit Code: 137. Then compare kubectl top pod <name> --containers with the container’s memory limit. If usage was under the limit, check the node for kernel OOM events.

Q: Is OOMKilled the same as a pod being evicted?
A: No. OOMKilled means the kernel killed a container, and the kubelet restarts it in the same pod. Eviction means the kubelet removed the whole pod because the node was short on memory, and the pod status shows Evicted.

Q: Can I increase a pod’s memory limit without restarting it?
A: Yes, on Kubernetes v1.35 or later, with kubectl v1.32 or later. Patch the pod’s resize subresource. But for Deployment-managed pods, also update the Deployment, or the change is lost when the pod is replaced.

Q: Why does my Java app get OOMKilled when the heap is smaller than the limit?
A: The JVM uses memory outside the heap for metaspace, thread stacks, the code cache and direct buffers. Heap plus non-heap can exceed the container limit. Leave headroom, for example -XX:MaxRAMPercentage=75 instead of 100.

Quick Summary:
– Exit code 137 is 128 + 9 (SIGKILL); only the OOMKilled reason ties it to memory
– A container limit kill, a node OOM kill and an eviction look similar but need different fixes
– On node OOM, BestEffort pods die first (oom_score_adj 1000) and Guaranteed pods last (-997)
– The JVM defaults to a max heap of 25 percent of the container limit, and non-heap memory still counts
– In-place resize (stable since v1.35) raises memory on a running pod, but update the Deployment too

Next time a pod shows Kubernetes OOMKilled, run kubectl describe pod and kubectl top pod --containers side by side before touching any limits. Those two outputs tell you which of the three cases you’re in.

Related guides

Leave a Reply