A Deployment’s Pods restart every few minutes. kubectl get pods shows the restart counter climbing, and describe tells you why:
$ kubectl get pods -n shop
NAME READY STATUS RESTARTS AGE
api-7c9d8b6f5-x2k4q 0/1 Running 5 (40s ago) 6m10s
$ kubectl describe pod api-7c9d8b6f5-x2k4q -n shop
Last State: Terminated
Reason: OOMKilled
Exit Code: 137
Limits:
memory: 256Mi
Requests:
memory: 256MiExit code 137 is 128 + 9: the container’s process was killed with SIGKILL, and there is only one common reason for that inside Kubernetes.
Quick fix checklist
- Confirm the reason:
kubectl get pod <pod> -o jsonpath='{.status.containerStatuses[*].lastState.terminated.reason}'printsOOMKilled(notErrororCompleted). - Check whether it is the container or the node: a container limit kill sets
lastState.terminated.reason: OOMKilled, while node pressure showsstatus.reason: Evictedandstatus.messagenamingMemoryPressure. - Find the limit actually in force (
Limits: memory:indescribe); remember a container with no limit can still be killed by node pressure. - Compare current usage with
kubectl top pod <pod>(or the container’s own metrics) against the limit. - Raise
limits.memoryabove the true peak and setrequests.memoryto the typical steady state. - For JVM and Node apps, cap the heap below the container limit so the runtime does not over-commit.
- Then investigate the leak or the request that spiked; a bigger limit only delays the crash.
Before you start
You need kubectl access to the namespace, the Pod’s last termination state, and metrics (metrics-server for kubectl top, or your monitoring). The examples use Kubernetes 1.30 and later with cgroup v2; the field names are stable across recent versions. On cgroup v1 the accounting is slightly different but OOMKilled and exit code 137 mean the same thing.
Why it happens
Every container runs inside a cgroup with a memory limit. The Linux kernel tracks its resident memory (RSS plus page cache); when the group tries to exceed its limit and there is nothing to reclaim, the kernel OOM killer picks a process in the group and sends SIGKILL. A process killed by signal 9 has no chance to log or shut down cleanly, and the shell’s convention 128 + signal gives exit code 137. Kubernetes reads that exit and records Reason: OOMKilled.
Two different memory limits can trigger it:
- The container limit.
resources.limits.memorysets the cgroup limit. This is the common case, and it is a per-container problem: the node may have plenty of free memory. Because containers normally run without swap, the limit is a hard wall. - Node pressure. Every node has an
--eviction-hardthreshold (oftenmemory.available<100Mi). When node memory runs low, the kubelet evicts Pods, and the kernel may OOM-kill processes first. This looks different: the Pod’sstatus.reasonisEvicted,status.messagementions memory pressure, and several Pods on the same node fail together.
The cgroup limit is enforced on the whole group, not just your process. The page cache a container builds up (file reads, writes not yet flushed) counts toward the limit, and the kernel usually reclaims that before killing anything; a kill means live, unreclaimable memory genuinely exceeded the ceiling.
Step-by-step walkthrough
Step 1: Read the termination state
kubectl describe pod <pod> | sed -n '/Last State/,/Restart Count/p'
kubectl get pod <pod> -o jsonpath='{.status.containerStatuses[*].lastState.terminated}'reason: OOMKilled with exitCode: 137 is the container-limit case. If instead you see reason: Error with exitCode: 137, the process was killed by something else (a SIGKILL from a health check or a manual kill) and you should read the events.
Step 2: Confirm it was the container, not the node
kubectl get pod <pod> -o jsonpath='{.status.reason}{"\n"}{.status.message}'
kubectl get events -n shop --field-selector involvedObject.name=<pod>A status.reason of Evicted and a message about MemoryPressure points at the node, not the limit. In that case the fix is about requests and node capacity, not one container’s limit.
Step 3: Measure the real working set
kubectl top pod <pod> --containers
kubectl get pod <pod> -o jsonpath='{.spec.containers[*].resources}'kubectl top shows current usage; a spike at kill time is what matters, so use a dashboard or container_memory_working_set_bytes from cAdvisor if you have it. Working set (not RSS alone) is what the kernel compares against the cgroup limit, because it includes the memory the group is actively using and cannot easily reclaim.
Step 4: Size requests and limits
resources:
requests:
memory: "256Mi"
limits:
memory: "384Mi"Set requests.memory from the typical steady state so the scheduler places the Pod correctly, and limits.memory above the observed peak with headroom for garbage-collection lag. Avoid a limit equal to the request (a Guaranteed container) unless you have measured it, and never set a limit far below what the app needs “to make it fit”, since that just makes the kills more frequent.
Step 5: Fix the runtime and the leak
- JVM: the container-aware heap defaults are usually fine, but an explicit
-Xmxabove the limit (or an unordered-Xms) causes OOM. Keep total JVM memory (heap plus metaspace, stacks, direct buffers) belowlimits.memory. - Node.js: set
--max-old-space-sizebelow the container limit; the default heap is sized from total machine memory, not the cgroup. - Leak or spike: if usage climbs steadily until the kill, it is a leak; if it spikes on a specific request, it is load. Either way, raise the limit only as a stopgap and fix the cause with a heap dump or a load test.
Worked scenario
An API Pod is killed every few minutes at exactly the same restart cadence. describe shows OOMKilled, exit code 137, limit 256Mi. kubectl top pod sits at 240Mi just before each kill, and a memory graph shows a staircase: 180Mi after start, rising with each batch job, then a kill and a reset. The staircase means a leak (objects retained across requests), not a too-small limit. Raising the limit to 512Mi pushes the kill from 4 minutes to 9 minutes. The durable fix is a heap profile that finds the growing map, plus a right-sized limit once the leak is gone.
Common mistake
Treating exit code 137 as “just restart it” or “give it more memory” and moving on. A 137 with OOMKilled is real memory exhaustion; if usage is climbing it will exhaust any limit eventually, and if it spikes it will do so again under the same load. The restart also hides the symptom in dashboards that only chart the latest container. Fix the memory use, then size the limit from data.
Verify the behavior
After resizing (and fixing the cause), the restart count should stop climbing and usage should stay under the limit:
kubectl get pods -n shop -l app=api -w
kubectl top pod -n shop -l app=api --containers
kubectl get pod -n shop -l app=api -o jsonpath='{.items[*].status.containerStatuses[*].restartCount}'api-5b8f9d7c6-p4w2n 1/1 Running 0 9mNo new OOMKilled reason appears in lastState, and restartCount stays constant for longer than the worst previous interval.
Interview exercise
“A Pod shows OOMKilled with exit code 137, but kubectl top says it uses only 200Mi against a 256Mi limit. How can it be OOM-killed below its limit?”
Answer and reasoning
Several effects can produce a kill that a point-in-time kubectl top reading does not show. First, kubectl top samples at a moment; the kill happens at the peak, which may be a short burst between scrapes. Second, the kernel compares the cgroup’s working set (RSS plus active page cache) against the limit, not the process RSS that top-style tools often display. Third, if the Pod has multiple containers or init containers, memory is attributed per container and one may exceed its own limit. Fourth, a 137 with reason: Error rather than OOMKilled is a SIGKILL from something else, not memory at all. I would confirm the reason field, look at a time series of working-set bytes around the kill, and check every container’s limit before changing anything. The fix follows the evidence: a spike means right-sizing or a burst buffer, a staircase means a leak.
Continue learning
- Kubernetes interview questions and Kubernetes MCQs
- Kubernetes resource requests and limits
- Kubernetes CrashLoopBackOff
- Kubernetes resource quotas
- Kubernetes node placement
- Docker resource limits, for the same cgroup limits on a single host
- Official reference: Managing resources for containers, Pod eviction and Assign memory resources