Your Deployment rolled out, but the Pod never becomes ready and the restart counter keeps climbing (Kubernetes 1.30 and later; output checked against the v1.37 documentation):
$ kubectl get pods -n shop
NAME READY STATUS RESTARTS AGE
api-7c9d8b6f5-x2k4q 0/1 CrashLoopBackOff 6 (2m1s ago) 8m40s
$ kubectl describe pod api-7c9d8b6f5-x2k4q -n shop
State: Waiting
Reason: CrashLoopBackOff
Last State: Terminated
Reason: Error
Exit Code: 1
Restart Count: 6
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning BackOff 15s (x34 over 8m) kubelet Back-off restarting failed container api in pod api-7c9d8b6f5-x2k4q_shop(...)CrashLoopBackOff is not a failure reason by itself. It means the container started, exited, was restarted, exited again, and the kubelet is now waiting an increasing delay before the next attempt. The real error is in the run that just ended: its exit code, its logs and the events around it.
Quick fix checklist
- Run
kubectl describe pod <pod>and readLast State: Terminated(Reason and Exit Code) plus the Events at the bottom. - Run
kubectl logs <pod> --previous(add-c <container>for multi-container Pods) to see output from the crashed run. - Exit code 1 or 2: application error, usually a missing env var, bad config file or unreachable dependency at startup.
- Exit code 0 with Reason
Completed: the process finished; a Deployment expects it to run forever. StartErrororexecutable file not found in $PATH:commandorargsin the manifest is wrong.- Exit code 137 or 143 with
failed liveness probeevents: the probe is killing the app; add astartupProbe. - Exit code 137 with Reason
OOMKilled: the memory limit is too low; see the OOMKilled guide linked below.
Before you start
You need kubectl access to the namespace with permission to get, describe and read logs for Pods (the view ClusterRole is enough for everything except kubectl debug). Commands use kubectl 1.30 or newer. Know which controller owns the Pod (Deployment, StatefulSet, Job), because you fix the template there, not on the Pod.
Resist the urge to delete the Pod first. A new Pod from the same template crashes the same way, and you lose the Last State and the previous container’s logs that tell you why.
Why it happens
Every Pod has a restartPolicy. Deployments, StatefulSets and DaemonSets force Always, so whenever the main process exits, for any reason and with any exit code, the kubelet restarts the container on the same node. To avoid hammering the node with a container that dies instantly, the kubelet applies an exponential backoff: 10 seconds, then 20, 40, 80, 160, and then a cap of 300 seconds (5 minutes). While it waits, the container’s state is Waiting with reason CrashLoopBackOff. If a container runs for 10 minutes without problems, the backoff resets and the next crash is treated as the first.
That timing explains a common confusion. After an hour of crashing, the Pod spends almost all of its time in CrashLoopBackOff and only a few seconds running, so kubectl get pods nearly always shows the backoff state rather than Error or Running.
Note
Newer kubelets can tune this.
ReduceDefaultCrashLoopBackOffDecay(alpha since 1.33, off by default) changes the defaults to 1s initial and 60s maximum.KubeletCrashLoopBackOffMax(alpha in 1.32, beta and on by default since 1.35) lets an administrator lower the per-node cap withcrashLoopBackOff.maxContainerRestartPeriodin the kubelet configuration. Neither changes the diagnosis.
The process exits for one of a handful of reasons:
- The application fails at startup: a required environment variable is empty, a config file is missing, the database URL is wrong, or a migration fails. Exit code is usually 1.
- The command is wrong:
commandreplaces the imageENTRYPOINTandargsreplacesCMD. A typo, a shell built-in used without a shell, or a binary that does not exist in the image all fail before your code runs. - The process has nothing to do: a container whose command finishes (a script,
echo, a CLI) exits 0. UnderrestartPolicy: Alwaysthat is still a crash loop. - Kubernetes kills it: a failing liveness probe makes the kubelet stop the container, and the memory limit makes the kernel kill it (
OOMKilled).
Step-by-step walkthrough
Step 1: Read the last termination, not the current state
describe is long; pull out just the fields that matter:
kubectl get pod api-7c9d8b6f5-x2k4q -n shop \
-o jsonpath='{range .status.containerStatuses[*]}{.name}{" reason="}{.lastState.terminated.reason}{" exit="}{.lastState.terminated.exitCode}{" restarts="}{.restartCount}{"\n"}{end}'api reason=Error exit=1 restarts=6Map the exit code before guessing. Codes above 128 mean a signal: 128 plus the signal number.
| Exit code | Usual meaning in a Pod |
|---|---|
| 0 | Process finished normally; Reason is Completed |
| 1, 2 | Application or shell reported an error |
| 126 | Command found but not executable (permissions) |
| 127 | Command not found, when a shell runs it |
128 with Reason StartError |
Runtime could not start the process at all |
| 137 | SIGKILL (9): OOM kill, or forced kill after a grace period |
| 139 | SIGSEGV (11): native crash |
| 143 | SIGTERM (15): the process was asked to stop and did |
You can confirm the signal arithmetic in any shell: sh -c 'kill -9 $$'; echo $? prints 137, and the same with -TERM prints 143.
Step 2: Read the logs of the run that died
kubectl logs api-7c9d8b6f5-x2k4q -n shop --previous
kubectl logs api-7c9d8b6f5-x2k4q -n shop -c api --previous --tail=50--previous (short form -p) asks the kubelet for the last terminated instance of the container. Without it you may get the log of a brand-new instance that has printed nothing yet. For an init container, pass its name with -c; init containers crash-loop too, and the Pod shows Init:CrashLoopBackOff.
Typical startup failures look like this:
Error: environment variable DATABASE_URL is requiredIf the previous log is empty, the process died before writing anything. That points at the command, the image, or the runtime, so go to the events.
Step 3: Check events for probe kills and start errors
kubectl get events -n shop --field-selector involvedObject.name=api-7c9d8b6f5-x2k4q --sort-by=.lastTimestampTwo patterns are worth recognising. A wrong command produces a Last State with Reason StartError and a message containing executable file not found in $PATH (the full text comes from the container runtime and varies between containerd and CRI-O). A liveness kill produces these events:
Warning Unhealthy kubelet Liveness probe failed: Get "http://10.244.1.7:8080/healthz": dial tcp 10.244.1.7:8080: connect: connection refused
Normal Killing kubelet Container api failed liveness probe, will be restartedAfter a liveness kill the exit code is 143 if the app handled SIGTERM and exited, or 137 if it ignored SIGTERM and was killed when the grace period ran out.
Step 4: Reproduce interactively when the container dies too fast
If you cannot tell what the command does inside the image, start a copy of the Pod with a shell instead of the real command:
kubectl debug pod/api-7c9d8b6f5-x2k4q -n shop -it --copy-to=api-debug --container=api -- shInside, run the original entrypoint by hand, check env, and look for the files the app expects. Delete api-debug when finished. This needs permission to create Pods, and the image must contain a shell.
Step 5: Fix the template, then watch one rollout
Fix the Deployment, not the Pod. For a wrong command, remember the mapping:
containers:
- name: api
image: ghcr.io/acme/shop-api:1.8.0
command: ["node"] # replaces ENTRYPOINT
args: ["dist/server.js", "--port=8080"] # replaces CMDFor missing configuration, add the env var or ConfigMap reference. For a process that legitimately finishes, use a Job or CronJob instead of a Deployment.
Worked scenario
A Spring Boot service takes about 70 seconds to start on a busy node. The Deployment has only a liveness probe:
livenessProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
initialDelaySeconds: 10
periodSeconds: 10
failureThreshold: 3The previous logs show a normal startup cut off mid-way, with no stack trace. describe shows Exit Code: 143 and repeated Container api failed liveness probe, will be restarted events. The arithmetic gives it away: the first probe runs at 10s, three consecutive failures at 10-second intervals end at about 30 to 40 seconds, long before the app listens. The JVM receives SIGTERM, shuts down cleanly (hence 143), and the kubelet restarts it into the same trap with an ever longer delay.
The fix gives startup its own budget and keeps liveness strict for the running app:
startupProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
periodSeconds: 5
failureThreshold: 30 # up to 150s to start
livenessProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
periodSeconds: 10
failureThreshold: 3
readinessProbe:
httpGet:
path: /actuator/health/readiness
port: 8080
periodSeconds: 5While the startup probe has not succeeded, the kubelet does not run liveness or readiness checks. Once it succeeds, liveness takes over with its normal 30-second tolerance.
Common mistake
The tempting fix is to delete the Pod or run kubectl rollout restart. Neither changes the template, so the new Pod crashes for the same reason, and deleting the old one throws away the previous logs.
The second tempting fix, for probe kills, is removing the liveness probe or setting initialDelaySeconds: 300. Removing it means a deadlocked process is never restarted. A five-minute initial delay means that, after every future restart, a genuinely hung container also gets five minutes of free time. A startupProbe solves only the slow start.
A third trap: wrapping the command in sh -c "... || sleep infinity" to stop the loop. The Pod looks Running, the app is not, and every alert that watches restarts goes quiet.
Verify the behavior
After the fixed rollout, the restart count should stay at zero and the rollout should complete:
kubectl rollout status deployment/api -n shop
kubectl get pods -n shop -l app=api -wdeployment "api" successfully rolled out
api-5b8f9d7c6-p4w2n 1/1 Running 0 2mFor a probe fix, also check that no Unhealthy events appear during startup: kubectl get events -n shop --field-selector reason=Unhealthy. Restart counts that stay at zero for more than 10 minutes confirm the backoff will not resume.
Interview exercise
A Pod shows CrashLoopBackOff with 12 restarts. kubectl logs --previous shows a healthy startup that stops abruptly with no error. describe reports Exit Code: 143. What is happening, and what would you change?
Answer and reasoning
Exit code 143 is 128 + 15, so the process received SIGTERM and exited cleanly. Applications rarely send that to themselves, so something outside asked the container to stop. In a Pod that is not being deleted, the usual sender is the kubelet acting on a failed liveness probe; I would confirm with Unhealthy and failed liveness probe, will be restarted events. The logs look healthy because the app was healthy, just not yet listening when the probe gave up. I would add a startupProbe sized to the worst observed startup time (failureThreshold × periodSeconds), keep the liveness probe tight, and make sure liveness checks only the process itself, not downstream dependencies, so a database outage cannot cause a restart storm. If the code were 137 instead, I would check lastState.terminated.reason to separate an OOMKilled from a SIGTERM that the app ignored.
Continue learning
- Kubernetes interview questions and Kubernetes MCQs
- Kubernetes Startup, Readiness and Liveness Probes
- Kubernetes Troubleshooting From Symptoms to Evidence
- Kubernetes OOMKilled and exit code 137 and CreateContainerConfigError
- An image built for the wrong CPU crash-loops too: exec format error in Docker
- Official reference: Pod lifecycle and container restarts and Debug running Pods