Ch. 14 · Kubernetes

Kubernetes CrashLoopBackOff: Back-off restarting failed container

Fix Kubernetes CrashLoopBackOff: read the exit code, check kubectl logs --previous, fix command, config and probes, and understand backoff timing.

~9 min readintermediateupdated Oct 4, 2026

Your Deployment rolled out, but the Pod never becomes ready and the restart counter keeps climbing (Kubernetes 1.30 and later; output checked against the v1.37 documentation):

$ kubectl get pods -n shop
NAME                   READY   STATUS             RESTARTS      AGE
api-7c9d8b6f5-x2k4q    0/1     CrashLoopBackOff   6 (2m1s ago)  8m40s

$ kubectl describe pod api-7c9d8b6f5-x2k4q -n shop
    State:          Waiting
      Reason:       CrashLoopBackOff
    Last State:     Terminated
      Reason:       Error
      Exit Code:    1
    Restart Count:  6
Events:
  Type     Reason   Age                 From     Message
  ----     ------   ----                ----     -------
  Warning  BackOff  15s (x34 over 8m)   kubelet  Back-off restarting failed container api in pod api-7c9d8b6f5-x2k4q_shop(...)
Text

CrashLoopBackOff is not a failure reason by itself. It means the container started, exited, was restarted, exited again, and the kubelet is now waiting an increasing delay before the next attempt. The real error is in the run that just ended: its exit code, its logs and the events around it.

Quick fix checklist

  • Run kubectl describe pod <pod> and read Last State: Terminated (Reason and Exit Code) plus the Events at the bottom.
  • Run kubectl logs <pod> --previous (add -c <container> for multi-container Pods) to see output from the crashed run.
  • Exit code 1 or 2: application error, usually a missing env var, bad config file or unreachable dependency at startup.
  • Exit code 0 with Reason Completed: the process finished; a Deployment expects it to run forever.
  • StartError or executable file not found in $PATH: command or args in the manifest is wrong.
  • Exit code 137 or 143 with failed liveness probe events: the probe is killing the app; add a startupProbe.
  • Exit code 137 with Reason OOMKilled: the memory limit is too low; see the OOMKilled guide linked below.

Before you start

You need kubectl access to the namespace with permission to get, describe and read logs for Pods (the view ClusterRole is enough for everything except kubectl debug). Commands use kubectl 1.30 or newer. Know which controller owns the Pod (Deployment, StatefulSet, Job), because you fix the template there, not on the Pod.

Resist the urge to delete the Pod first. A new Pod from the same template crashes the same way, and you lose the Last State and the previous container’s logs that tell you why.

Why it happens

Every Pod has a restartPolicy. Deployments, StatefulSets and DaemonSets force Always, so whenever the main process exits, for any reason and with any exit code, the kubelet restarts the container on the same node. To avoid hammering the node with a container that dies instantly, the kubelet applies an exponential backoff: 10 seconds, then 20, 40, 80, 160, and then a cap of 300 seconds (5 minutes). While it waits, the container’s state is Waiting with reason CrashLoopBackOff. If a container runs for 10 minutes without problems, the backoff resets and the next crash is treated as the first.

That timing explains a common confusion. After an hour of crashing, the Pod spends almost all of its time in CrashLoopBackOff and only a few seconds running, so kubectl get pods nearly always shows the backoff state rather than Error or Running.

Note

Newer kubelets can tune this. ReduceDefaultCrashLoopBackOffDecay (alpha since 1.33, off by default) changes the defaults to 1s initial and 60s maximum. KubeletCrashLoopBackOffMax (alpha in 1.32, beta and on by default since 1.35) lets an administrator lower the per-node cap with crashLoopBackOff.maxContainerRestartPeriod in the kubelet configuration. Neither changes the diagnosis.

The process exits for one of a handful of reasons:

  1. The application fails at startup: a required environment variable is empty, a config file is missing, the database URL is wrong, or a migration fails. Exit code is usually 1.
  2. The command is wrong: command replaces the image ENTRYPOINT and args replaces CMD. A typo, a shell built-in used without a shell, or a binary that does not exist in the image all fail before your code runs.
  3. The process has nothing to do: a container whose command finishes (a script, echo, a CLI) exits 0. Under restartPolicy: Always that is still a crash loop.
  4. Kubernetes kills it: a failing liveness probe makes the kubelet stop the container, and the memory limit makes the kernel kill it (OOMKilled).

Step-by-step walkthrough

Step 1: Read the last termination, not the current state

describe is long; pull out just the fields that matter:

kubectl get pod api-7c9d8b6f5-x2k4q -n shop \
  -o jsonpath='{range .status.containerStatuses[*]}{.name}{"  reason="}{.lastState.terminated.reason}{"  exit="}{.lastState.terminated.exitCode}{"  restarts="}{.restartCount}{"\n"}{end}'
Terminal
api  reason=Error  exit=1  restarts=6
Text

Map the exit code before guessing. Codes above 128 mean a signal: 128 plus the signal number.

Exit code Usual meaning in a Pod
0 Process finished normally; Reason is Completed
1, 2 Application or shell reported an error
126 Command found but not executable (permissions)
127 Command not found, when a shell runs it
128 with Reason StartError Runtime could not start the process at all
137 SIGKILL (9): OOM kill, or forced kill after a grace period
139 SIGSEGV (11): native crash
143 SIGTERM (15): the process was asked to stop and did

You can confirm the signal arithmetic in any shell: sh -c 'kill -9 $$'; echo $? prints 137, and the same with -TERM prints 143.

Step 2: Read the logs of the run that died

kubectl logs api-7c9d8b6f5-x2k4q -n shop --previous
kubectl logs api-7c9d8b6f5-x2k4q -n shop -c api --previous --tail=50
Terminal

--previous (short form -p) asks the kubelet for the last terminated instance of the container. Without it you may get the log of a brand-new instance that has printed nothing yet. For an init container, pass its name with -c; init containers crash-loop too, and the Pod shows Init:CrashLoopBackOff.

Typical startup failures look like this:

Error: environment variable DATABASE_URL is required
Text

If the previous log is empty, the process died before writing anything. That points at the command, the image, or the runtime, so go to the events.

Step 3: Check events for probe kills and start errors

kubectl get events -n shop --field-selector involvedObject.name=api-7c9d8b6f5-x2k4q --sort-by=.lastTimestamp
Terminal

Two patterns are worth recognising. A wrong command produces a Last State with Reason StartError and a message containing executable file not found in $PATH (the full text comes from the container runtime and varies between containerd and CRI-O). A liveness kill produces these events:

Warning  Unhealthy  kubelet  Liveness probe failed: Get "http://10.244.1.7:8080/healthz": dial tcp 10.244.1.7:8080: connect: connection refused
Normal   Killing    kubelet  Container api failed liveness probe, will be restarted
Text

After a liveness kill the exit code is 143 if the app handled SIGTERM and exited, or 137 if it ignored SIGTERM and was killed when the grace period ran out.

Step 4: Reproduce interactively when the container dies too fast

If you cannot tell what the command does inside the image, start a copy of the Pod with a shell instead of the real command:

kubectl debug pod/api-7c9d8b6f5-x2k4q -n shop -it --copy-to=api-debug --container=api -- sh
Terminal

Inside, run the original entrypoint by hand, check env, and look for the files the app expects. Delete api-debug when finished. This needs permission to create Pods, and the image must contain a shell.

Step 5: Fix the template, then watch one rollout

Fix the Deployment, not the Pod. For a wrong command, remember the mapping:

containers:
  - name: api
    image: ghcr.io/acme/shop-api:1.8.0
    command: ["node"]                 # replaces ENTRYPOINT
    args: ["dist/server.js", "--port=8080"]  # replaces CMD
yaml

For missing configuration, add the env var or ConfigMap reference. For a process that legitimately finishes, use a Job or CronJob instead of a Deployment.

Worked scenario

A Spring Boot service takes about 70 seconds to start on a busy node. The Deployment has only a liveness probe:

livenessProbe:
  httpGet:
    path: /actuator/health/liveness
    port: 8080
  initialDelaySeconds: 10
  periodSeconds: 10
  failureThreshold: 3
yaml

The previous logs show a normal startup cut off mid-way, with no stack trace. describe shows Exit Code: 143 and repeated Container api failed liveness probe, will be restarted events. The arithmetic gives it away: the first probe runs at 10s, three consecutive failures at 10-second intervals end at about 30 to 40 seconds, long before the app listens. The JVM receives SIGTERM, shuts down cleanly (hence 143), and the kubelet restarts it into the same trap with an ever longer delay.

The fix gives startup its own budget and keeps liveness strict for the running app:

startupProbe:
  httpGet:
    path: /actuator/health/liveness
    port: 8080
  periodSeconds: 5
  failureThreshold: 30      # up to 150s to start
livenessProbe:
  httpGet:
    path: /actuator/health/liveness
    port: 8080
  periodSeconds: 10
  failureThreshold: 3
readinessProbe:
  httpGet:
    path: /actuator/health/readiness
    port: 8080
  periodSeconds: 5
yaml

While the startup probe has not succeeded, the kubelet does not run liveness or readiness checks. Once it succeeds, liveness takes over with its normal 30-second tolerance.

Common mistake

The tempting fix is to delete the Pod or run kubectl rollout restart. Neither changes the template, so the new Pod crashes for the same reason, and deleting the old one throws away the previous logs.

The second tempting fix, for probe kills, is removing the liveness probe or setting initialDelaySeconds: 300. Removing it means a deadlocked process is never restarted. A five-minute initial delay means that, after every future restart, a genuinely hung container also gets five minutes of free time. A startupProbe solves only the slow start.

A third trap: wrapping the command in sh -c "... || sleep infinity" to stop the loop. The Pod looks Running, the app is not, and every alert that watches restarts goes quiet.

Verify the behavior

After the fixed rollout, the restart count should stay at zero and the rollout should complete:

kubectl rollout status deployment/api -n shop
kubectl get pods -n shop -l app=api -w
Terminal
deployment "api" successfully rolled out
api-5b8f9d7c6-p4w2n   1/1   Running   0   2m
Text

For a probe fix, also check that no Unhealthy events appear during startup: kubectl get events -n shop --field-selector reason=Unhealthy. Restart counts that stay at zero for more than 10 minutes confirm the backoff will not resume.

Interview exercise

A Pod shows CrashLoopBackOff with 12 restarts. kubectl logs --previous shows a healthy startup that stops abruptly with no error. describe reports Exit Code: 143. What is happening, and what would you change?

Answer and reasoning

Exit code 143 is 128 + 15, so the process received SIGTERM and exited cleanly. Applications rarely send that to themselves, so something outside asked the container to stop. In a Pod that is not being deleted, the usual sender is the kubelet acting on a failed liveness probe; I would confirm with Unhealthy and failed liveness probe, will be restarted events. The logs look healthy because the app was healthy, just not yet listening when the probe gave up. I would add a startupProbe sized to the worst observed startup time (failureThreshold × periodSeconds), keep the liveness probe tight, and make sure liveness checks only the process itself, not downstream dependencies, so a database outage cannot cause a restart storm. If the code were 137 instead, I would check lastState.terminated.reason to separate an OOMKilled from a SIGTERM that the app ignored.

Continue learning

More in Kubernetes

esc