Ch. 14 · Kubernetes

Kubernetes ImagePullBackOff: ErrImagePull, Failed to pull image

Fix Kubernetes ImagePullBackOff and ErrImagePull: read the pull error, then fix tag typos, imagePullSecrets, CPU architecture or Docker Hub rate limits.

~8 min readbeginnerupdated Oct 4, 2026

The Pod was scheduled to a node, but its container never starts (Kubernetes 1.30 and later with containerd; checked against the v1.37 documentation):

$ kubectl get pods
NAME                    READY   STATUS             RESTARTS   AGE
web-6d8f7c9b5d-q7m2z    0/1     ImagePullBackOff   0          2m14s

$ kubectl describe pod web-6d8f7c9b5d-q7m2z
Events:
  Normal   Pulling  95s (x4 over 2m13s)  kubelet  Pulling image "nginx:1.27.99"
  Warning  Failed   94s (x4 over 2m12s)  kubelet  Failed to pull image "nginx:1.27.99": rpc error: code = NotFound desc = failed to pull and unpack image "docker.io/library/nginx:1.27.99": failed to resolve reference "docker.io/library/nginx:1.27.99": docker.io/library/nginx:1.27.99: not found
  Warning  Failed   94s (x4 over 2m12s)  kubelet  Error: ErrImagePull
  Normal   BackOff  8s (x9 over 2m11s)   kubelet  Back-off pulling image "nginx:1.27.99"
  Warning  Failed   8s (x9 over 2m11s)   kubelet  Error: ImagePullBackOff
Text

The kubelet asked the container runtime to download the image and the registry said no. ErrImagePull is that failure; ImagePullBackOff means the kubelet is waiting before it tries again. The useful part is the end of the Failed to pull image message. The rpc error: code = ... desc = wrapper and the exact phrasing come from the runtime and vary between containerd versions, CRI-O and kubelet releases, so match on the last clause.

Quick fix checklist

  • Copy the image reference from the event and check it character by character: registry host, repository path, tag.
  • not found: the tag or digest does not exist in that repository.
  • pull access denied, repository does not exist or may require authorization: wrong repository name, or a private repository with no credentials.
  • 401 Unauthorized or 403 Forbidden: credentials are missing, wrong, expired, or lack pull permission.
  • no match for platform in manifest: the image was not built for the node’s CPU architecture.
  • 429 Too Many Requests: Docker Hub rate limit; authenticate or use a mirror.
  • Check imagePullSecrets exists in the same namespace as the Pod and is listed on the Pod or its ServiceAccount.

Before you start

You need kubectl access to read Pods, events and (for the credential steps) Secrets in the namespace. It helps to have a machine where you can run docker pull, crane or skopeo against the same registry, so you can test the reference and credentials outside Kubernetes. Node shell access is not required; everything below works through the API.

Know where the image is supposed to live. “It works on my laptop” often means your laptop is logged in to a registry the cluster has never heard of, or already has the image cached.

Why it happens

When a Pod lands on a node, the kubelet checks the container’s imagePullPolicy. With Always it contacts the registry every time a container starts; with IfNotPresent it pulls only when the node lacks the image; with Never it never pulls. If you omit the field, the API server fills it in: Always for :latest or for no tag, IfNotPresent for any other tag or a digest.

A pull is several HTTP round trips: resolve the tag to a manifest (or an image index listing one manifest per platform), pick the manifest matching the node’s OS and architecture, then download layers. Each step has its own failure:

  • Resolve fails: the repository or tag does not exist, or the registry refuses anonymous access. Registries such as Docker Hub deliberately answer “repository does not exist or may require authorization” for both cases, so they do not reveal which private repositories exist. A typo in the repository name therefore looks exactly like an authentication problem.
  • Authentication fails: the kubelet sends credentials from the Pod’s imagePullSecrets, from the ServiceAccount’s imagePullSecrets, or from a node-level credential provider (common on EKS, GKE and AKS for their own registries). Wrong, expired or missing credentials give 401 or 403.
  • Platform selection fails: the index has no entry for linux/amd64 (or linux/arm64) and the runtime reports no match for platform in manifest.
  • The registry throttles you: Docker Hub returns 429 with You have reached your pull rate limit once you exceed the anonymous or account limit.

After each failure the kubelet backs off, doubling the delay up to a compiled-in limit of 300 seconds. That is why fixing the cause does not start the container instantly: the next attempt may be minutes away.

Note

A malformed reference fails before any network call. Uppercase letters in the repository, for example, give the status InvalidImageName with invalid reference format in the message, and no backoff loop.

Step-by-step walkthrough

Step 1: Get the exact reference and the error tail

kubectl get pod web-6d8f7c9b5d-q7m2z -o jsonpath='{.spec.containers[*].image}{"\n"}'
kubectl get events --field-selector involvedObject.name=web-6d8f7c9b5d-q7m2z,reason=Failed
Terminal

Note the fully qualified form in the error. nginx:1.27.99 becomes docker.io/library/nginx:1.27.99, which tells you a short name defaulted to Docker Hub. If you meant your company registry, the host is missing from the manifest.

Step 2: Test the same reference outside Kubernetes

docker manifest inspect nginx:1.27.99
docker buildx imagetools inspect ghcr.io/acme/shop-api:1.4.2
Terminal

If this fails from your laptop with the same error, the reference is wrong. If it works from your laptop but not in the cluster, compare credentials and network: you may be logged in locally, or the nodes may not reach the registry (look for i/o timeout, no such host or x509: certificate signed by unknown authority in the event instead).

imagetools inspect also lists the platforms in the index, which settles architecture questions.

Step 3: Fix credentials for a private registry

Create a registry Secret in the Pod’s namespace and reference it:

kubectl create secret docker-registry ghcr-pull -n shop \
  --docker-server=ghcr.io \
  --docker-username=ci-bot \
  --docker-password="$GHCR_READ_TOKEN"
Terminal
spec:
  imagePullSecrets:
    - name: ghcr-pull
  containers:
    - name: api
      image: ghcr.io/acme/shop-api:1.4.2
yaml

To avoid repeating this on every workload, add it to the ServiceAccount the Pods use with kubectl patch serviceaccount default -n shop -p '{"imagePullSecrets":[{"name":"ghcr-pull"}]}'. Note that changing a ServiceAccount does not alter existing Pods; recreate them.

Tokens with short lifetimes (Amazon ECR’s are valid for 12 hours) need refreshing. Prefer the cloud’s kubelet credential provider or a controller that rotates the Secret over a hand-made Secret that silently expires.

Step 4: Fix architecture mismatches

An index that lacks your node’s platform fails with no match for platform in manifest. This happens when an image is built on an Apple Silicon laptop (linux/arm64) and run on linux/amd64 nodes, or the reverse on Graviton or Ampere nodes. Build for every node architecture:

docker buildx build --platform linux/amd64,linux/arm64 -t ghcr.io/acme/shop-api:1.4.3 --push .
Terminal

If the image was pushed as a single-platform manifest rather than an index, the pull can succeed and the container then fails at start with exec format error, which shows up as CrashLoopBackOff instead.

Step 5: Deal with Docker Hub rate limits

Nodes behind one NAT gateway share a public IP, so a scale-up that pulls the same public image on 30 nodes can hit the anonymous limit. Docker’s published limits at the time of writing are 100 pulls per 6 hours per IPv4 address (or IPv6 /64) for unauthenticated users and 200 for a Personal account; check Docker’s usage page for current numbers. Fixes, best first: mirror the images you depend on into your own registry, configure a pull-through cache, or pull with an authenticated paid account.

Worked scenario

The shop team moves its API image to a private GitHub Container Registry repository. The new Deployment runs in namespace shop, but the pull secret was created months ago in default. Events read:

Warning  FailedToRetrieveImagePullSecret  kubelet  Unable to retrieve some image pull secrets (ghcr-pull); attempting to pull the image may not succeed.
Warning  Failed  kubelet  Failed to pull image "ghcr.io/acme/shop-api:1.4.2": ... failed to authorize: failed to fetch anonymous token: ... 401 Unauthorized
Text

The first event is the clue most people scroll past. The kubelet looked for ghcr-pull in shop, did not find it, and pulled anonymously. anonymous token in the second message confirms that no credentials were sent.

kubectl get secret ghcr-pull -n shop
# Error from server (NotFound): secrets "ghcr-pull" not found
kubectl get secret ghcr-pull -n default
# NAME        TYPE                             DATA   AGE
# ghcr-pull   kubernetes.io/dockerconfigjson   1      212d
Terminal

The fix is to create the Secret in shop (from the source of truth, not by copying the old one, which may hold an expired token), then delete the stuck Pod so the ReplicaSet creates one that retries immediately instead of waiting out the backoff.

Common mistake

The most tempting fix is adding imagePullPolicy: Always or IfNotPresent and redeploying. Pull policy decides when the kubelet pulls, not whether the registry will allow it. If the tag does not exist or the credentials are wrong, every policy except Never fails the same way, and Never only works if someone pre-loaded the image onto every node.

The reverse mistake hides problems: with IfNotPresent and a mutable tag such as :1.4 or :main, a node that pulled the tag last week keeps running the old image while a new node pulls the new one. Pods behave differently depending on placement. Use immutable tags or digests (image@sha256:...) for anything deployed.

Also, do not assume a cached image is a free pass for Pods without credentials. With the KubeletEnsureSecretPulledImages feature (beta and on by default since 1.35), the kubelet checks that a Pod is authorized for a private image even when it is already on the node.

Verify the behavior

Watch the Pod move through the pull:

kubectl get pod -n shop -l app=api -w
kubectl describe pod -n shop -l app=api | grep -E 'Pulling|Pulled|Failed'
Terminal

Success looks like:

Normal  Pulling  kubelet  Pulling image "ghcr.io/acme/shop-api:1.4.2"
Normal  Pulled   kubelet  Successfully pulled image "ghcr.io/acme/shop-api:1.4.2" in 3.2s (3.2s including waiting). Image size: 48213765 bytes.
Text

Then confirm what actually runs, by digest: kubectl get pod -n shop -l app=api -o jsonpath='{.items[*].status.containerStatuses[*].imageID}'. Every replica should show the same sha256.

Interview exercise

Pods of one Deployment run on two of five nodes. On the other three they show ImagePullBackOff with pull access denied. The image reference has no typo. What explains the split, and how do you fix it properly?

Answer and reasoning

A split by node means the difference is node state, not the manifest. The likeliest cause is that the image is cached on two nodes (pulled earlier, perhaps by someone with credentials or before the repository went private) and the Pod has no working pull credentials. With imagePullPolicy: IfNotPresent and an older kubelet, those two nodes start the container from cache without contacting the registry, while the others must pull and are refused. I would confirm with kubectl describe on a failing Pod (look for FailedToRetrieveImagePullSecret or anonymous-token wording) and by checking imagePullSecrets on the Pod and its ServiceAccount. The proper fix is credentials that every node uses: an imagePullSecrets entry in the Pod’s namespace, or a node credential provider. Relying on cache is fragile because a node replacement or image garbage collection breaks it, and on 1.35+ the credential verification feature removes it anyway.

Continue learning

More in Kubernetes

esc