Ch. 14 · Kubernetes

Kubernetes Horizontal Pod Autoscaling

Scale replicas from metrics with the HPA, set sensible targets, and stabilize against flapping.

~2 min readadvancedupdated Oct 5, 2026

The Horizontal Pod Autoscaler adjusts a Deployment’s replica count based on observed metrics. It compares current utilization to a target and adds or removes pods, but it only works when each pod declares resource requests, because utilization is a ratio against the request.

Before you start

You should be comfortable with Deployments, resource requests and metrics. This article covers the HPA and its prerequisites.

Step-by-step walkthrough

Step 1: Provide a metrics source

The HPA reads metrics from the metrics server for CPU and memory, or from an adapter for custom metrics such as queue depth. Without a metrics source, the HPA cannot compute a ratio and reports unknown metrics.

Step 2: Set a target and requests

A target like 70% CPU means the HPA scales so average CPU stays near seventy percent of the pod’s request. If pods have no CPU request, utilization is undefined and the HPA does nothing. Requests are therefore a prerequisite, not an optional detail.

Step 3: Stabilize against flapping

Scaling is deliberately smooth: the HPA waits for a stabilization window before scaling down and uses a scale-up policy, so it does not thrash with brief spikes. Tune behavior so the HPA reacts to sustained load rather than every blip.

Worked scenario

The HPA targets CPU and holds a floor and ceiling.

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: api
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: api
  minReplicas: 2
  maxReplicas: 20
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70
yaml

Walk through the example

The HPA keeps the api Deployment between two and twenty replicas, scaling so average CPU stays near 70% of the request. Missing requests would disable CPU scaling, which is why each pod must declare them. The min and max bound the reaction.

Common mistake

Autoscaling on CPU only for an I/O-bound service, which shows low CPU while requests queue; use a custom metric such as concurrency or queue depth instead. Another is forgetting requests, so the HPA has nothing to compute a ratio against.

Verify the behavior

Generate load and confirm the replica count rises toward the target utilization. Remove load and confirm it scales down after the stabilization window. Check the HPA status for metric values and any unknown-metric errors.

Interview exercise

Why does the HPA need resource requests to scale on CPU?

Answer and reasoning

Utilization is defined as current usage divided by the request for that resource. Without a request, the divisor is undefined, so the HPA cannot compute a percentage to compare against the target. Requests also drive scheduling, so they serve double duty: placement and autoscaling.

Continue learning

Compare capacity in Resource requests and Autoscaling signals. Read the Kubernetes HPA documentation and try the Kubernetes interview questions.

More in Kubernetes

read ✓Kubernetes · mid

Kubernetes ConfigMap Update Behavior

Understand why ConfigMap changes reach volumes but not environment variables, and how to roll a Deployment deliberately.

~2 min readread →
read ✓Kubernetes · mid

Kubernetes emptyDir Volumes

Share scratch space between containers in a pod with emptyDir, choose the backing medium, and bound its size.

~2 min readread →
esc