The Horizontal Pod Autoscaler adjusts a Deployment’s replica count based on observed metrics. It compares current utilization to a target and adds or removes pods, but it only works when each pod declares resource requests, because utilization is a ratio against the request.
Before you start
You should be comfortable with Deployments, resource requests and metrics. This article covers the HPA and its prerequisites.
Step-by-step walkthrough
Step 1: Provide a metrics source
The HPA reads metrics from the metrics server for CPU and memory, or from an adapter for custom metrics such as queue depth. Without a metrics source, the HPA cannot compute a ratio and reports unknown metrics.
Step 2: Set a target and requests
A target like 70% CPU means the HPA scales so average CPU stays near seventy percent of the pod’s request. If pods have no CPU request, utilization is undefined and the HPA does nothing. Requests are therefore a prerequisite, not an optional detail.
Step 3: Stabilize against flapping
Scaling is deliberately smooth: the HPA waits for a stabilization window before scaling down and uses a scale-up policy, so it does not thrash with brief spikes. Tune behavior so the HPA reacts to sustained load rather than every blip.
Worked scenario
The HPA targets CPU and holds a floor and ceiling.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api
minReplicas: 2
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70Walk through the example
The HPA keeps the api Deployment between two and twenty replicas, scaling so average CPU stays near 70% of the request. Missing requests would disable CPU scaling, which is why each pod must declare them. The min and max bound the reaction.
Common mistake
Autoscaling on CPU only for an I/O-bound service, which shows low CPU while requests queue; use a custom metric such as concurrency or queue depth instead. Another is forgetting requests, so the HPA has nothing to compute a ratio against.
Verify the behavior
Generate load and confirm the replica count rises toward the target utilization. Remove load and confirm it scales down after the stabilization window. Check the HPA status for metric values and any unknown-metric errors.
Interview exercise
Why does the HPA need resource requests to scale on CPU?
Answer and reasoning
Utilization is defined as current usage divided by the request for that resource. Without a request, the divisor is undefined, so the HPA cannot compute a percentage to compare against the target. Requests also drive scheduling, so they serve double duty: placement and autoscaling.
Continue learning
Compare capacity in Resource requests and Autoscaling signals. Read the Kubernetes HPA documentation and try the Kubernetes interview questions.