The Kubernetes objects, defaults and failure modes interviewers ask about, from the control plane down to a pod stuck in CrashLoopBackOff.
Architecture
- Declarative model: you store desired state in the API; controllers run loops that observe the actual state and act to close the gap (reconciliation).
| Component | Runs on | Job |
|---|---|---|
kube-apiserver |
control plane | front door: authentication → authorization (RBAC) → admission → validation → etcd. The only component that talks to etcd |
etcd |
control plane | consistent key-value store (Raft) for all cluster state; run 3 or 5 members for quorum and back it up |
kube-scheduler |
control plane | assigns pending pods to nodes: filter feasible nodes, then score them |
kube-controller-manager |
control plane | built-in controllers: Deployment, ReplicaSet, Node, Job, EndpointSlice… |
cloud-controller-manager |
control plane | cloud glue: load balancers, routes, node lifecycle |
kubelet |
every node | starts pods through the CRI, runs probes, reports node and pod status |
kube-proxy |
every node | programs Service virtual IPs (iptables/nftables); some CNIs replace it |
| container runtime | every node | containerd or CRI-O; dockershim was removed in v1.24 (Docker-built images still run) |
- Add-ons: CoreDNS (service discovery), a CNI plugin (Calico, Cilium, Flannel), metrics-server (
kubectl top, HPA), an Ingress or Gateway controller. - Networking model: every pod gets its own IP and can reach every other pod without NAT; containers in a pod share
localhost.
Core objects
| Object | What it is | Remember |
|---|---|---|
| Pod | 1+ containers sharing network and volumes | ephemeral; a bare pod isn’t recreated if its node dies |
| ReplicaSet | keeps N identical pods running | created and owned by Deployments |
| Deployment | stateless app with rolling updates | each template change creates a new ReplicaSet |
| StatefulSet | stable names (db-0), ordered start, a PVC per pod |
databases, brokers; pairs with a headless Service |
| DaemonSet | one pod per (matching) node | log shippers, node exporters, CNI agents |
| Job / CronJob | run to completion / on a schedule | backoffLimit default 6; concurrencyPolicy Allow, Forbid, Replace |
| Service | stable virtual IP + DNS for pods matching a selector | ClusterIP (default), NodePort (30000–32767), LoadBalancer, ExternalName, headless (clusterIP: None) |
| Ingress | HTTP(S) host/path routing to Services | needs a controller; API frozen, Gateway API is the successor |
| ConfigMap | non-secret config as env vars or files | 1 MiB max |
| Secret | sensitive data, base64-encoded | not encrypted in etcd unless you enable encryption at rest |
| PV / PVC / StorageClass | storage / a claim on it / a dynamic provisioning template | modes RWO, ROX, RWX, RWOP; reclaim Retain or Delete |
| Namespace | scope for names, RBAC and quotas | nodes, PVs and ClusterRoles are cluster-scoped |
| ServiceAccount | identity for pods calling the API | token is mounted, bound and auto-rotated |
- Service DNS:
web.prod.svc.cluster.local(justwebfrom the same namespace). StatefulSet pods getdb-0.db.prod.svc.cluster.local. externalTrafficPolicy: Localkeeps the client source IP and skips the extra hop, but only nodes with a ready pod receive traffic.
Deployment + Service
apiVersion: apps/v1
kind: Deployment
metadata: { name: web }
spec:
replicas: 3
selector: { matchLabels: { app: web } }
template:
metadata: { labels: { app: web } }
spec:
containers:
- { name: web, image: nginx:1.27, ports: [{ containerPort: 80 }] }apiVersion: v1
kind: Service
metadata: { name: web }
spec:
selector: { app: web }
ports: [{ port: 80, targetPort: 80 }]selector.matchLabelsmust match the template labels and is immutable after creation. The Service finds pods by label, not by Deployment.portis the Service’s port,targetPortthe container’s,nodePortthe port opened on every node (NodePort/LoadBalancer only).- Defaults:
strategy: RollingUpdatewithmaxSurge: 25%andmaxUnavailable: 25%,revisionHistoryLimit: 10,progressDeadlineSeconds: 600.
Probes
| Probe | Asks | On failure | Tip |
|---|---|---|---|
startupProbe |
has it finished booting? | container restarted; the others wait until it passes | failureThreshold × periodSeconds ≥ worst-case startup |
readinessProbe |
can it take traffic now? | removed from Service endpoints, not restarted | runs for the pod’s whole life |
livenessProbe |
is it stuck? | container killed and restarted | cheap and local: never check the database |
- Handlers:
httpGet(status 200–399 passes),tcpSocket,exec(exit 0 passes),grpc. - Defaults:
initialDelaySeconds 0,periodSeconds 10,timeoutSeconds 1,successThreshold 1(must be 1 for liveness and startup),failureThreshold 3.
# per container
resources:
requests: { cpu: 250m, memory: 256Mi }
limits: { memory: 256Mi }
startupProbe:
httpGet: { path: /healthz, port: 8080 }
failureThreshold: 30
periodSeconds: 10
readinessProbe:
httpGet: { path: /ready, port: 8080 }
livenessProbe:
httpGet: { path: /healthz, port: 8080 }Requests, limits & QoS
- Requests are what the scheduler reserves (a pod fits if the node’s allocatable minus other requests covers it). Limits are enforced at run time.
- Over the CPU limit the container is throttled; over the memory limit it is OOMKilled (exit 137).
- Units: CPU
1= one vCPU/core,250m= a quarter. MemoryMi/Giare binary;128mmeans 0.128 bytes. - Only a limit set? The request defaults to the limit.
| QoS class | Rule | Evicted under node pressure |
|---|---|---|
| Guaranteed | every container has CPU and memory requests equal to limits | last |
| Burstable | at least one request or limit, but not Guaranteed | after BestEffort |
| BestEffort | no requests or limits anywhere | first |
LimitRangesets per-container defaults and min/max in a namespace;ResourceQuotacaps a namespace’s totals (and then pods must declare those resources).- In-place resize (stable since v1.35) changes a running pod’s CPU/memory through the
resizesubresource without recreating it.
Scaling
| Tool | Scales | Notes |
|---|---|---|
HPA (built in, autoscaling/v2) |
replica count | CPU/memory utilization (% of requests), custom or external metrics; needs metrics-server |
| VPA (add-on) | pod requests/limits | modes Off (recommend only), Initial, Recreate, InPlaceOrRecreate; Auto is deprecated |
| Cluster Autoscaler / Karpenter | nodes | add nodes when pods are unschedulable, remove underused ones |
| KEDA (add-on) | replicas from events | queue length, cron, etc.; can scale to zero |
- HPA formula:
desired = ceil(current × currentMetric / targetMetric). It re-evaluates every 15 s and by default waits a 5-minute stabilization window before scaling down. kubectl autoscale deployment web --min=2 --max=10 --cpu=70%(older kubectl:--cpu-percent=70). Manual:kubectl scale deployment web --replicas=5.- Don’t let HPA and VPA both act on the same CPU/memory metric.
Scheduling
nodeSelector: { disktype: ssd }is the simplest node constraint.- Node affinity:
requiredDuringSchedulingIgnoredDuringExecution(hard) orpreferredDuringScheduling…(soft, weighted). “IgnoredDuringExecution”: later label changes don’t evict running pods. - Pod (anti-)affinity with a
topologyKey(kubernetes.io/hostname,topology.kubernetes.io/zone) co-locates or separates pods;topologySpreadConstraints(maxSkew,whenUnsatisfiable) spread replicas evenly across zones or nodes. - Taints repel, tolerations allow:
kubectl taint nodes n1 gpu=true:NoSchedule; remove withkubectl taint nodes n1 gpu:NoSchedule-.
| Taint effect | Meaning |
|---|---|
NoSchedule |
new pods without a matching toleration aren’t placed |
PreferNoSchedule |
avoid if possible (soft) |
NoExecute |
also evicts running pods without the toleration (after tolerationSeconds if set) |
- A toleration doesn’t attract pods: dedicate nodes with a taint plus node affinity.
RBAC & NetworkPolicy
Role(namespaced) andClusterRole(cluster-wide or reusable) list rules:apiGroups,resources,verbs(get,list,watch,create,update,patch,delete).RoleBinding/ClusterRoleBindinggrant a role to subjects:User,GrouporServiceAccount. A RoleBinding may reference a ClusterRole to grant it in one namespace only.- Permissions are purely additive (no deny rules). Users aren’t API objects; they come from certificates or OIDC.
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata: { name: pod-reader, namespace: dev }
rules:
- apiGroups: [""]
resources: ["pods", "pods/log"]
verbs: ["get", "list", "watch"]- Bind it and test it:
kubectl create rolebinding ci-read --role=pod-reader --serviceaccount=dev:ci -n dev, thenkubectl auth can-i list pods --as=system:serviceaccount:dev:ci -n dev. - NetworkPolicy (L3/L4 only) needs a CNI that enforces it; otherwise it silently does nothing. Pods are open until a policy selects them for a direction; policies then allow the union of their rules.
- Default deny:
podSelector: {}withpolicyTypes: [Ingress]and no rules. Remember to allow DNS (port 53) when you deny egress.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata: { name: api-from-frontend }
spec:
podSelector: { matchLabels: { app: api } }
policyTypes: [Ingress]
ingress:
- from: [{ podSelector: { matchLabels: { app: frontend } } }]
ports: [{ protocol: TCP, port: 8080 }]Rollouts & deployment strategies
kubectl set image deployment/web web=nginx:1.28 # template change = rollout
kubectl rollout status deployment/web
kubectl rollout history deployment/web
kubectl rollout undo deployment/web --to-revision=2
kubectl rollout restart deployment/web # new pods, same spec
kubectl rollout pause deployment/web # batch edits, then resume| Strategy | How | Trade-off |
|---|---|---|
| Rolling update | default; replace pods gradually | no downtime, but two versions run at once |
| Recreate | stop all old pods, then start new | downtime; for versions that can’t coexist |
| Blue/green | two Deployments, flip the Service selector | instant switch and rollback; double capacity |
| Canary | small traffic share to the new version | replica ratio, or weighted routing (Gateway API, mesh, Argo Rollouts, Flagger) |
- Zero-downtime needs readiness probes, enough replicas, a
PodDisruptionBudgetfor drains, and graceful shutdown: the pod getsSIGTERM, thenSIGKILLafterterminationGracePeriodSeconds(default 30).
kubectl cheat table
| Task | Command |
|---|---|
| List | kubectl get pods -o wide, -A (all namespaces), -w (watch), -l app=web |
| Details and events | kubectl describe pod web-abc, kubectl events --for pod/web-abc |
| Logs | kubectl logs web-abc -c app -f, --previous (last crash), -l app=web |
| Shell | kubectl exec -it web-abc -- sh |
| Debug container | kubectl debug -it web-abc --image=busybox --target=app |
| Local access | kubectl port-forward svc/web 8080:80 |
| Apply / preview | kubectl apply -f k8s/, kubectl diff -f k8s/, kubectl apply -k overlays/prod |
| Generate YAML | kubectl create deployment web --image=nginx --dry-run=client -o yaml |
| Output | -o yaml, -o jsonpath='{.status.phase}', --sort-by=.metadata.creationTimestamp |
| Usage | kubectl top pods, kubectl top nodes (metrics-server) |
| Schema | kubectl explain deployment.spec.strategy |
| Context | kubectl config get-contexts, use-context prod, set-context --current --namespace=dev |
| Node maintenance | kubectl cordon n1, kubectl drain n1 --ignore-daemonsets --delete-emptydir-data, kubectl uncordon n1 |
Troubleshooting
Start with kubectl get pods → describe pod (read the Events) → logs --previous → exec or debug.
| Symptom | Likely cause | Check |
|---|---|---|
CrashLoopBackOff |
app exits at start (config, env, command), or liveness kills it | logs --previous, describe (Last State, exit code); restarts back off 10 s, 20 s, 40 s… up to 5 min |
ImagePullBackOff / ErrImagePull |
wrong name or tag, private registry without imagePullSecrets, rate limit |
describe pod events |
Pending |
requests bigger than any node’s free capacity, taint without toleration, affinity mismatch, unbound PVC, quota | events (“0/3 nodes are available…”), describe node, get pvc |
OOMKilled (137) |
memory limit too low, leak, heap sized above the limit | describe Last State, kubectl top pod |
| Running, not Ready | readiness failing: wrong path or port, slow start, dependency down | events, get endpointslices -l kubernetes.io/service-name=web |
| Restarts under load | liveness timeout too tight or no startup probe | raise timeoutSeconds, add a startupProbe |
CreateContainerConfigError |
missing ConfigMap, Secret or key | describe pod |
| Service gets no traffic | selector doesn’t match pod labels, wrong targetPort, NetworkPolicy |
endpoints empty? port-forward to the pod directly |
Helm basics
- Chart = templated package (
Chart.yaml,values.yaml,templates/); release = an installed instance with revision history; charts come from repos or OCI registries. helm install web ./chart -f values-prod.yaml --set image.tag=1.4.2;helm upgrade --install web ./chart --waitis the idempotent CI form.helm list,helm history web,helm rollback web 3,helm uninstall web,helm template(render locally),helm lint,helm show values repo/chart.- Value precedence: chart defaults <
-ffiles (the last one wins) <--set. - Helm 3 dropped Tiller and stores release state as Secrets in the release’s namespace. Helm 4 (Nov 2025) uses server-side apply for new releases and renames
--atomicto--rollback-on-failure. - Helm templates and versions packages; Kustomize patches plain YAML with overlays (
kubectl apply -k).
Security best practices
- Least-privilege RBAC with one ServiceAccount per app;
automountServiceAccountToken: falsewhen the pod never calls the API. securityContext:runAsNonRoot: true,allowPrivilegeEscalation: false,readOnlyRootFilesystem: true,capabilities: { drop: ["ALL"] },seccompProfile: { type: RuntimeDefault }.- Enforce Pod Security Standards (
privileged,baseline,restricted) with the namespace labelpod-security.kubernetes.io/enforce: restricted. - Default-deny NetworkPolicies, then allow specific flows.
- Encrypt Secrets at rest (a KMS provider), or sync them from an external manager; restrict
geton Secrets, which exposes their values. - Pin images by digest, scan them, and enforce policy at admission (ValidatingAdmissionPolicy, Kyverno, OPA Gatekeeper).
- Keep the API server private, enable audit logs, upgrade regularly, and never mount the host’s container runtime socket.
Gotcha
Base64 is encoding, not encryption: anyone who can read a Secret object, or etcd, can read its value. The same goes for anyone who can create pods in the namespace, since a pod can mount any Secret there.
Quick answers
- Pod vs container? A pod is the scheduling unit: one or more containers that share an IP, ports and volumes and always run on the same node.
- Deployment vs StatefulSet? Deployments run interchangeable pods; StatefulSets give each pod a stable name, stable storage and ordered rollout.
- Why a Service? Pod IPs change; a Service gives a stable virtual IP and DNS name and balances across ready pods.
- ClusterIP vs NodePort vs LoadBalancer? Internal only; plus a port on every node; plus a cloud load balancer in front.
- Liveness vs readiness? Liveness failure restarts the container; readiness failure only takes it out of load balancing.
- Requests vs limits? Requests drive scheduling; limits cap usage (CPU throttled, memory OOM-killed).
- ConfigMap vs Secret? Same shape; Secrets are for sensitive data, get tighter RBAC and can be encrypted at rest.
- What happens on
kubectl apply? API server authenticates, authorizes, runs admission and stores the object in etcd; the Deployment controller creates a ReplicaSet, which creates pods; the scheduler binds them; the kubelet starts them. - How do you roll back?
kubectl rollout undo deployment/web(optionally--to-revision); old ReplicaSets are kept for that. - What is etcd? The Raft-based key-value store holding all cluster state; lose it without a backup and you lose the cluster.
- Ingress vs Gateway API? Both route HTTP into the cluster; Ingress is frozen, and Gateway API is its role-oriented, more expressive successor.
- Does a namespace isolate workloads? Only by name; add RBAC, NetworkPolicies and quotas for real isolation.
- Why is my pod Pending? The scheduler can’t place it: check the events for insufficient resources, taints, affinity or an unbound PVC.
Gotchas & traps
- Re-pushing the same tag doesn’t roll anything out: the pod template didn’t change. Use immutable tags or digests (or
rollout restart). - An omitted tag or
:latestdefaultsimagePullPolicytoAlways; any other tag defaults toIfNotPresent. - A liveness probe that checks the database restarts every pod when the database blips.
- CPU limits throttle even on an idle node; many teams set CPU requests without CPU limits but always set memory limits.
- ConfigMap and Secret changes reach mounted files eventually, but never env vars or
subPathmounts: restart the pods. - An RWO volume plus
RollingUpdatecan deadlock: the new pod lands on another node and hits a Multi-Attach error. UseRecreateor a StatefulSet. - After
SIGTERM, endpoint removal and shutdown race, so a pod may still get requests: handle them, or add a shortpreStopsleep. - Deleting a namespace deletes everything in it; it hangs in
Terminatingif a finalizer can’t complete. - A NetworkPolicy on a cluster whose CNI doesn’t enforce it is silently ignored.