Kubernetes interview questions & answers
Container orchestration: architecture, pods, deployments, services and ingress, config and secrets, probes, autoscaling, storage and troubleshooting.
Top 58 Kubernetes interview questions most asked first
1.What is Kubernetes, and what problems does it solve compared with running containers by hand?easy
Kubernetes is an open-source container orchestration platform, originally designed at Google and now hosted by the CNCF. You describe the desired state of your application declaratively, usually in YAML: which images, how many replicas, how they're exposed, how much CPU and memory they need. Kubernetes then works continuously to make the actual state match.
That solves the problems you hit once you have more than a handful of containers on more than one machine:
- Scheduling: placing containers on nodes that have capacity
- Self-healing: restarting crashed containers and replacing pods from failed nodes
- Scaling: changing replica counts, manually or automatically
- Service discovery and load balancing: stable names and virtual IPs in front of changing pods
- Rolling updates and rollbacks without downtime
- Configuration and secrets kept out of the image
The core idea is the reconciliation loop: controllers watch the desired state stored in the API server and keep correcting any drift.
What interviewers listen for- Orchestrates containers across a cluster of machines
- Declarative desired state, usually YAML manifests
- Controllers reconcile actual state toward desired state
- Self-healing, scaling, service discovery, rolling updates
Likely follow-up: What does "declarative" mean compared with imperative commands? · When would Kubernetes be overkill?
2.Explain the Kubernetes architecture. What runs on the control plane, and what runs on each worker node?easy
A cluster has a control plane that makes decisions and worker nodes that run the workloads.
Control plane:
- kube-apiserver: the front door. kubectl, controllers and kubelets all talk to it; it authenticates, authorizes, runs admission and persists objects.
- etcd: a consistent key-value store holding all cluster state. Only the API server talks to it.
- kube-scheduler: picks a node for each new pod, based on resources, affinity, taints and so on.
- kube-controller-manager: runs the built-in controllers (Deployment, ReplicaSet, Node, Job…), each reconciling desired and actual state.
- cloud-controller-manager (optional): integrates with a cloud provider for load balancers, routes and nodes.
Each node runs:
- kubelet: the agent that makes sure the pods assigned to its node are running, and reports their status.
- Container runtime: containerd or CRI-O, driven by the kubelet through the CRI.
- kube-proxy (optional): programs Service routing rules; some network plugins replace it.
What interviewers listen for- API server is the only component that talks to etcd
- etcd stores all cluster state
- Scheduler places pods; controllers reconcile state
- kubelet runs pods through a CRI container runtime
- kube-proxy programs Service routing on each node
Likely follow-up: What happens to running pods if the whole control plane goes down? · Why was dockershim removed?
3.What is a Pod, and why does Kubernetes schedule Pods instead of individual containers?easy
A Pod is the smallest deployable unit in Kubernetes: one or more containers that are always scheduled together on the same node and share a context. Containers in a pod share a network namespace (one IP address and port space, so they reach each other on
localhost) and can share volumes.Kubernetes schedules pods rather than containers because some processes are tightly coupled and must live and die together, like an app and its log shipper or proxy. The pod gives them a common lifecycle, IP and storage without merging them into one image.
Pods are ephemeral. If a pod dies or its node fails, it isn't resurrected: a controller creates a new pod with a new name and a new IP. That's why you rarely create bare pods. You use a Deployment, StatefulSet, DaemonSet or Job to manage them, and put a Service in front for a stable address.
What interviewers listen for- Smallest deployable unit: one or more containers
- Containers share one IP and can share volumes
- Co-scheduled on the same node with a shared lifecycle
- Pods are ephemeral; replacements get new names and IPs
- Manage pods with controllers, not as bare pods
Likely follow-up: How do two containers in the same pod communicate? · What are the pod phases?
4.What's the relationship between a Deployment, a ReplicaSet and a Pod? Why not just create a ReplicaSet directly?easy
They form a chain of controllers:
- A ReplicaSet keeps a set number of identical pods running. It counts the pods matching its label selector and creates or deletes pods from its pod template to reach
replicas. - A Deployment manages ReplicaSets. Each time you change the pod template (new image, env var, resources), it creates a new ReplicaSet and gradually scales it up while scaling the old one down. That's what gives you rolling updates, rollout history and
kubectl rollout undo.
You rarely create a ReplicaSet yourself because it has no update strategy: editing its template doesn't touch pods that already exist. Old ReplicaSets are kept at zero replicas as revision history, 10 by default (
revisionHistoryLimit), and that's what a rollback scales back up.The selector must match the template's labels, and in
apps/v1a Deployment's selector is immutable after creation.apiVersion: apps/v1 kind: Deployment metadata: { name: web } spec: replicas: 3 selector: matchLabels: { app: web } # must match the template labels template: metadata: labels: { app: web } spec: containers: - name: web image: nginx:1.27What interviewers listen for- ReplicaSet keeps N pods matching a selector
- Deployment manages ReplicaSets, one per template revision
- Template change creates a new ReplicaSet and rolls over
- Old ReplicaSets kept for rollback (default 10)
- Selector must match template labels and is immutable
Likely follow-up: What happens if you manually create a pod whose labels match a ReplicaSet selector?
- A ReplicaSet keeps a set number of identical pods running. It counts the pods matching its label selector and creates or deletes pods from its pod template to reach
5.What is a Service, and how do the
ClusterIP,NodePort,LoadBalancerandExternalNametypes differ?easyPods come and go with changing IPs, so a Service gives a set of pods, chosen by a label selector, a stable virtual IP and DNS name, and spreads traffic across the ready ones.
- ClusterIP (the default): a virtual IP reachable only from inside the cluster. Used for service-to-service traffic.
- NodePort: a ClusterIP plus a port opened on every node, by default from the 30000–32767 range. Traffic to any
nodeIP:nodePortreaches the Service. - LoadBalancer: asks the cloud provider, or a load-balancer controller such as MetalLB on bare metal, to provision an external load balancer for the Service. By default it's built on top of a NodePort. It's the usual way to expose one service publicly.
- ExternalName: no selector and no proxying. Cluster DNS just returns a CNAME to an external hostname such as
db.example.com.
A headless Service (
clusterIP: None) is a variation with no virtual IP: DNS returns the pod IPs directly.apiVersion: v1 kind: Service metadata: name: web spec: type: ClusterIP # or NodePort / LoadBalancer selector: app: web ports: - port: 80 # port clients connect to targetPort: 8080 # port the container listens onWhat interviewers listen for- Stable virtual IP and DNS name for a set of pods
- ClusterIP is the default and internal-only
- NodePort opens a port (30000–32767) on every node
- LoadBalancer provisions an external load balancer
- ExternalName is just a DNS CNAME, no proxying
Likely follow-up: Why would you put an Ingress in front instead of one LoadBalancer per service? · What is
targetPortversusport?6.How does a Deployment rolling update work? What do
maxSurgeandmaxUnavailablecontrol, and how do you roll back?midWhen the pod template changes, the Deployment creates a new ReplicaSet and shifts pods over step by step: scale the new one up, wait for new pods to become available, scale the old one down, repeat.
maxSurge: how many pods may exist above the desired count during the rollout.maxUnavailable: how many pods may be unavailable below the desired count.
Both take a number or a percentage, both default to 25%, and they can't both be 0.
maxSurge: 1, maxUnavailable: 0keeps full capacity but needs room for an extra pod; a highermaxUnavailableis faster and needs no spare capacity.Readiness probes are what make this safe: a new pod that never becomes ready stalls the rollout instead of taking traffic. After
progressDeadlineSeconds(600 by default) the Deployment reportsProgressDeadlineExceeded, but it does not roll back on its own.To roll back, check
kubectl rollout history deployment/web, then runkubectl rollout undo deployment/web, optionally with--to-revision=N.spec: replicas: 4 strategy: type: RollingUpdate rollingUpdate: maxSurge: 1 # at most 5 pods during the rollout maxUnavailable: 0 # never fewer than 4 availableWhat interviewers listen for- New ReplicaSet scaled up, old one scaled down
- maxSurge: extra pods; maxUnavailable: missing pods
- Both default to 25%; cannot both be 0
- Readiness gates progress; stalled rollouts are not auto-reverted
kubectl rollout undo, optionally--to-revision
Likely follow-up: What is the
Recreatestrategy and when would you use it? · How would you pause a rollout halfway?7.What's the difference between liveness, readiness and startup probes, and what happens when each one fails?mid
They answer different questions, and failures have different consequences:
- Readiness: can this container serve traffic right now? On failure the pod is marked not Ready and removed from Service endpoints, but it is not restarted. It runs for the container's whole life, so it also covers temporary overload.
- Liveness: is the container stuck beyond recovery, like a deadlock? After
failureThresholdconsecutive failures (3 by default) the kubelet restarts the container. - Startup: has the app finished starting? Until it succeeds, liveness and readiness probes don't run, so slow starters aren't killed early. If it keeps failing for
failureThreshold × periodSeconds, the container is restarted.
Each probe can be
httpGet(any status from 200 to 399 passes),tcpSocket,execorgrpc.A classic mistake is a liveness check that tests the database. If the database blips, every pod restarts at once. Liveness should only check the process itself.
startupProbe: httpGet: { path: /healthz, port: 8080 } failureThreshold: 30 periodSeconds: 10 # allows up to 300s to start livenessProbe: httpGet: { path: /healthz, port: 8080 } readinessProbe: httpGet: { path: /ready, port: 8080 } periodSeconds: 5What interviewers listen for- Readiness failure removes the pod from endpoints, no restart
- Liveness failure restarts the container
- Startup probe holds off the other probes until it passes
- Mechanisms: httpGet, tcpSocket, exec, grpc
- Don't check dependencies in liveness probes
Likely follow-up: Why is
initialDelaySecondsoften replaced by a startup probe? · Do you need a readiness probe to drain traffic on shutdown?8.How do ConfigMaps and Secrets differ, and what are the ways a Pod can consume them?easy
Both are namespaced key-value objects that keep configuration out of the image.
- ConfigMap: non-sensitive settings such as feature flags, URLs or whole config files.
- Secret: sensitive data such as passwords, tokens and TLS keys. Values are base64-encoded, which is encoding, not encryption. Secrets can be locked down separately with RBAC, and the kubelet keeps mounted secrets in
tmpfs(memory) rather than on disk.
A pod can consume either one as:
- Environment variables, one key with
valueFromor every key withenvFrom - Files in a volume, one file per key
- Command arguments built from those env vars
The practical difference: env vars are read once at container start, so changes need a restart. Mounted volumes are eventually updated by the kubelet (except with
subPath), so an app that re-reads its files can pick up changes. Both objects are limited to 1 MiB and can be markedimmutable: true.# in the container spec env: - name: DB_PASSWORD valueFrom: secretKeyRef: { name: db-creds, key: password } envFrom: - configMapRef: { name: app-config } volumeMounts: - { name: config, mountPath: /etc/app } # in the pod spec volumes: - name: config configMap: { name: app-config }What interviewers listen for- ConfigMap for plain config, Secret for sensitive data
- Secret values are base64-encoded, not encrypted
- Consume as env vars or as mounted files
- Env vars need a restart; volumes update eventually
- 1 MiB limit;
immutable: trueavailable
Likely follow-up: How would you roll pods automatically when a ConfigMap changes?
9.Are Kubernetes Secrets actually secure? Why aren't they encrypted by default, and how would you harden them?mid
Not by default. Secret values are only base64-encoded, and the API server stores them unencrypted in etcd. Anyone with access to etcd, an etcd backup, or
get/list/watchon Secrets can read them. And anyone who can create a pod in a namespace can mount any Secret in it, so "create pods" effectively implies reading Secrets.Encryption at rest isn't on by default because it needs keys that someone has to generate, store and rotate, which is an operator decision. Many managed services configure it for you.
To harden:
- Enable encryption at rest with an
EncryptionConfigurationpassed to the API server via--encryption-provider-config, ideally the KMS v2 provider so the key lives in an external KMS. Then rewrite existing Secrets so they're stored encrypted. - Use least-privilege RBAC; remember
listandwatchreveal values too. - Protect etcd: TLS, restricted access, encrypted backups.
- Consider an external secret store (Vault, a cloud secret manager) via the Secrets Store CSI Driver or External Secrets Operator.
What interviewers listen for- base64 is encoding; stored unencrypted in etcd by default
- Pod-create permission implies Secret read in the namespace
- Enable encryption at rest, preferably with KMS v2
- Least-privilege RBAC;
list/watchexpose values too - Protect etcd and its backups; consider external stores
Likely follow-up: After enabling encryption at rest, why are old Secrets still in plaintext? · How does the KMS provider use envelope encryption?
- Enable encryption at rest with an
10.When would you use a StatefulSet instead of a Deployment, and what guarantees does it give you?mid
Use a StatefulSet when each replica has its own identity and data: databases, Kafka, ZooKeeper, Elasticsearch. It adds three guarantees a Deployment doesn't give:
- Stable names: pods are
db-0,db-1,db-2, and a replacement keeps the same name. - Stable network identity: through the headless Service named in
serviceName, each pod gets its own DNS record, likedb-0.db-headless.default.svc.cluster.local, so peers can find each other. - Stable storage:
volumeClaimTemplatescreate one PVC per pod, and a rescheduled pod reattaches to the same volume.
With the default
OrderedReadypolicy, pods are created in order from 0 up, each waiting for the previous one to be Running and Ready, and removed in reverse. Rolling updates go from the highest ordinal down, andpartitionallows staged updates.Gotchas: by default, deleting or scaling down a StatefulSet does not delete its PVCs, and you have to create the headless Service yourself.
apiVersion: apps/v1 kind: StatefulSet metadata: { name: db } spec: serviceName: db-headless # headless Service you create yourself replicas: 3 selector: { matchLabels: { app: db } } template: metadata: { labels: { app: db } } spec: { containers: [{ name: db, image: postgres:17 }] } volumeClaimTemplates: # one PVC per pod: data-db-0, data-db-1... - metadata: { name: data } spec: { accessModes: [ReadWriteOnce], resources: { requests: { storage: 10Gi } } }What interviewers listen for- For stateful apps where replicas are not interchangeable
- Stable ordinal names: db-0, db-1, …
- Per-pod DNS via a headless Service
- Per-pod PVCs from volumeClaimTemplates survive rescheduling
- Ordered create, delete and update by default
Likely follow-up: How would you do a canary update of a StatefulSet? · What does
podManagementPolicy: Parallelchange?- Stable names: pods are
11.What are resource requests and limits, and how does Kubernetes use each of them?easy
They're set per container, for CPU and memory:
- Requests are what the container is guaranteed. The scheduler uses them: a pod only lands on a node whose unreserved allocatable capacity covers the sum of its requests. Requests also decide how CPU is shared under contention.
- Limits are the ceiling, enforced by the kernel through cgroups on the node. A container that hits its CPU limit is throttled; one that exceeds its memory limit is OOM-killed.
CPU is measured in cores (
250mis a quarter core) and memory in bytes (Mi,Gi). If you set only a limit, the request defaults to the limit. Together, requests and limits decide the pod's QoS class.In practice, always set requests, because without them scheduling and autoscaling are guesswork. Set a memory limit to contain leaks. Many teams skip CPU limits for latency-sensitive services to avoid throttling.
LimitRangecan inject defaults andResourceQuotacan cap totals per namespace.resources: requests: cpu: 250m # a quarter of a CPU core memory: 256Mi limits: cpu: "1" memory: 512MiWhat interviewers listen for- Requests drive scheduling and are guaranteed
- Limits are hard ceilings enforced by cgroups
- Over CPU limit: throttled; over memory limit: OOM-killed
- Only a limit set: request defaults to the limit
- Requests and limits determine the QoS class
Likely follow-up: Should you set CPU limits at all? · What do
LimitRangeandResourceQuotado?12.A pod is in
CrashLoopBackOff. What does that mean, and how do you troubleshoot it?midCrashLoopBackOffisn't a phase or a root cause. It means a container keeps exiting, and the kubelet is waiting before restarting it again. The delay grows exponentially (10s, 20s, 40s…) up to five minutes, and resets after the container runs cleanly for 10 minutes.My steps:
kubectl describe podfor the Last State: the reason and exit code, plus events. Exit code 1 is usually an app error; 137 means SIGKILL, often OOMKilled (the reason says so); 127 or 126 point at a bad command or entrypoint.kubectl logs --previousto read the logs of the instance that crashed, since the current one may have just started.- Check config: missing env vars, a Secret or ConfigMap key that doesn't exist, wrong args, a database it can't reach.
- Check probes: a too-aggressive liveness probe kills a healthy but slow app; the events show "Liveness probe failed".
- If the image has a shell, override the command with
sleep, or usekubectl debugto inspect it.
kubectl get pod web-7d9c -o wide kubectl describe pod web-7d9c # Last State, Reason, Exit Code, Events kubectl logs web-7d9c --previous # output of the crashed instance kubectl logs web-7d9c -c migrate # a specific container kubectl events --for pod/web-7d9cWhat interviewers listen for- Container keeps exiting; kubelet backs off restarts
- Backoff doubles from 10s, capped at 5 minutes
describeshows exit code and reason;logs --previous- Exit 137 often OOMKilled; exit 1 app error
- Check config, dependencies and liveness probes
Likely follow-up: How do you debug a container that crashes before you can exec into it?
13.What is an Ingress, and why does it need an ingress controller?easy
An Ingress is a set of layer-7 HTTP(S) routing rules: send
shop.example.com/apito theapiService and everything else toweb, and terminate TLS with a given certificate Secret.pathTypecan beExact,PrefixorImplementationSpecific.The Ingress object alone does nothing. An ingress controller (Traefik, HAProxy, an NGINX-based controller, or a cloud one like the AWS Load Balancer Controller) watches Ingress objects and configures a real proxy or load balancer to match.
ingressClassNamepicks which controller handles it. Usually the controller itself sits behind one LoadBalancer Service, so dozens of services share a single external IP instead of paying for one load balancer each.Worth knowing: the Ingress API is frozen. It's GA and isn't going away, but it won't get new features, and the Kubernetes project now recommends Gateway API. The popular community
ingress-nginxcontroller was retired in March 2026.apiVersion: networking.k8s.io/v1 kind: Ingress metadata: { name: shop } spec: ingressClassName: nginx tls: - { hosts: [shop.example.com], secretName: shop-tls } rules: - host: shop.example.com http: paths: - path: /api pathType: Prefix backend: { service: { name: api, port: { number: 80 } } }What interviewers listen for- Host- and path-based HTTP routing plus TLS termination
- Needs a controller to actually implement the rules
ingressClassNameselects the controller- Many services share one external load balancer
- Ingress API is frozen; Gateway API recommended
Likely follow-up: How is Gateway API different from Ingress? · Where would you terminate TLS?
14.What are namespaces used for, and are they a security boundary?easy
Namespaces divide one cluster into named scopes. Names must be unique within a namespace, not across the cluster, so
devandprodcan each have awebDeployment. They're also the unit that most policies attach to:- RBAC Roles and RoleBindings grant access per namespace
- ResourceQuota and LimitRange cap and default resource usage
- NetworkPolicies and Pod Security labels are applied per namespace
A new cluster starts with
default,kube-system,kube-publicandkube-node-lease. Some resources aren't namespaced at all: Nodes, PersistentVolumes, StorageClasses, ClusterRoles and Namespaces themselves (kubectl api-resources --namespaced=falselists them). Services in other namespaces are reached by DNS asname.namespace.On their own they're a soft boundary. Pods in different namespaces share nodes and kernels, and can talk to each other over the network unless NetworkPolicies say otherwise. For real multi-tenancy you add RBAC, NetworkPolicies, quotas and Pod Security Standards, or use separate clusters.
What interviewers listen for- Scope for names; unique within a namespace
- Unit for RBAC, quotas, LimitRanges, NetworkPolicies
- Some resources are cluster-scoped (nodes, PVs, ClusterRoles)
- Cross-namespace DNS:
service.namespace - Soft isolation only, without extra policies
Likely follow-up: How would you stop one team from using all the cluster resources?
15.What is the difference between labels and annotations, and how do selectors use labels?easy
Both are key-value metadata on objects, but they serve different purposes.
Labels identify and group objects, like
app=web,tier=frontend,env=prod. They're short (values up to 63 characters) and they're what the system selects on. A Service finds its pods by label, a ReplicaSet counts pods by label, a NetworkPolicy targets pods by label, andkubectl get pods -l app=webfilters by label.Selectors come in two styles: equality-based (
app=web,env!=prod) and set-based (env in (prod,staging),!canary). In manifests that'smatchLabelsandmatchExpressionswithIn,NotIn,ExistsandDoesNotExist.Annotations hold non-identifying data that you can't select on and that can be much larger: build info, a git SHA, contact details, or tool configuration such as ingress controller settings or Prometheus scrape hints.
kubectl applystores its last-applied configuration in one.A classic bug is a Service selector that doesn't match the pod labels: the Service exists but has no endpoints.
What interviewers listen for- Labels identify and group; selectors query them
- Services, ReplicaSets, NetworkPolicies select by labels
- Equality-based and set-based selectors
- Annotations: non-identifying metadata, not selectable
- Mismatched selector means a Service with no endpoints
16.What is a DaemonSet, and what would you use one for?easy
A DaemonSet runs one copy of a pod on every node, or on every node matching a selector. When a node joins the cluster it gets the pod automatically, and when a node is removed its pod is garbage-collected. You don't set a replica count.
Typical uses are node-level agents:
- Log collectors such as Fluent Bit or Fluentd
- Monitoring agents such as Prometheus node-exporter
- Networking: CNI agents and kube-proxy itself
- Storage and security daemons, like CSI node plugins
You limit it to some nodes with
nodeSelectoror node affinity, and add tolerations if it must run on tainted nodes such as the control plane. The controller also adds some tolerations automatically, which is why DaemonSet pods can run on cordoned nodes, and whykubectl drainneeds--ignore-daemonsets.Updates use
RollingUpdateby default, one node at a time unless you raisemaxUnavailable, orOnDelete, where pods only update when you delete them.What interviewers listen for- One pod per eligible node, no replica count
- New nodes get the pod automatically
- Used for log, monitoring, network and storage agents
- Target nodes with nodeSelector, affinity and tolerations
- RollingUpdate or OnDelete update strategies
Likely follow-up: How is a DaemonSet different from a Deployment with pod anti-affinity?
17.How does the Horizontal Pod Autoscaler decide how many replicas to run?mid
The HPA is a control loop in the controller manager, running every 15 seconds by default. It reads metrics, computes a replica count and updates the target's
scalesubresource:desired = ceil(current × currentMetric / targetMetric)So 4 pods averaging 90% CPU against a 70% target gives
ceil(4 × 90/70) = 6. It skips changes within a 10% tolerance, clamps the result tominReplicas/maxReplicas, and uses a 5-minute scale-down stabilization window by default so it doesn't flap.Key details:
- CPU and memory come from the Metrics API, usually metrics-server. Custom and external metrics need an adapter, such as Prometheus Adapter or KEDA.
- Utilization is relative to requests, so containers without requests can't be scaled on utilization.
- The
behaviorfield tunes scale-up and scale-down rates. - Don't hard-code
replicasin a Deployment the HPA manages, or everykubectl applyfights it.
apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: { name: web } spec: scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: web } minReplicas: 2 maxReplicas: 10 metrics: - type: Resource resource: name: cpu target: { type: Utilization, averageUtilization: 70 }What interviewers listen forceil(current × currentMetric / target)- Needs metrics-server; adapters for custom metrics
- Utilization is a percentage of the resource request
- Tolerance and 5-minute scale-down stabilization prevent flapping
- Bounded by minReplicas and maxReplicas
Likely follow-up: How would you scale on queue length instead of CPU? · Can HPA and VPA target the same Deployment?
18.Explain PersistentVolumes, PersistentVolumeClaims and StorageClasses, and how dynamic provisioning works.mid
They separate what storage an app needs from how it's provided:
- A PersistentVolume (PV) is a cluster-scoped piece of real storage, like a cloud disk or an NFS export, whose lifecycle is independent of any pod.
- A PersistentVolumeClaim (PVC) is a namespaced request: "20Gi, ReadWriteOnce, class fast-ssd". It binds one-to-one to a matching PV, and pods mount the claim, not the PV.
- A StorageClass describes a kind of storage: its CSI provisioner, parameters,
reclaimPolicyandvolumeBindingMode.
With dynamic provisioning, a PVC that names a class (or falls back to the default class) makes the provisioner create a PV on demand, so no admin pre-creates volumes.
WaitForFirstConsumerdelays that until a pod is scheduled, so the disk is created in the pod's zone; the default isImmediate.Access modes are
ReadWriteOnce(one node),ReadOnlyMany,ReadWriteManyandReadWriteOncePod(one pod). Dynamically provisioned PVs default to theDeletereclaim policy; useRetainfor data you can't lose.apiVersion: v1 kind: PersistentVolumeClaim metadata: { name: data } spec: accessModes: [ReadWriteOnce] storageClassName: fast-ssd resources: { requests: { storage: 20Gi } } --- apiVersion: storage.k8s.io/v1 kind: StorageClass metadata: { name: fast-ssd } provisioner: ebs.csi.aws.com volumeBindingMode: WaitForFirstConsumer reclaimPolicy: RetainWhat interviewers listen for- PV is real storage; PVC is a namespaced request
- Pods mount PVCs; a PVC binds to one PV
- StorageClass enables dynamic provisioning via a CSI driver
WaitForFirstConsumeravoids zone mismatches- Access modes RWO/ROX/RWX/RWOP; reclaim Delete or Retain
Likely follow-up: What does
ReadWriteOnceactually guarantee? · How do you resize a PVC?19.How does traffic sent to a Service ClusterIP actually reach a pod? What does kube-proxy do?mid
A ClusterIP is virtual: in the default mode no network interface owns it. Three pieces make it work:
- The EndpointSlice controller keeps a list of the IPs and ports of the ready pods matching the Service's selector.
- kube-proxy, running on every node, watches Services and EndpointSlices and programs packet-forwarding rules in the node's kernel.
- When a pod connects to
ClusterIP:port, those rules on the client's node DNAT the packet to one backend pod IP, picked at random (or bysessionAffinity: ClientIP). Connection tracking keeps the rest of that connection on the same pod, and the pod network carries it to whichever node the pod is on.
So kube-proxy isn't in the data path; the kernel does the forwarding. On Linux the default mode is iptables, nftables is the newer replacement and is expected to become the default, and IPVS mode is deprecated. Some CNIs, such as Cilium, replace kube-proxy with eBPF.
A consequence: balancing is per connection, not per request, so long-lived HTTP/2 or gRPC connections can pile onto a few pods.
What interviewers listen for- ClusterIP is virtual; EndpointSlices list ready pods
- kube-proxy programs kernel rules on every node
- DNAT to a backend pod on the client node
- Modes: iptables default, nftables newer, IPVS deprecated
- Load balancing is per connection, not per request
Likely follow-up: Why might gRPC traffic be unevenly balanced across pods? · What happens to traffic when a pod fails its readiness probe?
20.Walk me through how an HTTPS request from a user on the internet reaches a container in your cluster.hard
A typical path, with an ingress controller or gateway in front:
- DNS resolves
shop.example.comto the external load balancer the cloud created for the controller'sLoadBalancerService. - The cloud load balancer forwards to a node's NodePort, or straight to pod IPs if it supports IP-target mode.
- On the node, kube-proxy rules route it to an ingress controller pod, possibly on another node. With
externalTrafficPolicy: Localonly nodes running a controller pod receive traffic, which also preserves the client IP. - The ingress controller terminates TLS (unless the load balancer already did), matches host and path, and picks a backend. Many controllers send traffic directly to pod IPs from the Service's EndpointSlices rather than through the ClusterIP.
- The pod network (the CNI plugin) delivers the packet to the target pod's IP. Pod-to-pod traffic needs no NAT.
- The container receives it on its
containerPort. Only ready pods are listed as endpoints, which is why readiness probes matter for zero-downtime deploys.
The response returns along the same connection path.
What interviewers listen for- DNS to cloud load balancer to node or pod
- kube-proxy rules or IP mode reach the ingress pods
- Ingress controller terminates TLS and routes host/path
- Controllers often target pod IPs from EndpointSlices
- CNI delivers to the pod; only ready pods get traffic
Likely follow-up: Where can you lose the client source IP along this path? · What changes with a service mesh?
- DNS resolves
21.What are the three Pod QoS classes, how is a pod’s class decided, and why does it matter?mid
You don't set the QoS class; Kubernetes derives it from requests and limits and records it in
status.qosClass:- Guaranteed: every container has CPU and memory requests and limits, and each request equals its limit.
- Burstable: not Guaranteed, but at least one container has a CPU or memory request or limit.
- BestEffort: no container has any requests or limits.
It matters when a node runs short of memory. The kubelet sets each container's
oom_score_adjfrom its class (-997 for Guaranteed, 1000 for BestEffort, in between for Burstable), so the kernel OOM killer picks BestEffort processes first.For node-pressure eviction, the kubelet ranks pods by whether their usage exceeds their requests, then by priority, then by how far over requests they are. It doesn't use the QoS class directly, but the result is roughly BestEffort first and Guaranteed last.
So critical workloads should have honest memory requests, ideally Guaranteed. Guaranteed pods with whole-number CPU requests can also get exclusive cores under the static CPU manager policy.
What interviewers listen for- Derived from requests and limits, not set directly
- Guaranteed: requests equal limits for CPU and memory
- BestEffort: no requests or limits at all
- Drives OOM score; BestEffort is killed first
- Eviction ranks by usage over requests, then priority
Likely follow-up: Can a pod with only CPU requests be Guaranteed?
22.One container keeps getting
OOMKilled, and another is slow even though its average CPU usage looks low. What is going on in each case?midThey're the two ways limits bite, because memory is incompressible and CPU is compressible.
OOMKilled: when a container goes over its memory limit, the kernel OOM killer kills a process in its cgroup. The status shows
Reason: OOMKilledand exit code 137, and the container restarts, often into CrashLoopBackOff. A pod can also be killed or evicted under its limit if the whole node runs out of memory. Fix it by sizing the limit from real usage (kubectl top, working-set metrics, VPA recommendations), fixing leaks, and making runtimes container-aware, such as the JVM's-XX:MaxRAMPercentage.CPU throttling: a CPU limit is enforced as a CFS quota: CPU time per scheduling period, usually 100 ms. A multi-threaded app can burn its whole quota early in a period and then stall until the next one. Average usage looks low while tail latency spikes. The throttling counters in cAdvisor metrics show it. Fix it by raising or removing the CPU limit while keeping an accurate request, or by using fewer threads.
What interviewers listen for- Memory over limit: OOM-killed, exit code 137
- Node-level memory pressure can kill pods too
- CPU over limit: throttled by CFS quota, not killed
- Throttling causes latency spikes despite low averages
- Size from real usage; consider dropping CPU limits
Likely follow-up: How would you confirm throttling with metrics? · Why might a JVM app get OOMKilled with a heap well below the limit?
23.A pod has been stuck in
Pendingfor ten minutes. How do you find out why?midPendingmeans the pod was accepted but isn't running yet. Usually it hasn't been scheduled; sometimes it's scheduled and still pulling images.kubectl describe podtells you which: aFailedSchedulingevent from the scheduler explains why each group of nodes was rejected.Common causes:
- Insufficient CPU or memory: no node has enough unreserved allocatable capacity for the pod's requests. Actual usage doesn't matter. Lower the requests or add nodes; with a cluster autoscaler, look for scale-up events.
- Taints the pod doesn't tolerate, or a nodeSelector or affinity that matches no node.
- Anti-affinity or topology spread rules that can't be satisfied, for example more replicas than nodes with required per-node anti-affinity.
- An unbound PVC: no default StorageClass, no matching PV, or a volume in a different zone.
If the pod already has a node assigned, look at container states instead: it's probably pulling a large image or waiting on a volume mount. And if a ResourceQuota is the problem, the pod is never created at all, so check the ReplicaSet's events.
kubectl describe pod api-5f7b # read the Events at the bottom # e.g. FailedScheduling: 0/6 nodes are available: 3 Insufficient cpu, # 3 node(s) had untolerated taint {dedicated: gpu} kubectl describe node worker-1 # Allocated resources, Taints kubectl get pvc -n shop # an unbound claim blocks schedulingWhat interviewers listen forkubectl describe podand the FailedScheduling event- Scheduling uses requests, not actual usage
- Taints, selectors and affinity can exclude every node
- Unbound PVCs keep pods Pending
- Quota failures show on the ReplicaSet, not the pod
Likely follow-up: How does the cluster autoscaler react to Pending pods?
24.What does
ImagePullBackOffmean, and what are the usual causes?easyThe kubelet couldn't pull the container image. You first see
ErrImagePull, thenImagePullBackOffwhile it keeps retrying with an increasing delay, capped at five minutes. The container never starts, so there are no logs; the answer is inkubectl describe podevents, which show the registry's actual error.Usual causes:
- A typo in the image name, or a tag that doesn't exist ("manifest unknown", "not found").
- A private registry without credentials ("unauthorized"). Create a
docker-registrySecret and reference it inimagePullSecretson the pod or its ServiceAccount, or grant the nodes registry access through cloud IAM. - Registry rate limits, like Docker Hub's anonymous pull limits.
- Network problems from the node: DNS, egress firewall, proxy ("i/o timeout").
- Wrong architecture: the image has no variant for the node's platform, such as arm64.
To avoid surprises, pin images by tag or digest instead of
:latest, and remember that:latestwith no explicitimagePullPolicymeansAlways.What interviewers listen for- Kubelet cannot pull the image; retries with backoff
- Check
kubectl describe podevents for the registry error - Typos, missing tags, missing
imagePullSecrets - Rate limits, network or architecture mismatches
- Pin tags or digests instead of
:latest
25.How does RBAC work in Kubernetes? Explain Role, ClusterRole, RoleBinding and ClusterRoleBinding.mid
RBAC decides whether a subject may perform a verb on a resource. It's purely additive: there are no deny rules, so anything not granted by some binding is refused.
- Role: a list of rules (API groups, resources, verbs) inside one namespace.
- ClusterRole: the same, but not namespaced. Use it for cluster-scoped resources like nodes, for non-resource URLs, or as a reusable template.
- RoleBinding: grants a Role or a ClusterRole to subjects (users, groups, ServiceAccounts) within one namespace. Binding a ClusterRole this way limits it to that namespace; that's how the built-in
view,editandadminroles are usually handed out. - ClusterRoleBinding: grants a ClusterRole across the whole cluster.
Kubernetes has no User objects; users and groups come from authentication, such as certificates or OIDC. A binding's
roleRefcan't be changed; you recreate the binding. Test withkubectl auth can-i list pods -n dev --as=system:serviceaccount:dev:ci-bot, and avoid wildcards,cluster-admin, and broad access to Secrets orpods/exec.apiVersion: rbac.authorization.k8s.io/v1 kind: Role metadata: { name: pod-reader, namespace: dev } rules: - apiGroups: [""] # "" is the core API group resources: [pods, pods/log] verbs: [get, list, watch] --- apiVersion: rbac.authorization.k8s.io/v1 kind: RoleBinding metadata: { name: ci-reads-pods, namespace: dev } subjects: [{ kind: ServiceAccount, name: ci-bot, namespace: dev }] roleRef: { kind: Role, name: pod-reader, apiGroup: rbac.authorization.k8s.io }What interviewers listen for- Additive only; no deny rules
- Role is namespaced; ClusterRole is cluster-wide
- RoleBinding can bind a ClusterRole within one namespace
- Subjects: users, groups, ServiceAccounts
- Verify with
kubectl auth can-i --as
Likely follow-up: Why is permission to create pods close to permission to read Secrets? · What do the
escalateandbindverbs protect against?26.What is a ServiceAccount, and how does a pod authenticate to the Kubernetes API?mid
A ServiceAccount is an identity for workloads, as opposed to human users. It's namespaced, every namespace gets one called
default, and a pod usesdefaultunless you setserviceAccountName.By default the kubelet mounts a short-lived, automatically rotated token into the pod at
/var/run/secrets/kubernetes.io/serviceaccount/, together with the cluster CA and the namespace name. The token comes from the TokenRequest API and is bound to the pod, so it stops working once the pod is deleted. Since v1.24, Kubernetes no longer auto-creates long-lived token Secrets for each ServiceAccount.Client libraries use this "in-cluster config": they find the API server through the
KUBERNETES_SERVICE_HOSTvariable and send the token as a bearer token. The API server authenticates it assystem:serviceaccount:<namespace>:<name>and RBAC authorizes it.Good practice: one ServiceAccount per app with minimal RBAC, no permissions on
default, andautomountServiceAccountToken: falsefor pods that never call the API. Cloud workload identity features map ServiceAccounts to cloud IAM roles.What interviewers listen for- Namespaced identity for pods;
defaultif unspecified - Projected, short-lived, auto-rotated token
- No auto-created long-lived token Secrets since v1.24
- Authenticated as
system:serviceaccount:ns:name, then RBAC - Disable automount when the API is not needed
Likely follow-up: How would a pod get AWS or GCP credentials without static keys?
- Namespaced identity for pods;
27.What are Jobs and CronJobs, and which settings control retries, parallelism and overlapping runs?easy
A Job runs pods to completion, unlike a Deployment, which keeps pods running forever. It's done when enough pods succeed:
completionsandparallelismcontrol how many successes you need and how many pods run at once.completionMode: Indexedgives each pod an index for sharded work.backoffLimitsets how many failures are retried before the Job is marked failed (default 6).activeDeadlineSecondscaps total runtime.ttlSecondsAfterFinishedcleans up finished Jobs automatically.- The pod's
restartPolicymust beOnFailureorNever.
A CronJob creates Jobs on a cron schedule.
concurrencyPolicydecides what happens if the previous run is still going:Allow(the default),Forbid(skip the new run) orReplace(cancel the old one). You can also setstartingDeadlineSecondsfor late starts, history limits (3 successful and 1 failed Job kept by default),suspendandtimeZone.A CronJob creates a Job roughly once per scheduled time, but occasionally two or none, so jobs should be idempotent.
apiVersion: batch/v1 kind: CronJob metadata: { name: nightly-report } spec: schedule: "0 2 * * *" # 02:00 every day timeZone: "Etc/UTC" concurrencyPolicy: Forbid # skip a run if the last one is still going jobTemplate: spec: backoffLimit: 3 template: spec: restartPolicy: OnFailure containers: [{ name: report, image: example/report:1.4 }]What interviewers listen for- Job: run to completion; Deployment: run forever
completions,parallelism,backoffLimit(default 6)- Pod restartPolicy must be OnFailure or Never
- CronJob
concurrencyPolicy: Allow, Forbid, Replace - Scheduling is approximate, so make jobs idempotent
Likely follow-up: How would you run a database migration before a deployment?
28.How do taints and tolerations work, and how would you dedicate a pool of GPU nodes to specific workloads?mid
A taint on a node repels pods; a toleration on a pod lets it ignore a matching taint. That's the opposite of affinity: affinity attracts, taints repel.
Taint effects:
NoSchedule: pods without a matching toleration won't be scheduled there; running pods stay.PreferNoSchedule: a soft version; the scheduler tries to avoid the node.NoExecute: also evicts running pods that don't tolerate it.tolerationSecondslets a pod stay for a limited time.
Tolerations match with
operator: Equal(key, value and effect) orExists(any value).The key point for dedicated nodes: a toleration only allows a pod onto a node, it doesn't send it there. So taint the GPU nodes to keep everything else off, and give the GPU workloads both the toleration and a
nodeSelectoror node affinity for those nodes.Kubernetes uses taints itself: the node controller adds
not-readyandunreachableNoExecute taints, and pods get default 300-second tolerations for them, which is why pods on a dead node are evicted after about five minutes.# kubectl taint nodes gpu-1 dedicated=gpu:NoSchedule tolerations: - key: dedicated operator: Equal value: gpu effect: NoSchedule nodeSelector: pool: gpu # a toleration permits the node; this attracts the podWhat interviewers listen for- Taints repel; tolerations allow, they do not attract
- NoSchedule, PreferNoSchedule, NoExecute
- NoExecute evicts running pods;
tolerationSecondsdelays it - Dedicated nodes: taint plus toleration plus node selector
- Default 300s tolerations for not-ready and unreachable
Likely follow-up: Why do DaemonSet pods still run on cordoned nodes?
29.Compare
nodeSelector, node affinity and pod anti-affinity. How would you keep replicas of a service off the same node?midAll three constrain where the scheduler places a pod:
- nodeSelector: the simplest. The node must have all the listed labels.
- Node affinity: the same idea with operators (
In,NotIn,Exists,DoesNotExist,Gt,Lt) and two strengths:required…(hard) andpreferred…(weighted, soft). "IgnoredDuringExecution" means pods already running aren't evicted if node labels change later. - Pod affinity and anti-affinity: rules relative to other pods' labels within a topology domain set by
topologyKey, such askubernetes.io/hostnamefor a node ortopology.kubernetes.io/zonefor a zone. Affinity co-locates, say a cache next to its app; anti-affinity spreads.
To keep replicas apart, add pod anti-affinity on the app's own label with
topologyKey: kubernetes.io/hostname. Prefer the preferred form: with the required form, replicas beyond the number of nodes stay Pending. Inter-pod affinity is also expensive in large clusters, andtopologySpreadConstraintsis often the better tool for even spreading.affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - { key: topology.kubernetes.io/zone, operator: In, values: [eu-west-1a, eu-west-1b] } podAntiAffinity: preferredDuringSchedulingIgnoredDuringExecution: - weight: 100 podAffinityTerm: labelSelector: { matchLabels: { app: web } } topologyKey: kubernetes.io/hostnameWhat interviewers listen for- nodeSelector: exact label match
- Node affinity: operators plus required or preferred rules
- Pod anti-affinity spreads relative to other pods
topologyKeydefines node, zone or region domains- Required anti-affinity can leave extra replicas Pending
Likely follow-up: How do
topologySpreadConstraintsdiffer from anti-affinity?30.What is Helm, and what are charts, values and releases?easy
Helm is the package manager for Kubernetes. It saves you from hand-editing near-identical YAML for every environment.
- A chart is a package: templated manifests written in Go templates, a
Chart.yamlwith metadata, and avalues.yamlwith defaults. Charts can depend on other charts and are shared through repositories or OCI registries. - Values are the configuration you supply at install time, with
-f values-prod.yamlor--set image.tag=1.4, so one chart serves dev, staging and prod. - A release is an installed instance of a chart, with a name and a numbered revision history. Every upgrade creates a new revision, and
helm rollbackreturns to an earlier one.
Everyday commands are
helm upgrade --install(idempotent, good for CI),helm rollback,helm list, andhelm templateto render manifests locally for review.Since Helm 3 there's no server-side Tiller: the CLI talks to the API server with your own credentials and stores release state in Secrets in the release's namespace. The main alternative is Kustomize, which uses template-free overlays and is built into kubectl as
kubectl apply -k.What interviewers listen for- Package manager: templated, reusable manifests
- Chart = templates + Chart.yaml + default values
- Values customize a chart per environment
- Release = installed instance with revision history
- No Tiller since Helm 3; Kustomize is the alternative
Likely follow-up: When would you pick Kustomize over Helm? · How do you manage secrets in Helm values?
- A chart is a package: templated manifests written in Go templates, a
31.What are init containers, and when would you use them?easy
Init containers, listed under
spec.initContainers, run before the app containers start. They run one at a time, in order, and each must exit successfully before the next one starts. If one fails, the kubelet retries it according to the pod'srestartPolicy; withNever, the pod fails. While they run,kubectl get podsshows a status likeInit:0/2.Typical uses:
- Wait for a dependency, for example loop until the database's DNS name resolves.
- Prepare files: fetch config, clone a repo or render a template into a shared
emptyDirvolume. - Fix permissions on a mounted volume.
- One-off setup per pod. Schema migrations are usually better as a separate Job, because an init container runs once for every replica.
The advantage is separation: setup tools like
gitorcurlstay out of the app image, and the init container can have different permissions. Regular init containers don't support liveness, readiness or startup probes. An init container withrestartPolicy: Alwaysis a different thing: a native sidecar.What interviewers listen for- Run before app containers, sequentially, to completion
- A failed init container is retried per restartPolicy
- Used for waiting, file preparation and permissions
- Keeps setup tools out of the app image
- No probes;
restartPolicy: Alwaysmakes it a sidecar
Likely follow-up: Why is running migrations in an init container risky with many replicas?
32.When would you run more than one container in a Pod, and how do those containers communicate?easy
Only when containers are tightly coupled: they must run on the same node, start and stop together, and scale together. The common patterns:
- Sidecar: extends the main app, like a log shipper, a service-mesh proxy, or a config or certificate reloader.
- Ambassador: a local proxy for outbound connections, such as a database or cloud SQL proxy on
localhost. - Adapter: translates the app's output, for example an exporter that turns app stats into Prometheus metrics.
Containers in a pod communicate cheaply because they share context:
- The same network namespace: they talk over
localhostand can't both bind the same port. - Shared volumes, usually an
emptyDir, for exchanging files. - Optionally a shared process namespace (
shareProcessNamespace: true), so they can see and signal each other's processes.
The classic anti-pattern is putting an app and its database in one pod. They scale differently, a restart takes both down, and the database can't survive the pod. Those belong in separate workloads connected by a Service.
What interviewers listen for- Only for tightly coupled, co-scaled containers
- Sidecar, ambassador and adapter patterns
- Shared network: communicate over localhost
- Shared volumes such as emptyDir for files
- Don't pack an app and its database together
33.What problems do native sidecar containers solve, and how do you declare one?mid
Sidecars used to be just extra entries in
containers, which caused three problems:- No startup order: the app could start before its proxy was ready, so early requests failed.
- No shutdown order: the sidecar could exit before the app finished draining.
- Jobs never finished: the Job waited for every container, and a log shipper or proxy never exits.
Native sidecars, stable since v1.33, are init containers with
restartPolicy: Always:- They start in order with the other init containers, and the kubelet moves on as soon as the sidecar has started, gated by its startup probe if it has one. So they're running before the app containers.
- They keep running for the pod's lifetime and are restarted if they exit, whatever the pod's restart policy.
- On termination they're stopped after the main containers, in reverse order.
- They don't block Job completion.
- They support probes, and a sidecar's readiness counts toward the pod's readiness.
Service-mesh proxies and log or metrics agents are the main users.
spec: initContainers: - name: log-shipper image: example/log-shipper:1.0 restartPolicy: Always # this makes it a sidecar volumeMounts: [{ name: logs, mountPath: /var/log/app }] containers: - name: app image: example/app:2.0 volumeMounts: [{ name: logs, mountPath: /var/log/app }] volumes: - { name: logs, emptyDir: {} }What interviewers listen for- Declared as init containers with
restartPolicy: Always - Start before app containers and keep running
- Stopped after main containers on shutdown
- Don't block Job completion
- Stable since v1.33; support probes
Likely follow-up: How did people stop sidecars in Jobs before this feature existed?
34.How do pods discover other services in the cluster? What DNS names does Kubernetes create?easy
Through cluster DNS, usually CoreDNS running in
kube-systembehind a Service calledkube-dns. It watches Services and serves records for them:my-svc.my-ns.svc.cluster.localresolves to the Service's ClusterIP.cluster.localis the default cluster domain, but it's configurable.- Named ports get SRV records such as
_http._tcp.my-svc.my-ns.svc.cluster.local. - A headless Service resolves to the IPs of its ready pods, and StatefulSet pods get names like
db-0.db.my-ns.svc.cluster.local.
The kubelet writes each pod's
/etc/resolv.confto point at the DNS Service and adds search domains, starting withmy-ns.svc.cluster.local. That's why a pod can just callhttp://my-svcin its own namespace, andmy-svc.other-nsfor another namespace.One gotcha: the default
ndots:5makes short external names likeapi.example.comtry every search domain first, which multiplies DNS queries. A trailing dot or a tuneddnsConfigavoids that.There's also an older mechanism: environment variables such as
MY_SVC_SERVICE_HOST, but only for Services that existed when the pod started.What interviewers listen for- CoreDNS serves records for every Service
service.namespace.svc.cluster.localresolves to the ClusterIP- Search domains allow the short name within a namespace
- Headless Services resolve to pod IPs
ndots:5can multiply lookups for external names
Likely follow-up: How would you debug a pod that cannot resolve a Service name?
35.What is a headless Service, and when do you need one?mid
A headless Service sets
clusterIP: None. It gets no virtual IP, kube-proxy ignores it, and nothing load-balances it. Instead, a DNS lookup of the Service name returns the IPs of all its ready pods directly.You need one when clients must see or address individual pods:
- StatefulSets: the governing Service in
serviceNamegives each pod a stable name likedb-0.db.shop.svc.cluster.local, so clients can reach the primary specifically, and cluster members like Kafka brokers or etcd peers can find each other. - Client-side load balancing: a gRPC client can resolve every pod IP and balance per request, which works around the per-connection balancing of a normal ClusterIP.
- Discovery for clustered software that wants the full member list.
publishNotReadyAddresses: truepublishes pods before they're ready, which helps members find each other while a cluster is bootstrapping.The trade-off is that clients take on the work: they must re-resolve DNS as pods come and go, and not cache old IPs for too long.
What interviewers listen forclusterIP: None: no virtual IP, no kube-proxy- DNS returns the ready pod IPs directly
- Gives StatefulSet pods stable per-pod DNS names
- Enables client-side load balancing, e.g. gRPC
- Clients must handle DNS changes and caching
- StatefulSets: the governing Service in
36.How do NetworkPolicies work? How would you lock down traffic to an API so only the frontend can reach it?mid
By default the pod network is flat: every pod can reach every other pod, across namespaces. A NetworkPolicy is a layer 3/4 firewall for pods, selected by labels.
- It's enforced by the network plugin, such as Calico or Cilium. If the plugin doesn't support NetworkPolicy, the objects are accepted but do nothing.
- A pod becomes isolated for ingress or egress as soon as any policy of that type selects it. From then on, only traffic some policy allows gets through. Policies are additive: there's no deny rule and no ordering. Replies to allowed connections are allowed automatically.
- Peers are chosen with
podSelector,namespaceSelector(both in one entry means AND) oripBlock, plus ports.
The usual pattern is a default-deny policy per namespace (
podSelector: {}with no allow rules), then narrow allow policies like the one shown. If you also deny egress, remember to allow DNS tokube-dnson port 53, or everything breaks.For layer-7 rules, explicit denies or cluster-wide policies you need CNI-specific resources, such as Cilium's or Calico's own policy types.
apiVersion: networking.k8s.io/v1 kind: NetworkPolicy metadata: { name: api-allow-frontend, namespace: shop } spec: podSelector: { matchLabels: { app: api } } policyTypes: [Ingress] ingress: - from: - podSelector: { matchLabels: { app: frontend } } ports: - { protocol: TCP, port: 8080 }What interviewers listen for- Default: all pod traffic allowed
- Enforced by the CNI; unsupported plugins ignore it
- Selected pods become isolated; allowed traffic is a union
- Start with default deny, then allow specific flows
- Allow DNS egress when restricting egress
Likely follow-up: What is the difference between
namespaceSelectorandpodSelectorin the samefromentry versus separate entries?37.Compare rolling, blue-green and canary deployments. How would you implement each on Kubernetes?mid
- Rolling update: the Deployment default. Pods are replaced gradually under
maxSurgeandmaxUnavailable. It needs no extra infrastructure, but both versions serve traffic during the rollout, so changes must be backward compatible, and a rollback is just another rollout. (Recreateis the blunt alternative: stop everything, then start the new version, with downtime.) - Blue-green: run the complete new version (green) next to the old one (blue), test it, then switch all traffic at once, for example by changing a Service selector from
version: bluetoversion: green, or repointing a route. Rollback is an instant switch back. The cost is double capacity during the switch. - Canary: send a small share of real traffic to the new version, watch error rates and latency, then increase step by step. The crude way is two Deployments behind one Service, where traffic roughly follows the replica ratio. The precise way is weighted routing: Gateway API
backendRefsweights, a service mesh, or ingress controller features.
Tools like Argo Rollouts and Flagger automate blue-green and canary, including metric analysis and automatic rollback.
What interviewers listen for- Rolling: gradual, built in, versions overlap
- Blue-green: full parallel stack, instant switch and rollback
- Canary: small traffic share, promote on healthy metrics
- Weighted routing via Gateway API or a mesh
- Argo Rollouts or Flagger automate analysis and rollback
Likely follow-up: How do database schema changes affect these strategies?
- Rolling update: the Deployment default. Pods are replaced gradually under
38.Why is etcd so critical, and how do you back it up and restore it?hard
etcd stores all cluster state: every object, including Secrets. Running containers keep going without it for a while, but nothing can be created, changed or rescheduled. Losing it without a backup means rebuilding the cluster.
Backup: take a snapshot from a live member with
etcdctl snapshot save, using the etcd client TLS certificates, or snapshot the etcd data volume. Check it withetcdutl snapshot status. Run it on a schedule, store copies off the cluster and encrypted (a snapshot holds your Secrets, in plaintext unless encryption at rest is on), and rehearse restores.Restore: stop all API servers, restore the snapshot into a new data directory on each etcd member with
etcdutl snapshot restore, point etcd at that directory (on kubeadm, in the etcd static pod manifest), start etcd, then the API servers. Restarting the scheduler, controller manager and kubelets is recommended so nothing acts on stale data.On managed services like EKS, GKE or AKS you can't reach etcd. You back up at the resource level instead, with Velero or GitOps, plus volume snapshots.
# kubeadm certificate paths shown ETCDCTL_API=3 etcdctl --endpoints=https://127.0.0.1:2379 \ --cacert=/etc/kubernetes/pki/etcd/ca.crt \ --cert=/etc/kubernetes/pki/etcd/server.crt \ --key=/etc/kubernetes/pki/etcd/server.key \ snapshot save /backup/etcd-snapshot.db etcdutl --write-out=table snapshot status /backup/etcd-snapshot.db # restore into a NEW data dir, with all API servers stopped etcdutl --data-dir /var/lib/etcd-restored snapshot restore /backup/etcd-snapshot.dbWhat interviewers listen for- etcd holds all cluster state, including Secrets
etcdctl snapshot savewith TLS certificates- Store encrypted and off-cluster; test restores
- Restore with
etcdutlinto a new data dir, API servers stopped - Managed clusters: Velero or GitOps instead
Likely follow-up: Why should an etcd cluster have an odd number of members?
39.How does the kube-scheduler decide which node a pod runs on?mid
The scheduler watches for pods with no
spec.nodeNameand handles each in two phases:- Filtering removes nodes that can't run the pod: not enough unreserved allocatable CPU or memory for its requests, a nodeSelector or required affinity that doesn't match, taints it doesn't tolerate, host port conflicts, volume zone constraints, required pod anti-affinity, unschedulable nodes.
- Scoring ranks the remaining nodes with scoring plugins: spreading replicas, balancing resource usage, preferred affinity weights, topology spread, whether the image is already on the node.
It picks the highest-scoring node, breaking ties randomly, and binds the pod by setting its node through the API server. The kubelet on that node sees the assignment and starts the containers.
If no node passes filtering, the pod stays Pending with a
FailedSchedulingevent. A higher-priority pod, set through a PriorityClass, can trigger preemption, evicting lower-priority pods to make room.The scheduler is built from plugins, supports profiles and can run alongside other schedulers (
schedulerName). It only places pods once; it never moves running pods. The separate descheduler project does that.What interviewers listen for- Filtering removes infeasible nodes
- Scoring ranks feasible nodes; the best one wins
- Binding sets the node; the kubelet starts the pod
- Decisions use requests, taints, affinity and topology
- Preemption for priority; no rebalancing of running pods
Likely follow-up: How does pod priority and preemption work?
40.What is a PodDisruptionBudget, and what does it protect against and not protect against?mid
A PDB limits how many pods of an application can be taken down at once by voluntary disruptions: node drains, cluster upgrades, autoscaler scale-downs, anything that goes through the Eviction API. You set
minAvailableormaxUnavailable, as a number or a percentage, and a selector.If an eviction would break the budget, the API server refuses it with
429 Too Many Requests, andkubectl drainkeeps retrying until replacement pods are healthy somewhere else. That's how a three-replica service survives a rolling node upgrade.What it doesn't do:
- It can't prevent involuntary disruptions like node crashes or OOM kills, though those count against the budget.
- Deleting pods or a Deployment directly bypasses it.
- Deployment and StatefulSet rolling updates aren't limited by PDBs;
maxUnavailablein the workload controls those.
Pitfalls:
minAvailableequal to the replica count, ormaxUnavailable: 0, blocks drains forever, and so doesminAvailable: 1on a single-replica app. Crash-looping pods can also block drains;unhealthyPodEvictionPolicy: AlwaysAllowlets them be evicted.apiVersion: policy/v1 kind: PodDisruptionBudget metadata: { name: web-pdb } spec: minAvailable: 2 # or maxUnavailable: 1 selector: matchLabels: { app: web }What interviewers listen for- Limits voluntary disruptions through the Eviction API
minAvailableormaxUnavailable, plus a selector- Blocked evictions return 429; drain retries
- No protection from crashes or direct deletes
- Too-strict budgets block node drains
Likely follow-up: How would you drain a node that runs a single-replica app with a PDB?
41.Compare the Horizontal Pod Autoscaler, the Vertical Pod Autoscaler and the cluster autoscaler. How do they work together?mid
They scale different things:
- HPA changes the number of pods, based on CPU, memory or custom metrics. It's built in.
- VPA changes each pod's resource requests (and limits proportionally) based on observed usage. It's an add-on with its own CRD. Modes include
Off(recommendations only),Initial(set at pod creation),Recreate(evict pods to apply) andInPlaceOrRecreate, which uses in-place pod resize where it can. - Cluster autoscaler changes the number of nodes. It adds nodes to a node group when pods are Pending for lack of capacity, and removes underused nodes whose pods fit elsewhere, respecting PDBs. Karpenter is an alternative that provisions right-sized nodes directly.
They chain naturally: load rises, the HPA adds pods, some can't be scheduled, the cluster autoscaler adds nodes. The VPA keeps requests honest so both decide well, since scheduling works on requests, not usage.
The main rule: don't let the HPA and VPA act on the same CPU or memory metric for one workload, or they fight. Use VPA in
Offmode, or scale the HPA on custom metrics.What interviewers listen for- HPA scales replicas; VPA scales requests; CA scales nodes
- VPA is an add-on with Off, Initial and Recreate modes
- Cluster autoscaler reacts to Pending pods
- Pending pods from HPA trigger node scale-up
- Don't combine HPA and VPA on the same metric
Likely follow-up: How does Karpenter differ from the cluster autoscaler?
42.What are Custom Resource Definitions and Operators, and how does an operator work?hard
A CustomResourceDefinition extends the Kubernetes API with a new resource type, like
PostgresClusterorCertificate. It declares an OpenAPI v3 schema for validation, versions and optionally astatussubresource. After that, the API server serves the new type like a built-in one:kubectl get, RBAC and watches all work. But a CRD on its own just stores data.An operator is CRDs plus a custom controller that encodes the operational knowledge for one piece of software. Its reconcile loop:
- Watches its custom resources and the objects it creates.
- Compares the desired
specwith what actually exists. - Acts: creates or updates StatefulSets, Services and Secrets, runs backups, handles failover and version upgrades.
- Reports progress in
status.
Good operators are level-triggered and idempotent. They set owner references so garbage collection removes child objects, and use finalizers to clean up external resources before deletion. Examples include cert-manager, the Prometheus Operator, CloudNativePG and Strimzi. They're usually built with Kubebuilder or the Operator SDK.
The trade-offs: another privileged component to run and trust, and CRD version upgrades to manage.
What interviewers listen for- CRD adds a new API type with a schema
- A CRD alone only stores data
- Operator = CRD + controller with domain knowledge
- Reconcile loop: watch, compare, act, update status
- Owner references and finalizers handle cleanup
Likely follow-up: Why should reconcile logic be level-triggered instead of edge-triggered?
43.Which
kubectlcommands do you use day to day to inspect, change and debug workloads?easyGrouped by what I'm doing:
- Inspect:
kubectl get pods -n shop -o wide, filtered with-l app=web, across namespaces with-A, or live with-w.kubectl describe pod <name>is the first stop because of its events.kubectl get deploy web -o yamlshows the full object. - Logs and access:
kubectl logs <pod> -c <container> -f,--previousafter a crash,kubectl exec -it <pod> -- sh,kubectl port-forward svc/web 8080:80, andkubectl debugfor images without a shell. - Change:
kubectl apply -f, withkubectl diff -ffirst;kubectl rollout status,history,undoandrestart;kubectl scale deploy/web --replicas=5. - Cluster health:
kubectl top podsandkubectl top nodes(needs metrics-server),kubectl events,kubectl cordonandkubectl drain,kubectl auth can-i. - Discovery:
kubectl explain deployment.spec.strategy,kubectl api-resources. - Scaffolding:
kubectl create deployment web --image=nginx --dry-run=client -o yamlto generate a starting manifest.
And always check
kubectl config current-contextbefore changing anything, so you don't apply to production by accident.What interviewers listen forget,describe(events) and-o yamlto inspectlogs --previous,exec,port-forward,debugapply,diffandrolloutsubcommands to changetop,events,drain,auth can-ifor operations- Check the current context before changing anything
- Inspect:
44.How does Kubernetes compare with Docker Compose and Amazon ECS, and when would you pick each?easy
- Docker Compose describes a multi-container app in one YAML file and runs it on a single Docker host. It's ideal for local development and simple single-server setups, but there's no scheduling across machines, no rescheduling when a host dies, and no autoscaling.
- Amazon ECS is AWS's managed orchestrator. You write task definitions and services instead of pods and Deployments, and run them on EC2 or serverless on Fargate. There's no control plane to operate, and it integrates tightly with AWS: IAM roles per task, load balancers, CloudWatch. The trade-offs are AWS lock-in and a smaller, less extensible ecosystem.
- Kubernetes is the portable open standard across clouds and on-premises, with an extensible API and a huge ecosystem: Helm, operators, service meshes, GitOps. The cost is complexity: upgrades, networking, security and platform skills, even on EKS, GKE or AKS.
My rule of thumb: Compose for local development; ECS with Fargate for smaller teams all-in on AWS who want minimal operations; Kubernetes when you need portability, hybrid or multi-cloud, many teams and services, or its ecosystem.
What interviewers listen for- Compose: single host, great for local development
- ECS: managed, AWS-native, Fargate for serverless
- Kubernetes: portable, extensible, rich ecosystem
- Kubernetes costs more operational complexity
- Choose based on scale, team skills and cloud strategy
Likely follow-up: What would a migration from Compose to Kubernetes involve?
45.What are the most important security best practices for workloads running on Kubernetes?hard
I think in layers:
- Pod hardening: run as a non-root numeric UID, block privilege escalation, drop all capabilities, use a read-only root filesystem and the
RuntimeDefaultseccomp profile. Avoidprivileged,hostNetwork,hostPIDandhostPath. Enforce it with Pod Security Admission atrestricted. - Identity and access: least-privilege RBAC, one ServiceAccount per app, token automount off where unused, and no
cluster-adminfor CI pipelines. - Network: default-deny NetworkPolicies with explicit allows, and mTLS for sensitive traffic.
- Supply chain: minimal or distroless images, CVE scanning, pinning by digest, signing images with cosign and verifying signatures at admission, and only trusted registries.
- Secrets: encryption at rest, ideally with KMS, or an external secret store, and never secrets baked into images.
- Cluster: stay on a supported, patched version, keep the API endpoint private, enable audit logging, lock down etcd and kubelet access, and set resource limits so one workload can't starve a node.
securityContext: # per container runAsNonRoot: true runAsUser: 10001 allowPrivilegeEscalation: false readOnlyRootFilesystem: true capabilities: { drop: [ALL] } seccompProfile: { type: RuntimeDefault }What interviewers listen for- Non-root, no privilege escalation, drop capabilities
- Enforce with Pod Security Admission
restricted - Least-privilege RBAC and dedicated ServiceAccounts
- Default-deny NetworkPolicies
- Scan, sign and verify images; protect Secrets
Likely follow-up: How would you give a container a writable temp directory with a read-only root filesystem?
- Pod hardening: run as a non-root numeric UID, block privilege escalation, drop all capabilities, use a read-only root filesystem and the
46.What are the Pod Security Standards, and how does Pod Security Admission enforce them?hard
The Pod Security Standards define three profiles:
- Privileged: unrestricted, for trusted system workloads such as CNI or storage agents.
- Baseline: blocks known privilege escalations, such as privileged containers, host namespaces,
hostPathvolumes, host ports and dangerous capabilities. - Restricted: current hardening practice. On top of Baseline, pods must run as non-root, set
allowPrivilegeEscalation: false, dropALLcapabilities (onlyNET_BIND_SERVICEmay be added back), use aRuntimeDefaultorLocalhostseccomp profile, and use only a limited set of volume types.
Pod Security Admission is the built-in admission controller that replaced PodSecurityPolicy, which was removed in v1.25. You apply it per namespace with labels such as
pod-security.kubernetes.io/enforce: restricted. There are three modes:enforcerejects violating pods,auditrecords them in the audit log, andwarnreturns a warning to the user.A safe rollout is to enforce
baselinewithwarnandauditset torestricted, fix the violations, then enforcerestricted. Enforcement applies to pods, so a Deployment can be accepted while its pods are rejected; look at the ReplicaSet's events. For custom rules, use ValidatingAdmissionPolicy, Kyverno or Gatekeeper.What interviewers listen for- Three profiles: Privileged, Baseline, Restricted
- Restricted: non-root, no escalation, drop ALL, seccomp
- Applied per namespace with
pod-security.kubernetes.io/*labels - Modes: enforce, audit, warn
- Replaced PodSecurityPolicy, removed in v1.25
Likely follow-up: Why might a Deployment be created successfully but have no pods?
47.Pods restart every few minutes with
Liveness probe failed, or never become Ready. How do you debug probe failures?midStart with
kubectl describe pod: the events say exactly how the probe failed, and the message points to the cause.- "connection refused": nothing is listening on that port. The probe targets the wrong port, the app is still starting, or it's bound to
127.0.0.1while the kubelet probes the pod IP. - Timeout ("context deadline exceeded"):
timeoutSecondsdefaults to just 1 second. A slow health endpoint, a GC pause or CPU throttling makes it fail. Make the endpoint cheap, or raise the timeout. - Bad status code: anything outside 200–399 fails, for example a 401 because the health route requires authentication.
- Killed while starting: liveness runs before the app is up, causing a restart loop. Add a
startupProbewhosefailureThreshold × periodSecondscovers the worst-case startup. - Never Ready: check what the readiness endpoint depends on; a missing downstream service keeps it false.
Then reproduce the probe yourself:
kubectl execinto the pod and curl the endpoint, or port-forward to it.kubectl get endpointslices -l kubernetes.io/service-name=webshows which endpoints are ready.What interviewers listen for- Probe failure details are in the pod events
- Connection refused: wrong port, bind address or slow start
- Default
timeoutSecondsis only 1 second - Use a startup probe for slow starters
- Reproduce the probe with
execorport-forward
Likely follow-up: Should a readiness probe check the database?
- "connection refused": nothing is listening on that port. The probe targets the wrong port, the app is still starting, or it's bound to
48.What is Gateway API, and how does it improve on Ingress?hard
Gateway API is the successor to Ingress, developed by SIG Network. It's an add-on: you install its CRDs, and implementations such as Envoy Gateway, Istio, Cilium or cloud load balancers provide the controller.
What it improves:
- Role-oriented resources instead of one overloaded object. A GatewayClass (the infrastructure provider) says which controller to use. A Gateway (the cluster operator) defines listeners, ports, TLS and which namespaces may attach routes. Routes such as HTTPRoute and GRPCRoute (the app teams) attach to a Gateway through
parentRefs. - Expressive, portable features in the spec: header and query matching, weighted traffic splitting, redirects, rewrites, header changes and mirroring. With Ingress those needed controller-specific annotations.
- Safe cross-namespace sharing: a Gateway's
allowedRoutescontrols who can attach, and ReferenceGrant is required to reference backends in another namespace. - More protocols: HTTP and gRPC routes are stable; others exist at different maturity levels.
GatewayClass, Gateway, HTTPRoute and GRPCRoute are stable. With the Ingress API frozen, it's the recommended path, and the ingress2gateway tool helps migrate.
apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: { name: shop, namespace: shop } spec: parentRefs: [{ name: public-gw, namespace: infra }] hostnames: [shop.example.com] rules: - matches: [{ path: { type: PathPrefix, value: /api } }] backendRefs: - { name: api-v1, port: 80, weight: 90 } - { name: api-v2, port: 80, weight: 10 } # 10% canaryWhat interviewers listen for- Add-on CRDs; the successor to Ingress
- GatewayClass, Gateway and Routes map to roles
- Built-in weighted splits, header matching, rewrites
- Cross-namespace attachment needs explicit permission
- Recommended over the frozen Ingress API
Likely follow-up: Who would own the Gateway versus the HTTPRoute in your organization?
- Role-oriented resources instead of one overloaded object. A GatewayClass (the infrastructure provider) says which controller to use. A Gateway (the cluster operator) defines listeners, ports, TLS and which namespaces may attach routes. Routes such as HTTPRoute and GRPCRoute (the app teams) attach to a Gateway through
49.What is a service mesh, what does it give you, and when is it worth the cost?hard
A service mesh is an infrastructure layer that handles service-to-service traffic without changing application code. Its data plane is a set of proxies that intercept traffic: traditionally a sidecar in every pod (Envoy in Istio, Linkerd's own proxy), or per-node proxies as in Istio's ambient mode. Its control plane configures the proxies and issues workload certificates.
What you get:
- Security: automatic mTLS between services, with workload identities and authorization policies such as "only
checkoutmay callpayments". - Traffic control: retries, timeouts, circuit breaking, per-request load balancing (which fixes gRPC imbalance), weighted canaries, mirroring and fault injection.
- Observability: consistent latency, error and throughput metrics for every hop, plus tracing spans, though apps must still propagate trace headers.
The costs are real: another complex system to run and upgrade, extra latency and CPU and memory per hop, and one more layer to debug.
It's worth it with many services, a zero-trust or mTLS requirement, or advanced traffic management needs. For a handful of services, Gateway API, NetworkPolicies and good client libraries are usually enough.
What interviewers listen for- Proxies intercept traffic; control plane configures them
- Automatic mTLS and service-level authorization
- Retries, timeouts, traffic splitting, per-request balancing
- Uniform metrics and tracing across services
- Adds latency, resource cost and operational complexity
Likely follow-up: How does a sidecar-less mesh differ from the sidecar model?
- Security: automatic mTLS between services, with workload identities and authorization policies such as "only
50.What happens when a pod is terminated, and how do you make rolling deployments truly zero-downtime?hard
When a pod is deleted, by a rollout, a scale-down or a drain:
- It's marked terminating and its grace period starts (30 seconds by default).
- In parallel, the control plane marks its endpoints as not ready, so kube-proxy, ingress controllers and load balancers stop sending new traffic, while the kubelet runs any
preStophook and then sends SIGTERM to each container's PID 1. - Whatever is still running when the grace period ends gets SIGKILL.
The catch is that race. SIGTERM can arrive before every node and proxy has stopped routing to the pod, so an app that exits instantly drops requests.
Zero-downtime checklist:
- Handle SIGTERM: stop accepting connections, finish in-flight requests, then exit. Make sure your process is PID 1 or runs under an init that forwards signals; a shell-form entrypoint often swallows SIGTERM.
- Add a short preStop sleep so routing updates propagate first.
- Set
terminationGracePeriodSecondslonger than the preStop plus drain time, since preStop counts against it. - Use readiness probes,
maxUnavailable: 0, and PDBs for node drains.
spec: terminationGracePeriodSeconds: 45 containers: - name: web image: example/web:3.2 lifecycle: preStop: sleep: { seconds: 10 } # let endpoint removal propagate first readinessProbe: httpGet: { path: /ready, port: 8080 }What interviewers listen for- Endpoint removal and SIGTERM happen in parallel
- SIGKILL after the grace period (default 30s)
- App must handle SIGTERM and drain in-flight work
- preStop sleep covers routing propagation delay
- Grace period must cover preStop plus drain time
Likely follow-up: Why does
CMD npm startsometimes break graceful shutdown?51.Walk me through everything that happens after you run
kubectl apply -f deployment.yaml, until the containers are serving traffic.hard- kubectl reads the file, uses the current kubeconfig context for the server and credentials, and sends the object to the API server, as a patch computed from the last-applied state.
- The API server authenticates the request, authorizes it with RBAC, runs mutating admission (defaults, webhooks such as sidecar injection), validates the schema, runs validating admission (Pod Security, ValidatingAdmissionPolicy, webhooks), and persists the Deployment to etcd.
- The Deployment controller sees it through a watch and creates a ReplicaSet. The ReplicaSet controller creates Pod objects with no node assigned.
- The scheduler filters and scores nodes and binds each pod to one.
- The kubelet on that node sees the pod, has the container runtime pull images and create the pod sandbox, the CNI plugin assigns an IP, volumes are mounted, init containers run, then app containers start and probes begin.
- When readiness passes, the EndpointSlice controller adds the pod to its Services, and kube-proxy and ingress controllers route traffic to it.
No controller calls another directly; they coordinate through watches on the API server.
What interviewers listen for- Authentication, authorization, admission, then etcd
- Deployment controller creates a ReplicaSet, which creates pods
- Scheduler binds pods to nodes
- Kubelet, runtime, CNI and volumes start the containers
- Ready pods join EndpointSlices and receive traffic
Likely follow-up: Where would a mutating webhook failure show up? · What is the difference between client-side and server-side apply?
52.What does
externalTrafficPolicy: Localdo on a LoadBalancer or NodePort Service, and what are the trade-offs versusCluster?hardIt controls where a node may send traffic that arrives from outside the cluster.
Cluster(the default): any node can forward to any ready pod, including pods on other nodes. Load spreads evenly and every node can accept traffic, but the extra hop is source-NATed so replies come back the same way. The pod sees a node's IP, not the client's.Local: a node forwards only to pods on that same node. There's no second hop and no SNAT, so the client source IP is preserved (useful for allowlists, rate limiting and logs) and latency drops. Nodes with no local pod drop the traffic, so for LoadBalancer Services Kubernetes allocates ahealthCheckNodePortthat the cloud load balancer probes to find which nodes have endpoints.
The trade-offs of
Local: load can become imbalanced, because the load balancer spreads per node, not per pod, so a node with one pod gets as much traffic as a node with three. Pod moves also depend on the load balancer's health checks catching up. Spreading pods evenly helps.What interviewers listen for- Cluster: any node to any pod, SNAT hides client IP
- Local: only node-local pods, client IP preserved
- Local drops traffic on nodes without endpoints
healthCheckNodePorttells the LB which nodes have pods- Local risks imbalance: LB spreads per node, not per pod
53.How would you make sure only trusted, approved container images can run in your cluster?hard
I'd combine controls at build time and at admission time:
- Trusted registries only: reject pods whose images don't come from your approved registries. Enforce it at admission with ValidatingAdmissionPolicy (CEL rules, built in), Kyverno or OPA Gatekeeper.
- Immutable references: tags can be moved, digests can't. Require
image@sha256:…, or at least block:latestand untagged images. - Signing and provenance: sign images in CI with Sigstore cosign, attach SBOM and provenance attestations, and verify signatures at admission, with Kyverno's image verification, the Sigstore policy-controller, or Ratify with Gatekeeper.
- Vulnerability scanning: scan in CI with a tool like Trivy or Grype, fail builds above a severity threshold, and keep rescanning running images, since new CVEs appear after deployment.
- Minimal base images, distroless or scratch, to shrink the attack surface.
- Pull behavior in multi-tenant clusters: the
AlwaysPullImagesadmission plugin forces a registry check with the pod's own credentials, so one tenant can't reuse another's cached private image.
Admission policies are the enforcement point; everything else feeds them trustworthy data.
What interviewers listen for- Admission policy restricts allowed registries
- Pin images by digest; block
:latest - Sign with cosign and verify at admission
- Scan in CI and continuously for CVEs
- Minimal images reduce the attack surface
Likely follow-up: What happens to already-running pods when a new policy is added?
54.A worker node suddenly dies. What happens to its pods, step by step, and how long does recovery take?hard
- The kubelet stops renewing its Lease in
kube-node-lease. After a grace period (tens of seconds by default) the node controller sets the node'sReadycondition toUnknown, marks its pods not ready so they drop out of Service endpoints, and taints the nodenode.kubernetes.io/unreachable. - Pods get a default toleration for that taint of 300 seconds, so after about five minutes they're evicted.
- Deployment and ReplicaSet pods then get replacements on healthy nodes, because terminating pods no longer count toward the replica total.
- StatefulSet pods are not replaced automatically. Kubernetes can't tell a dead node from a network partition, and two
db-0pods could corrupt data. You confirm the node is really gone, then delete the Node, force-delete the pod, or add thenode.kubernetes.io/out-of-servicetaint. - ReadWriteOnce volumes stay attached to the dead node until they're detached. Kubernetes force-detaches after about six minutes if the node is still unhealthy.
- DaemonSet pods tolerate the taint and stay bound.
Shorter
tolerationSecondson critical stateless apps speeds up failover.What interviewers listen for- Missed Lease renewals mark the node Unknown
- Unreachable taint plus 300s default toleration
- ReplicaSets replace pods after about five minutes
- StatefulSets don't auto-replace; manual confirmation needed
- RWO volumes must detach before pods move
Likely follow-up: Why is force-deleting a StatefulSet pod dangerous?
- The kubelet stops renewing its Lease in
55.How do you debug a running pod whose image is distroless and has no shell? What does
kubectl debugdo?midkubectl execneeds a shell and tools inside the image, which distroless images deliberately leave out.kubectl debugbrings your own tools, in three ways:- Ephemeral container: adds a temporary container with a debug image to the running pod, without restarting it. With
--target, it shares the target container's process namespace, so you can see its processes, inspect/proc/<pid>/rootand use tools likenetstatornslookupin the pod's network namespace. Ephemeral containers are stable since v1.25. They can't have ports, probes or resources, they're never restarted, and they can't be removed from the pod afterwards. - Pod copy (
--copy-to): creates a copy of the pod, optionally with a debug container, a changed command or different images. That's useful for a container that crashes on startup, because you can start it with a different command. - Node debugging (
node/<name>): runs a pod in the node's host namespaces with the host filesystem at/host, for checking the kubelet or container runtime without SSH.
--profileselects preset security settings for the debug container, such asgeneral,baseline,restrictedorsysadmin.# ephemeral container attached to a running pod, sharing the app's process namespace kubectl debug -it api-6c9f --image=busybox:1.36 --target=api # copy of the pod plus a debug container (the original keeps serving) kubectl debug api-6c9f -it --image=busybox:1.36 --copy-to=api-debug --share-processes # pod in the node's host namespaces, host filesystem at /host kubectl debug node/worker-1 -it --image=busybox:1.36What interviewers listen for- Ephemeral containers add tools to a running pod
--targetshares the app container’s process namespace--copy-todebugs a modified copy of the podnode/<name>debugs the node without SSH- Ephemeral containers: no ports, probes or restarts
- Ephemeral container: adds a temporary container with a debug image to the running pod, without restarting it. With
56.How do you make the Kubernetes control plane highly available, and why does etcd need an odd number of members?hard
Run several control-plane nodes, typically three, spread across failure zones. Each component handles redundancy differently:
- kube-apiserver is stateless, so all replicas are active behind a load balancer. That load balancer's address is the cluster endpoint for kubectl and the kubelets.
- kube-scheduler and kube-controller-manager run on every control-plane node, but only one of each is active at a time. They use leader election with Lease objects in
kube-system, and a standby takes over when the leader stops renewing. - etcd uses the Raft consensus protocol: a write needs a majority of members. Three members tolerate one failure, five tolerate two. Four still tolerate only one, since the quorum is three, so an even count adds overhead without adding fault tolerance.
You can run etcd stacked on the control-plane nodes (fewer machines, but failures are coupled) or as a separate external cluster (more machines, better isolation).
If etcd loses quorum, the API stops accepting changes. Existing containers keep running on their nodes, but nothing is rescheduled or reconciled until quorum is restored. Managed services run all of this for you.
What interviewers listen for- Multiple control-plane nodes across zones
- API servers active-active behind a load balancer
- Scheduler and controller manager use leader election
- etcd quorum is a majority; odd sizes like 3 or 5
- Lost quorum: no changes, but workloads keep running
Likely follow-up: What are the trade-offs between stacked and external etcd?
57.How do
emptyDir,hostPathand PersistentVolumeClaim volumes differ, and when would you use each?easyA container's own filesystem is lost when the container restarts. Volumes are declared at the pod level and mounted into containers, and they differ mainly in how long the data lives:
- emptyDir: created empty when the pod lands on a node and shared by all its containers. It survives container crashes but is deleted with the pod. Use it for scratch space, caches and handing files between containers.
medium: Memorymakes it a tmpfs that counts against the container's memory limit, andsizeLimitcaps it. - configMap, secret, downwardAPI and projected: inject configuration, credentials or pod metadata as files.
- persistentVolumeClaim: durable storage backed by a PersistentVolume, such as a cloud disk or NFS, that outlives the pod and can follow it to another node.
- hostPath: mounts a directory from the node itself. It ties the pod to that node's data and is a serious security risk, since it can expose host files. Pod Security Baseline forbids it; keep it for trusted system agents, like a log collector reading
/var/log.
What interviewers listen for- emptyDir lives and dies with the pod
- configMap and secret volumes inject files
- PVCs provide durable storage that outlives pods
- hostPath exposes the node filesystem; avoid it
- Memory-backed emptyDir counts against memory limits
- emptyDir: created empty when the pod lands on a node and shared by all its containers. It survives container crashes but is deleted with the pod. Use it for scratch space, caches and handing files between containers.
58.How do you spread replicas evenly across zones and nodes? Explain
topologySpreadConstraints.hardA topology spread constraint tells the scheduler how unevenly matching pods may be distributed across domains, defined by a node label in
topologyKey: zones, nodes or racks.maxSkew: the maximum allowed difference between the number of matching pods in a domain and the lowest count in any eligible domain. WithmaxSkew: 1and three zones, six replicas land 2/2/2, and seven land 3/2/2.whenUnsatisfiable:DoNotSchedulemakes it a hard rule, so the pod stays Pending;ScheduleAnywaymakes it a scoring preference.labelSelector: which pods count, usually the app's own labels.matchLabelKeyscan restrict counting to pods of the same revision during rollouts.
Compared with pod anti-affinity, which is essentially all-or-nothing ("not on a node that already has one"), spread constraints scale with the replica count. A common combination is a hard rule across zones and a soft one across nodes, as shown.
Caveat: they're only evaluated at scheduling time. Scale-downs or node failures can leave things skewed, and the descheduler project can rebalance. Without explicit constraints, the scheduler still applies soft built-in defaults for nodes and zones.
topologySpreadConstraints: - maxSkew: 1 topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: DoNotSchedule # hard rule labelSelector: { matchLabels: { app: web } } - maxSkew: 1 topologyKey: kubernetes.io/hostname whenUnsatisfiable: ScheduleAnyway # soft preference labelSelector: { matchLabels: { app: web } }What interviewers listen fortopologyKeydefines domains such as zones or nodesmaxSkewbounds the imbalance between domains- DoNotSchedule is hard; ScheduleAnyway is soft
- Scales better than anti-affinity for many replicas
- Checked only at scheduling; no automatic rebalancing
Likely follow-up: What happens to spreading when a Deployment scales down?
No questions match that filter.
Prefer multiple choice? All 20 Kubernetes MCQs with answers →