Ch. 13

Kubernetes interview questions & answers

Container orchestration: architecture, pods, deployments, services and ingress, config and secrets, probes, autoscaling, storage and troubleshooting.

58 interview questions20 quiz questions0 notes
your progress0%

Top 58 Kubernetes interview questions most asked first

  1. 1.What is Kubernetes, and what problems does it solve compared with running containers by hand?easy

    Kubernetes is an open-source container orchestration platform, originally designed at Google and now hosted by the CNCF. You describe the desired state of your application declaratively, usually in YAML: which images, how many replicas, how they're exposed, how much CPU and memory they need. Kubernetes then works continuously to make the actual state match.

    That solves the problems you hit once you have more than a handful of containers on more than one machine:

    • Scheduling: placing containers on nodes that have capacity
    • Self-healing: restarting crashed containers and replacing pods from failed nodes
    • Scaling: changing replica counts, manually or automatically
    • Service discovery and load balancing: stable names and virtual IPs in front of changing pods
    • Rolling updates and rollbacks without downtime
    • Configuration and secrets kept out of the image

    The core idea is the reconciliation loop: controllers watch the desired state stored in the API server and keep correcting any drift.

    What interviewers listen for
    • Orchestrates containers across a cluster of machines
    • Declarative desired state, usually YAML manifests
    • Controllers reconcile actual state toward desired state
    • Self-healing, scaling, service discovery, rolling updates

    Likely follow-up: What does "declarative" mean compared with imperative commands? · When would Kubernetes be overkill?

  2. 2.Explain the Kubernetes architecture. What runs on the control plane, and what runs on each worker node?easy

    A cluster has a control plane that makes decisions and worker nodes that run the workloads.

    Control plane:

    • kube-apiserver: the front door. kubectl, controllers and kubelets all talk to it; it authenticates, authorizes, runs admission and persists objects.
    • etcd: a consistent key-value store holding all cluster state. Only the API server talks to it.
    • kube-scheduler: picks a node for each new pod, based on resources, affinity, taints and so on.
    • kube-controller-manager: runs the built-in controllers (Deployment, ReplicaSet, Node, Job…), each reconciling desired and actual state.
    • cloud-controller-manager (optional): integrates with a cloud provider for load balancers, routes and nodes.

    Each node runs:

    • kubelet: the agent that makes sure the pods assigned to its node are running, and reports their status.
    • Container runtime: containerd or CRI-O, driven by the kubelet through the CRI.
    • kube-proxy (optional): programs Service routing rules; some network plugins replace it.
    What interviewers listen for
    • API server is the only component that talks to etcd
    • etcd stores all cluster state
    • Scheduler places pods; controllers reconcile state
    • kubelet runs pods through a CRI container runtime
    • kube-proxy programs Service routing on each node

    Likely follow-up: What happens to running pods if the whole control plane goes down? · Why was dockershim removed?

  3. 3.What is a Pod, and why does Kubernetes schedule Pods instead of individual containers?easy

    A Pod is the smallest deployable unit in Kubernetes: one or more containers that are always scheduled together on the same node and share a context. Containers in a pod share a network namespace (one IP address and port space, so they reach each other on localhost) and can share volumes.

    Kubernetes schedules pods rather than containers because some processes are tightly coupled and must live and die together, like an app and its log shipper or proxy. The pod gives them a common lifecycle, IP and storage without merging them into one image.

    Pods are ephemeral. If a pod dies or its node fails, it isn't resurrected: a controller creates a new pod with a new name and a new IP. That's why you rarely create bare pods. You use a Deployment, StatefulSet, DaemonSet or Job to manage them, and put a Service in front for a stable address.

    What interviewers listen for
    • Smallest deployable unit: one or more containers
    • Containers share one IP and can share volumes
    • Co-scheduled on the same node with a shared lifecycle
    • Pods are ephemeral; replacements get new names and IPs
    • Manage pods with controllers, not as bare pods

    Likely follow-up: How do two containers in the same pod communicate? · What are the pod phases?

  4. 4.What's the relationship between a Deployment, a ReplicaSet and a Pod? Why not just create a ReplicaSet directly?easy

    They form a chain of controllers:

    • A ReplicaSet keeps a set number of identical pods running. It counts the pods matching its label selector and creates or deletes pods from its pod template to reach replicas.
    • A Deployment manages ReplicaSets. Each time you change the pod template (new image, env var, resources), it creates a new ReplicaSet and gradually scales it up while scaling the old one down. That's what gives you rolling updates, rollout history and kubectl rollout undo.

    You rarely create a ReplicaSet yourself because it has no update strategy: editing its template doesn't touch pods that already exist. Old ReplicaSets are kept at zero replicas as revision history, 10 by default (revisionHistoryLimit), and that's what a rollback scales back up.

    The selector must match the template's labels, and in apps/v1 a Deployment's selector is immutable after creation.

    apiVersion: apps/v1
    kind: Deployment
    metadata: { name: web }
    spec:
      replicas: 3
      selector:
        matchLabels: { app: web }   # must match the template labels
      template:
        metadata:
          labels: { app: web }
        spec:
          containers:
            - name: web
              image: nginx:1.27
    What interviewers listen for
    • ReplicaSet keeps N pods matching a selector
    • Deployment manages ReplicaSets, one per template revision
    • Template change creates a new ReplicaSet and rolls over
    • Old ReplicaSets kept for rollback (default 10)
    • Selector must match template labels and is immutable

    Likely follow-up: What happens if you manually create a pod whose labels match a ReplicaSet selector?

  5. 5.What is a Service, and how do the ClusterIP, NodePort, LoadBalancer and ExternalName types differ?easy

    Pods come and go with changing IPs, so a Service gives a set of pods, chosen by a label selector, a stable virtual IP and DNS name, and spreads traffic across the ready ones.

    • ClusterIP (the default): a virtual IP reachable only from inside the cluster. Used for service-to-service traffic.
    • NodePort: a ClusterIP plus a port opened on every node, by default from the 30000–32767 range. Traffic to any nodeIP:nodePort reaches the Service.
    • LoadBalancer: asks the cloud provider, or a load-balancer controller such as MetalLB on bare metal, to provision an external load balancer for the Service. By default it's built on top of a NodePort. It's the usual way to expose one service publicly.
    • ExternalName: no selector and no proxying. Cluster DNS just returns a CNAME to an external hostname such as db.example.com.

    A headless Service (clusterIP: None) is a variation with no virtual IP: DNS returns the pod IPs directly.

    apiVersion: v1
    kind: Service
    metadata:
      name: web
    spec:
      type: ClusterIP        # or NodePort / LoadBalancer
      selector:
        app: web
      ports:
        - port: 80           # port clients connect to
          targetPort: 8080   # port the container listens on
    What interviewers listen for
    • Stable virtual IP and DNS name for a set of pods
    • ClusterIP is the default and internal-only
    • NodePort opens a port (30000–32767) on every node
    • LoadBalancer provisions an external load balancer
    • ExternalName is just a DNS CNAME, no proxying

    Likely follow-up: Why would you put an Ingress in front instead of one LoadBalancer per service? · What is targetPort versus port?

  6. 6.How does a Deployment rolling update work? What do maxSurge and maxUnavailable control, and how do you roll back?mid

    When the pod template changes, the Deployment creates a new ReplicaSet and shifts pods over step by step: scale the new one up, wait for new pods to become available, scale the old one down, repeat.

    • maxSurge: how many pods may exist above the desired count during the rollout.
    • maxUnavailable: how many pods may be unavailable below the desired count.

    Both take a number or a percentage, both default to 25%, and they can't both be 0. maxSurge: 1, maxUnavailable: 0 keeps full capacity but needs room for an extra pod; a higher maxUnavailable is faster and needs no spare capacity.

    Readiness probes are what make this safe: a new pod that never becomes ready stalls the rollout instead of taking traffic. After progressDeadlineSeconds (600 by default) the Deployment reports ProgressDeadlineExceeded, but it does not roll back on its own.

    To roll back, check kubectl rollout history deployment/web, then run kubectl rollout undo deployment/web, optionally with --to-revision=N.

    spec:
      replicas: 4
      strategy:
        type: RollingUpdate
        rollingUpdate:
          maxSurge: 1         # at most 5 pods during the rollout
          maxUnavailable: 0   # never fewer than 4 available
    What interviewers listen for
    • New ReplicaSet scaled up, old one scaled down
    • maxSurge: extra pods; maxUnavailable: missing pods
    • Both default to 25%; cannot both be 0
    • Readiness gates progress; stalled rollouts are not auto-reverted
    • kubectl rollout undo, optionally --to-revision

    Likely follow-up: What is the Recreate strategy and when would you use it? · How would you pause a rollout halfway?

  7. 7.What's the difference between liveness, readiness and startup probes, and what happens when each one fails?mid

    They answer different questions, and failures have different consequences:

    • Readiness: can this container serve traffic right now? On failure the pod is marked not Ready and removed from Service endpoints, but it is not restarted. It runs for the container's whole life, so it also covers temporary overload.
    • Liveness: is the container stuck beyond recovery, like a deadlock? After failureThreshold consecutive failures (3 by default) the kubelet restarts the container.
    • Startup: has the app finished starting? Until it succeeds, liveness and readiness probes don't run, so slow starters aren't killed early. If it keeps failing for failureThreshold × periodSeconds, the container is restarted.

    Each probe can be httpGet (any status from 200 to 399 passes), tcpSocket, exec or grpc.

    A classic mistake is a liveness check that tests the database. If the database blips, every pod restarts at once. Liveness should only check the process itself.

    startupProbe:
      httpGet: { path: /healthz, port: 8080 }
      failureThreshold: 30
      periodSeconds: 10          # allows up to 300s to start
    livenessProbe:
      httpGet: { path: /healthz, port: 8080 }
    readinessProbe:
      httpGet: { path: /ready, port: 8080 }
      periodSeconds: 5
    What interviewers listen for
    • Readiness failure removes the pod from endpoints, no restart
    • Liveness failure restarts the container
    • Startup probe holds off the other probes until it passes
    • Mechanisms: httpGet, tcpSocket, exec, grpc
    • Don't check dependencies in liveness probes

    Likely follow-up: Why is initialDelaySeconds often replaced by a startup probe? · Do you need a readiness probe to drain traffic on shutdown?

  8. 8.How do ConfigMaps and Secrets differ, and what are the ways a Pod can consume them?easy

    Both are namespaced key-value objects that keep configuration out of the image.

    • ConfigMap: non-sensitive settings such as feature flags, URLs or whole config files.
    • Secret: sensitive data such as passwords, tokens and TLS keys. Values are base64-encoded, which is encoding, not encryption. Secrets can be locked down separately with RBAC, and the kubelet keeps mounted secrets in tmpfs (memory) rather than on disk.

    A pod can consume either one as:

    • Environment variables, one key with valueFrom or every key with envFrom
    • Files in a volume, one file per key
    • Command arguments built from those env vars

    The practical difference: env vars are read once at container start, so changes need a restart. Mounted volumes are eventually updated by the kubelet (except with subPath), so an app that re-reads its files can pick up changes. Both objects are limited to 1 MiB and can be marked immutable: true.

    # in the container spec
    env:
      - name: DB_PASSWORD
        valueFrom:
          secretKeyRef: { name: db-creds, key: password }
    envFrom:
      - configMapRef: { name: app-config }
    volumeMounts:
      - { name: config, mountPath: /etc/app }
    # in the pod spec
    volumes:
      - name: config
        configMap: { name: app-config }
    What interviewers listen for
    • ConfigMap for plain config, Secret for sensitive data
    • Secret values are base64-encoded, not encrypted
    • Consume as env vars or as mounted files
    • Env vars need a restart; volumes update eventually
    • 1 MiB limit; immutable: true available

    Likely follow-up: How would you roll pods automatically when a ConfigMap changes?

  9. 9.Are Kubernetes Secrets actually secure? Why aren't they encrypted by default, and how would you harden them?mid

    Not by default. Secret values are only base64-encoded, and the API server stores them unencrypted in etcd. Anyone with access to etcd, an etcd backup, or get/list/watch on Secrets can read them. And anyone who can create a pod in a namespace can mount any Secret in it, so "create pods" effectively implies reading Secrets.

    Encryption at rest isn't on by default because it needs keys that someone has to generate, store and rotate, which is an operator decision. Many managed services configure it for you.

    To harden:

    • Enable encryption at rest with an EncryptionConfiguration passed to the API server via --encryption-provider-config, ideally the KMS v2 provider so the key lives in an external KMS. Then rewrite existing Secrets so they're stored encrypted.
    • Use least-privilege RBAC; remember list and watch reveal values too.
    • Protect etcd: TLS, restricted access, encrypted backups.
    • Consider an external secret store (Vault, a cloud secret manager) via the Secrets Store CSI Driver or External Secrets Operator.
    What interviewers listen for
    • base64 is encoding; stored unencrypted in etcd by default
    • Pod-create permission implies Secret read in the namespace
    • Enable encryption at rest, preferably with KMS v2
    • Least-privilege RBAC; list/watch expose values too
    • Protect etcd and its backups; consider external stores

    Likely follow-up: After enabling encryption at rest, why are old Secrets still in plaintext? · How does the KMS provider use envelope encryption?

  10. 10.When would you use a StatefulSet instead of a Deployment, and what guarantees does it give you?mid

    Use a StatefulSet when each replica has its own identity and data: databases, Kafka, ZooKeeper, Elasticsearch. It adds three guarantees a Deployment doesn't give:

    • Stable names: pods are db-0, db-1, db-2, and a replacement keeps the same name.
    • Stable network identity: through the headless Service named in serviceName, each pod gets its own DNS record, like db-0.db-headless.default.svc.cluster.local, so peers can find each other.
    • Stable storage: volumeClaimTemplates create one PVC per pod, and a rescheduled pod reattaches to the same volume.

    With the default OrderedReady policy, pods are created in order from 0 up, each waiting for the previous one to be Running and Ready, and removed in reverse. Rolling updates go from the highest ordinal down, and partition allows staged updates.

    Gotchas: by default, deleting or scaling down a StatefulSet does not delete its PVCs, and you have to create the headless Service yourself.

    apiVersion: apps/v1
    kind: StatefulSet
    metadata: { name: db }
    spec:
      serviceName: db-headless    # headless Service you create yourself
      replicas: 3
      selector: { matchLabels: { app: db } }
      template:
        metadata: { labels: { app: db } }
        spec: { containers: [{ name: db, image: postgres:17 }] }
      volumeClaimTemplates:       # one PVC per pod: data-db-0, data-db-1...
        - metadata: { name: data }
          spec: { accessModes: [ReadWriteOnce], resources: { requests: { storage: 10Gi } } }
    What interviewers listen for
    • For stateful apps where replicas are not interchangeable
    • Stable ordinal names: db-0, db-1, …
    • Per-pod DNS via a headless Service
    • Per-pod PVCs from volumeClaimTemplates survive rescheduling
    • Ordered create, delete and update by default

    Likely follow-up: How would you do a canary update of a StatefulSet? · What does podManagementPolicy: Parallel change?

  11. 11.What are resource requests and limits, and how does Kubernetes use each of them?easy

    They're set per container, for CPU and memory:

    • Requests are what the container is guaranteed. The scheduler uses them: a pod only lands on a node whose unreserved allocatable capacity covers the sum of its requests. Requests also decide how CPU is shared under contention.
    • Limits are the ceiling, enforced by the kernel through cgroups on the node. A container that hits its CPU limit is throttled; one that exceeds its memory limit is OOM-killed.

    CPU is measured in cores (250m is a quarter core) and memory in bytes (Mi, Gi). If you set only a limit, the request defaults to the limit. Together, requests and limits decide the pod's QoS class.

    In practice, always set requests, because without them scheduling and autoscaling are guesswork. Set a memory limit to contain leaks. Many teams skip CPU limits for latency-sensitive services to avoid throttling. LimitRange can inject defaults and ResourceQuota can cap totals per namespace.

    resources:
      requests:
        cpu: 250m        # a quarter of a CPU core
        memory: 256Mi
      limits:
        cpu: "1"
        memory: 512Mi
    What interviewers listen for
    • Requests drive scheduling and are guaranteed
    • Limits are hard ceilings enforced by cgroups
    • Over CPU limit: throttled; over memory limit: OOM-killed
    • Only a limit set: request defaults to the limit
    • Requests and limits determine the QoS class

    Likely follow-up: Should you set CPU limits at all? · What do LimitRange and ResourceQuota do?

  12. 12.A pod is in CrashLoopBackOff. What does that mean, and how do you troubleshoot it?mid

    CrashLoopBackOff isn't a phase or a root cause. It means a container keeps exiting, and the kubelet is waiting before restarting it again. The delay grows exponentially (10s, 20s, 40s…) up to five minutes, and resets after the container runs cleanly for 10 minutes.

    My steps:

    • kubectl describe pod for the Last State: the reason and exit code, plus events. Exit code 1 is usually an app error; 137 means SIGKILL, often OOMKilled (the reason says so); 127 or 126 point at a bad command or entrypoint.
    • kubectl logs --previous to read the logs of the instance that crashed, since the current one may have just started.
    • Check config: missing env vars, a Secret or ConfigMap key that doesn't exist, wrong args, a database it can't reach.
    • Check probes: a too-aggressive liveness probe kills a healthy but slow app; the events show "Liveness probe failed".
    • If the image has a shell, override the command with sleep, or use kubectl debug to inspect it.
    kubectl get pod web-7d9c -o wide
    kubectl describe pod web-7d9c          # Last State, Reason, Exit Code, Events
    kubectl logs web-7d9c --previous       # output of the crashed instance
    kubectl logs web-7d9c -c migrate       # a specific container
    kubectl events --for pod/web-7d9c
    What interviewers listen for
    • Container keeps exiting; kubelet backs off restarts
    • Backoff doubles from 10s, capped at 5 minutes
    • describe shows exit code and reason; logs --previous
    • Exit 137 often OOMKilled; exit 1 app error
    • Check config, dependencies and liveness probes

    Likely follow-up: How do you debug a container that crashes before you can exec into it?

  13. 13.What is an Ingress, and why does it need an ingress controller?easy

    An Ingress is a set of layer-7 HTTP(S) routing rules: send shop.example.com/api to the api Service and everything else to web, and terminate TLS with a given certificate Secret. pathType can be Exact, Prefix or ImplementationSpecific.

    The Ingress object alone does nothing. An ingress controller (Traefik, HAProxy, an NGINX-based controller, or a cloud one like the AWS Load Balancer Controller) watches Ingress objects and configures a real proxy or load balancer to match. ingressClassName picks which controller handles it. Usually the controller itself sits behind one LoadBalancer Service, so dozens of services share a single external IP instead of paying for one load balancer each.

    Worth knowing: the Ingress API is frozen. It's GA and isn't going away, but it won't get new features, and the Kubernetes project now recommends Gateway API. The popular community ingress-nginx controller was retired in March 2026.

    apiVersion: networking.k8s.io/v1
    kind: Ingress
    metadata: { name: shop }
    spec:
      ingressClassName: nginx
      tls:
        - { hosts: [shop.example.com], secretName: shop-tls }
      rules:
        - host: shop.example.com
          http:
            paths:
              - path: /api
                pathType: Prefix
                backend: { service: { name: api, port: { number: 80 } } }
    What interviewers listen for
    • Host- and path-based HTTP routing plus TLS termination
    • Needs a controller to actually implement the rules
    • ingressClassName selects the controller
    • Many services share one external load balancer
    • Ingress API is frozen; Gateway API recommended

    Likely follow-up: How is Gateway API different from Ingress? · Where would you terminate TLS?

  14. 14.What are namespaces used for, and are they a security boundary?easy

    Namespaces divide one cluster into named scopes. Names must be unique within a namespace, not across the cluster, so dev and prod can each have a web Deployment. They're also the unit that most policies attach to:

    • RBAC Roles and RoleBindings grant access per namespace
    • ResourceQuota and LimitRange cap and default resource usage
    • NetworkPolicies and Pod Security labels are applied per namespace

    A new cluster starts with default, kube-system, kube-public and kube-node-lease. Some resources aren't namespaced at all: Nodes, PersistentVolumes, StorageClasses, ClusterRoles and Namespaces themselves (kubectl api-resources --namespaced=false lists them). Services in other namespaces are reached by DNS as name.namespace.

    On their own they're a soft boundary. Pods in different namespaces share nodes and kernels, and can talk to each other over the network unless NetworkPolicies say otherwise. For real multi-tenancy you add RBAC, NetworkPolicies, quotas and Pod Security Standards, or use separate clusters.

    What interviewers listen for
    • Scope for names; unique within a namespace
    • Unit for RBAC, quotas, LimitRanges, NetworkPolicies
    • Some resources are cluster-scoped (nodes, PVs, ClusterRoles)
    • Cross-namespace DNS: service.namespace
    • Soft isolation only, without extra policies

    Likely follow-up: How would you stop one team from using all the cluster resources?

  15. 15.What is the difference between labels and annotations, and how do selectors use labels?easy

    Both are key-value metadata on objects, but they serve different purposes.

    Labels identify and group objects, like app=web, tier=frontend, env=prod. They're short (values up to 63 characters) and they're what the system selects on. A Service finds its pods by label, a ReplicaSet counts pods by label, a NetworkPolicy targets pods by label, and kubectl get pods -l app=web filters by label.

    Selectors come in two styles: equality-based (app=web, env!=prod) and set-based (env in (prod,staging), !canary). In manifests that's matchLabels and matchExpressions with In, NotIn, Exists and DoesNotExist.

    Annotations hold non-identifying data that you can't select on and that can be much larger: build info, a git SHA, contact details, or tool configuration such as ingress controller settings or Prometheus scrape hints. kubectl apply stores its last-applied configuration in one.

    A classic bug is a Service selector that doesn't match the pod labels: the Service exists but has no endpoints.

    What interviewers listen for
    • Labels identify and group; selectors query them
    • Services, ReplicaSets, NetworkPolicies select by labels
    • Equality-based and set-based selectors
    • Annotations: non-identifying metadata, not selectable
    • Mismatched selector means a Service with no endpoints
  16. 16.What is a DaemonSet, and what would you use one for?easy

    A DaemonSet runs one copy of a pod on every node, or on every node matching a selector. When a node joins the cluster it gets the pod automatically, and when a node is removed its pod is garbage-collected. You don't set a replica count.

    Typical uses are node-level agents:

    • Log collectors such as Fluent Bit or Fluentd
    • Monitoring agents such as Prometheus node-exporter
    • Networking: CNI agents and kube-proxy itself
    • Storage and security daemons, like CSI node plugins

    You limit it to some nodes with nodeSelector or node affinity, and add tolerations if it must run on tainted nodes such as the control plane. The controller also adds some tolerations automatically, which is why DaemonSet pods can run on cordoned nodes, and why kubectl drain needs --ignore-daemonsets.

    Updates use RollingUpdate by default, one node at a time unless you raise maxUnavailable, or OnDelete, where pods only update when you delete them.

    What interviewers listen for
    • One pod per eligible node, no replica count
    • New nodes get the pod automatically
    • Used for log, monitoring, network and storage agents
    • Target nodes with nodeSelector, affinity and tolerations
    • RollingUpdate or OnDelete update strategies

    Likely follow-up: How is a DaemonSet different from a Deployment with pod anti-affinity?

  17. 17.How does the Horizontal Pod Autoscaler decide how many replicas to run?mid

    The HPA is a control loop in the controller manager, running every 15 seconds by default. It reads metrics, computes a replica count and updates the target's scale subresource:

    desired = ceil(current × currentMetric / targetMetric)

    So 4 pods averaging 90% CPU against a 70% target gives ceil(4 × 90/70) = 6. It skips changes within a 10% tolerance, clamps the result to minReplicas/maxReplicas, and uses a 5-minute scale-down stabilization window by default so it doesn't flap.

    Key details:

    • CPU and memory come from the Metrics API, usually metrics-server. Custom and external metrics need an adapter, such as Prometheus Adapter or KEDA.
    • Utilization is relative to requests, so containers without requests can't be scaled on utilization.
    • The behavior field tunes scale-up and scale-down rates.
    • Don't hard-code replicas in a Deployment the HPA manages, or every kubectl apply fights it.
    apiVersion: autoscaling/v2
    kind: HorizontalPodAutoscaler
    metadata: { name: web }
    spec:
      scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: web }
      minReplicas: 2
      maxReplicas: 10
      metrics:
        - type: Resource
          resource:
            name: cpu
            target: { type: Utilization, averageUtilization: 70 }
    What interviewers listen for
    • ceil(current × currentMetric / target)
    • Needs metrics-server; adapters for custom metrics
    • Utilization is a percentage of the resource request
    • Tolerance and 5-minute scale-down stabilization prevent flapping
    • Bounded by minReplicas and maxReplicas

    Likely follow-up: How would you scale on queue length instead of CPU? · Can HPA and VPA target the same Deployment?

  18. 18.Explain PersistentVolumes, PersistentVolumeClaims and StorageClasses, and how dynamic provisioning works.mid

    They separate what storage an app needs from how it's provided:

    • A PersistentVolume (PV) is a cluster-scoped piece of real storage, like a cloud disk or an NFS export, whose lifecycle is independent of any pod.
    • A PersistentVolumeClaim (PVC) is a namespaced request: "20Gi, ReadWriteOnce, class fast-ssd". It binds one-to-one to a matching PV, and pods mount the claim, not the PV.
    • A StorageClass describes a kind of storage: its CSI provisioner, parameters, reclaimPolicy and volumeBindingMode.

    With dynamic provisioning, a PVC that names a class (or falls back to the default class) makes the provisioner create a PV on demand, so no admin pre-creates volumes. WaitForFirstConsumer delays that until a pod is scheduled, so the disk is created in the pod's zone; the default is Immediate.

    Access modes are ReadWriteOnce (one node), ReadOnlyMany, ReadWriteMany and ReadWriteOncePod (one pod). Dynamically provisioned PVs default to the Delete reclaim policy; use Retain for data you can't lose.

    apiVersion: v1
    kind: PersistentVolumeClaim
    metadata: { name: data }
    spec:
      accessModes: [ReadWriteOnce]
      storageClassName: fast-ssd
      resources: { requests: { storage: 20Gi } }
    ---
    apiVersion: storage.k8s.io/v1
    kind: StorageClass
    metadata: { name: fast-ssd }
    provisioner: ebs.csi.aws.com
    volumeBindingMode: WaitForFirstConsumer
    reclaimPolicy: Retain
    What interviewers listen for
    • PV is real storage; PVC is a namespaced request
    • Pods mount PVCs; a PVC binds to one PV
    • StorageClass enables dynamic provisioning via a CSI driver
    • WaitForFirstConsumer avoids zone mismatches
    • Access modes RWO/ROX/RWX/RWOP; reclaim Delete or Retain

    Likely follow-up: What does ReadWriteOnce actually guarantee? · How do you resize a PVC?

  19. 19.How does traffic sent to a Service ClusterIP actually reach a pod? What does kube-proxy do?mid

    A ClusterIP is virtual: in the default mode no network interface owns it. Three pieces make it work:

    • The EndpointSlice controller keeps a list of the IPs and ports of the ready pods matching the Service's selector.
    • kube-proxy, running on every node, watches Services and EndpointSlices and programs packet-forwarding rules in the node's kernel.
    • When a pod connects to ClusterIP:port, those rules on the client's node DNAT the packet to one backend pod IP, picked at random (or by sessionAffinity: ClientIP). Connection tracking keeps the rest of that connection on the same pod, and the pod network carries it to whichever node the pod is on.

    So kube-proxy isn't in the data path; the kernel does the forwarding. On Linux the default mode is iptables, nftables is the newer replacement and is expected to become the default, and IPVS mode is deprecated. Some CNIs, such as Cilium, replace kube-proxy with eBPF.

    A consequence: balancing is per connection, not per request, so long-lived HTTP/2 or gRPC connections can pile onto a few pods.

    What interviewers listen for
    • ClusterIP is virtual; EndpointSlices list ready pods
    • kube-proxy programs kernel rules on every node
    • DNAT to a backend pod on the client node
    • Modes: iptables default, nftables newer, IPVS deprecated
    • Load balancing is per connection, not per request

    Likely follow-up: Why might gRPC traffic be unevenly balanced across pods? · What happens to traffic when a pod fails its readiness probe?

  20. 20.Walk me through how an HTTPS request from a user on the internet reaches a container in your cluster.hard

    A typical path, with an ingress controller or gateway in front:

    • DNS resolves shop.example.com to the external load balancer the cloud created for the controller's LoadBalancer Service.
    • The cloud load balancer forwards to a node's NodePort, or straight to pod IPs if it supports IP-target mode.
    • On the node, kube-proxy rules route it to an ingress controller pod, possibly on another node. With externalTrafficPolicy: Local only nodes running a controller pod receive traffic, which also preserves the client IP.
    • The ingress controller terminates TLS (unless the load balancer already did), matches host and path, and picks a backend. Many controllers send traffic directly to pod IPs from the Service's EndpointSlices rather than through the ClusterIP.
    • The pod network (the CNI plugin) delivers the packet to the target pod's IP. Pod-to-pod traffic needs no NAT.
    • The container receives it on its containerPort. Only ready pods are listed as endpoints, which is why readiness probes matter for zero-downtime deploys.

    The response returns along the same connection path.

    What interviewers listen for
    • DNS to cloud load balancer to node or pod
    • kube-proxy rules or IP mode reach the ingress pods
    • Ingress controller terminates TLS and routes host/path
    • Controllers often target pod IPs from EndpointSlices
    • CNI delivers to the pod; only ready pods get traffic

    Likely follow-up: Where can you lose the client source IP along this path? · What changes with a service mesh?

  21. 21.What are the three Pod QoS classes, how is a pod’s class decided, and why does it matter?mid

    You don't set the QoS class; Kubernetes derives it from requests and limits and records it in status.qosClass:

    • Guaranteed: every container has CPU and memory requests and limits, and each request equals its limit.
    • Burstable: not Guaranteed, but at least one container has a CPU or memory request or limit.
    • BestEffort: no container has any requests or limits.

    It matters when a node runs short of memory. The kubelet sets each container's oom_score_adj from its class (-997 for Guaranteed, 1000 for BestEffort, in between for Burstable), so the kernel OOM killer picks BestEffort processes first.

    For node-pressure eviction, the kubelet ranks pods by whether their usage exceeds their requests, then by priority, then by how far over requests they are. It doesn't use the QoS class directly, but the result is roughly BestEffort first and Guaranteed last.

    So critical workloads should have honest memory requests, ideally Guaranteed. Guaranteed pods with whole-number CPU requests can also get exclusive cores under the static CPU manager policy.

    What interviewers listen for
    • Derived from requests and limits, not set directly
    • Guaranteed: requests equal limits for CPU and memory
    • BestEffort: no requests or limits at all
    • Drives OOM score; BestEffort is killed first
    • Eviction ranks by usage over requests, then priority

    Likely follow-up: Can a pod with only CPU requests be Guaranteed?

  22. 22.One container keeps getting OOMKilled, and another is slow even though its average CPU usage looks low. What is going on in each case?mid

    They're the two ways limits bite, because memory is incompressible and CPU is compressible.

    OOMKilled: when a container goes over its memory limit, the kernel OOM killer kills a process in its cgroup. The status shows Reason: OOMKilled and exit code 137, and the container restarts, often into CrashLoopBackOff. A pod can also be killed or evicted under its limit if the whole node runs out of memory. Fix it by sizing the limit from real usage (kubectl top, working-set metrics, VPA recommendations), fixing leaks, and making runtimes container-aware, such as the JVM's -XX:MaxRAMPercentage.

    CPU throttling: a CPU limit is enforced as a CFS quota: CPU time per scheduling period, usually 100 ms. A multi-threaded app can burn its whole quota early in a period and then stall until the next one. Average usage looks low while tail latency spikes. The throttling counters in cAdvisor metrics show it. Fix it by raising or removing the CPU limit while keeping an accurate request, or by using fewer threads.

    What interviewers listen for
    • Memory over limit: OOM-killed, exit code 137
    • Node-level memory pressure can kill pods too
    • CPU over limit: throttled by CFS quota, not killed
    • Throttling causes latency spikes despite low averages
    • Size from real usage; consider dropping CPU limits

    Likely follow-up: How would you confirm throttling with metrics? · Why might a JVM app get OOMKilled with a heap well below the limit?

  23. 23.A pod has been stuck in Pending for ten minutes. How do you find out why?mid

    Pending means the pod was accepted but isn't running yet. Usually it hasn't been scheduled; sometimes it's scheduled and still pulling images. kubectl describe pod tells you which: a FailedScheduling event from the scheduler explains why each group of nodes was rejected.

    Common causes:

    • Insufficient CPU or memory: no node has enough unreserved allocatable capacity for the pod's requests. Actual usage doesn't matter. Lower the requests or add nodes; with a cluster autoscaler, look for scale-up events.
    • Taints the pod doesn't tolerate, or a nodeSelector or affinity that matches no node.
    • Anti-affinity or topology spread rules that can't be satisfied, for example more replicas than nodes with required per-node anti-affinity.
    • An unbound PVC: no default StorageClass, no matching PV, or a volume in a different zone.

    If the pod already has a node assigned, look at container states instead: it's probably pulling a large image or waiting on a volume mount. And if a ResourceQuota is the problem, the pod is never created at all, so check the ReplicaSet's events.

    kubectl describe pod api-5f7b      # read the Events at the bottom
    # e.g. FailedScheduling: 0/6 nodes are available: 3 Insufficient cpu,
    #      3 node(s) had untolerated taint {dedicated: gpu}
    kubectl describe node worker-1     # Allocated resources, Taints
    kubectl get pvc -n shop            # an unbound claim blocks scheduling
    What interviewers listen for
    • kubectl describe pod and the FailedScheduling event
    • Scheduling uses requests, not actual usage
    • Taints, selectors and affinity can exclude every node
    • Unbound PVCs keep pods Pending
    • Quota failures show on the ReplicaSet, not the pod

    Likely follow-up: How does the cluster autoscaler react to Pending pods?

  24. 24.What does ImagePullBackOff mean, and what are the usual causes?easy

    The kubelet couldn't pull the container image. You first see ErrImagePull, then ImagePullBackOff while it keeps retrying with an increasing delay, capped at five minutes. The container never starts, so there are no logs; the answer is in kubectl describe pod events, which show the registry's actual error.

    Usual causes:

    • A typo in the image name, or a tag that doesn't exist ("manifest unknown", "not found").
    • A private registry without credentials ("unauthorized"). Create a docker-registry Secret and reference it in imagePullSecrets on the pod or its ServiceAccount, or grant the nodes registry access through cloud IAM.
    • Registry rate limits, like Docker Hub's anonymous pull limits.
    • Network problems from the node: DNS, egress firewall, proxy ("i/o timeout").
    • Wrong architecture: the image has no variant for the node's platform, such as arm64.

    To avoid surprises, pin images by tag or digest instead of :latest, and remember that :latest with no explicit imagePullPolicy means Always.

    What interviewers listen for
    • Kubelet cannot pull the image; retries with backoff
    • Check kubectl describe pod events for the registry error
    • Typos, missing tags, missing imagePullSecrets
    • Rate limits, network or architecture mismatches
    • Pin tags or digests instead of :latest
  25. 25.How does RBAC work in Kubernetes? Explain Role, ClusterRole, RoleBinding and ClusterRoleBinding.mid

    RBAC decides whether a subject may perform a verb on a resource. It's purely additive: there are no deny rules, so anything not granted by some binding is refused.

    • Role: a list of rules (API groups, resources, verbs) inside one namespace.
    • ClusterRole: the same, but not namespaced. Use it for cluster-scoped resources like nodes, for non-resource URLs, or as a reusable template.
    • RoleBinding: grants a Role or a ClusterRole to subjects (users, groups, ServiceAccounts) within one namespace. Binding a ClusterRole this way limits it to that namespace; that's how the built-in view, edit and admin roles are usually handed out.
    • ClusterRoleBinding: grants a ClusterRole across the whole cluster.

    Kubernetes has no User objects; users and groups come from authentication, such as certificates or OIDC. A binding's roleRef can't be changed; you recreate the binding. Test with kubectl auth can-i list pods -n dev --as=system:serviceaccount:dev:ci-bot, and avoid wildcards, cluster-admin, and broad access to Secrets or pods/exec.

    apiVersion: rbac.authorization.k8s.io/v1
    kind: Role
    metadata: { name: pod-reader, namespace: dev }
    rules:
      - apiGroups: [""]              # "" is the core API group
        resources: [pods, pods/log]
        verbs: [get, list, watch]
    ---
    apiVersion: rbac.authorization.k8s.io/v1
    kind: RoleBinding
    metadata: { name: ci-reads-pods, namespace: dev }
    subjects: [{ kind: ServiceAccount, name: ci-bot, namespace: dev }]
    roleRef: { kind: Role, name: pod-reader, apiGroup: rbac.authorization.k8s.io }
    What interviewers listen for
    • Additive only; no deny rules
    • Role is namespaced; ClusterRole is cluster-wide
    • RoleBinding can bind a ClusterRole within one namespace
    • Subjects: users, groups, ServiceAccounts
    • Verify with kubectl auth can-i --as

    Likely follow-up: Why is permission to create pods close to permission to read Secrets? · What do the escalate and bind verbs protect against?

  26. 26.What is a ServiceAccount, and how does a pod authenticate to the Kubernetes API?mid

    A ServiceAccount is an identity for workloads, as opposed to human users. It's namespaced, every namespace gets one called default, and a pod uses default unless you set serviceAccountName.

    By default the kubelet mounts a short-lived, automatically rotated token into the pod at /var/run/secrets/kubernetes.io/serviceaccount/, together with the cluster CA and the namespace name. The token comes from the TokenRequest API and is bound to the pod, so it stops working once the pod is deleted. Since v1.24, Kubernetes no longer auto-creates long-lived token Secrets for each ServiceAccount.

    Client libraries use this "in-cluster config": they find the API server through the KUBERNETES_SERVICE_HOST variable and send the token as a bearer token. The API server authenticates it as system:serviceaccount:<namespace>:<name> and RBAC authorizes it.

    Good practice: one ServiceAccount per app with minimal RBAC, no permissions on default, and automountServiceAccountToken: false for pods that never call the API. Cloud workload identity features map ServiceAccounts to cloud IAM roles.

    What interviewers listen for
    • Namespaced identity for pods; default if unspecified
    • Projected, short-lived, auto-rotated token
    • No auto-created long-lived token Secrets since v1.24
    • Authenticated as system:serviceaccount:ns:name, then RBAC
    • Disable automount when the API is not needed

    Likely follow-up: How would a pod get AWS or GCP credentials without static keys?

  27. 27.What are Jobs and CronJobs, and which settings control retries, parallelism and overlapping runs?easy

    A Job runs pods to completion, unlike a Deployment, which keeps pods running forever. It's done when enough pods succeed:

    • completions and parallelism control how many successes you need and how many pods run at once. completionMode: Indexed gives each pod an index for sharded work.
    • backoffLimit sets how many failures are retried before the Job is marked failed (default 6). activeDeadlineSeconds caps total runtime.
    • ttlSecondsAfterFinished cleans up finished Jobs automatically.
    • The pod's restartPolicy must be OnFailure or Never.

    A CronJob creates Jobs on a cron schedule. concurrencyPolicy decides what happens if the previous run is still going: Allow (the default), Forbid (skip the new run) or Replace (cancel the old one). You can also set startingDeadlineSeconds for late starts, history limits (3 successful and 1 failed Job kept by default), suspend and timeZone.

    A CronJob creates a Job roughly once per scheduled time, but occasionally two or none, so jobs should be idempotent.

    apiVersion: batch/v1
    kind: CronJob
    metadata: { name: nightly-report }
    spec:
      schedule: "0 2 * * *"         # 02:00 every day
      timeZone: "Etc/UTC"
      concurrencyPolicy: Forbid     # skip a run if the last one is still going
      jobTemplate:
        spec:
          backoffLimit: 3
          template:
            spec:
              restartPolicy: OnFailure
              containers: [{ name: report, image: example/report:1.4 }]
    What interviewers listen for
    • Job: run to completion; Deployment: run forever
    • completions, parallelism, backoffLimit (default 6)
    • Pod restartPolicy must be OnFailure or Never
    • CronJob concurrencyPolicy: Allow, Forbid, Replace
    • Scheduling is approximate, so make jobs idempotent

    Likely follow-up: How would you run a database migration before a deployment?

  28. 28.How do taints and tolerations work, and how would you dedicate a pool of GPU nodes to specific workloads?mid

    A taint on a node repels pods; a toleration on a pod lets it ignore a matching taint. That's the opposite of affinity: affinity attracts, taints repel.

    Taint effects:

    • NoSchedule: pods without a matching toleration won't be scheduled there; running pods stay.
    • PreferNoSchedule: a soft version; the scheduler tries to avoid the node.
    • NoExecute: also evicts running pods that don't tolerate it. tolerationSeconds lets a pod stay for a limited time.

    Tolerations match with operator: Equal (key, value and effect) or Exists (any value).

    The key point for dedicated nodes: a toleration only allows a pod onto a node, it doesn't send it there. So taint the GPU nodes to keep everything else off, and give the GPU workloads both the toleration and a nodeSelector or node affinity for those nodes.

    Kubernetes uses taints itself: the node controller adds not-ready and unreachable NoExecute taints, and pods get default 300-second tolerations for them, which is why pods on a dead node are evicted after about five minutes.

    # kubectl taint nodes gpu-1 dedicated=gpu:NoSchedule
    tolerations:
      - key: dedicated
        operator: Equal
        value: gpu
        effect: NoSchedule
    nodeSelector:
      pool: gpu    # a toleration permits the node; this attracts the pod
    What interviewers listen for
    • Taints repel; tolerations allow, they do not attract
    • NoSchedule, PreferNoSchedule, NoExecute
    • NoExecute evicts running pods; tolerationSeconds delays it
    • Dedicated nodes: taint plus toleration plus node selector
    • Default 300s tolerations for not-ready and unreachable

    Likely follow-up: Why do DaemonSet pods still run on cordoned nodes?

  29. 29.Compare nodeSelector, node affinity and pod anti-affinity. How would you keep replicas of a service off the same node?mid

    All three constrain where the scheduler places a pod:

    • nodeSelector: the simplest. The node must have all the listed labels.
    • Node affinity: the same idea with operators (In, NotIn, Exists, DoesNotExist, Gt, Lt) and two strengths: required… (hard) and preferred… (weighted, soft). "IgnoredDuringExecution" means pods already running aren't evicted if node labels change later.
    • Pod affinity and anti-affinity: rules relative to other pods' labels within a topology domain set by topologyKey, such as kubernetes.io/hostname for a node or topology.kubernetes.io/zone for a zone. Affinity co-locates, say a cache next to its app; anti-affinity spreads.

    To keep replicas apart, add pod anti-affinity on the app's own label with topologyKey: kubernetes.io/hostname. Prefer the preferred form: with the required form, replicas beyond the number of nodes stay Pending. Inter-pod affinity is also expensive in large clusters, and topologySpreadConstraints is often the better tool for even spreading.

    affinity:
      nodeAffinity:
        requiredDuringSchedulingIgnoredDuringExecution:
          nodeSelectorTerms:
            - matchExpressions:
                - { key: topology.kubernetes.io/zone, operator: In, values: [eu-west-1a, eu-west-1b] }
      podAntiAffinity:
        preferredDuringSchedulingIgnoredDuringExecution:
          - weight: 100
            podAffinityTerm:
              labelSelector: { matchLabels: { app: web } }
              topologyKey: kubernetes.io/hostname
    What interviewers listen for
    • nodeSelector: exact label match
    • Node affinity: operators plus required or preferred rules
    • Pod anti-affinity spreads relative to other pods
    • topologyKey defines node, zone or region domains
    • Required anti-affinity can leave extra replicas Pending

    Likely follow-up: How do topologySpreadConstraints differ from anti-affinity?

  30. 30.What is Helm, and what are charts, values and releases?easy

    Helm is the package manager for Kubernetes. It saves you from hand-editing near-identical YAML for every environment.

    • A chart is a package: templated manifests written in Go templates, a Chart.yaml with metadata, and a values.yaml with defaults. Charts can depend on other charts and are shared through repositories or OCI registries.
    • Values are the configuration you supply at install time, with -f values-prod.yaml or --set image.tag=1.4, so one chart serves dev, staging and prod.
    • A release is an installed instance of a chart, with a name and a numbered revision history. Every upgrade creates a new revision, and helm rollback returns to an earlier one.

    Everyday commands are helm upgrade --install (idempotent, good for CI), helm rollback, helm list, and helm template to render manifests locally for review.

    Since Helm 3 there's no server-side Tiller: the CLI talks to the API server with your own credentials and stores release state in Secrets in the release's namespace. The main alternative is Kustomize, which uses template-free overlays and is built into kubectl as kubectl apply -k.

    What interviewers listen for
    • Package manager: templated, reusable manifests
    • Chart = templates + Chart.yaml + default values
    • Values customize a chart per environment
    • Release = installed instance with revision history
    • No Tiller since Helm 3; Kustomize is the alternative

    Likely follow-up: When would you pick Kustomize over Helm? · How do you manage secrets in Helm values?

  31. 31.What are init containers, and when would you use them?easy

    Init containers, listed under spec.initContainers, run before the app containers start. They run one at a time, in order, and each must exit successfully before the next one starts. If one fails, the kubelet retries it according to the pod's restartPolicy; with Never, the pod fails. While they run, kubectl get pods shows a status like Init:0/2.

    Typical uses:

    • Wait for a dependency, for example loop until the database's DNS name resolves.
    • Prepare files: fetch config, clone a repo or render a template into a shared emptyDir volume.
    • Fix permissions on a mounted volume.
    • One-off setup per pod. Schema migrations are usually better as a separate Job, because an init container runs once for every replica.

    The advantage is separation: setup tools like git or curl stay out of the app image, and the init container can have different permissions. Regular init containers don't support liveness, readiness or startup probes. An init container with restartPolicy: Always is a different thing: a native sidecar.

    What interviewers listen for
    • Run before app containers, sequentially, to completion
    • A failed init container is retried per restartPolicy
    • Used for waiting, file preparation and permissions
    • Keeps setup tools out of the app image
    • No probes; restartPolicy: Always makes it a sidecar

    Likely follow-up: Why is running migrations in an init container risky with many replicas?

  32. 32.When would you run more than one container in a Pod, and how do those containers communicate?easy

    Only when containers are tightly coupled: they must run on the same node, start and stop together, and scale together. The common patterns:

    • Sidecar: extends the main app, like a log shipper, a service-mesh proxy, or a config or certificate reloader.
    • Ambassador: a local proxy for outbound connections, such as a database or cloud SQL proxy on localhost.
    • Adapter: translates the app's output, for example an exporter that turns app stats into Prometheus metrics.

    Containers in a pod communicate cheaply because they share context:

    • The same network namespace: they talk over localhost and can't both bind the same port.
    • Shared volumes, usually an emptyDir, for exchanging files.
    • Optionally a shared process namespace (shareProcessNamespace: true), so they can see and signal each other's processes.

    The classic anti-pattern is putting an app and its database in one pod. They scale differently, a restart takes both down, and the database can't survive the pod. Those belong in separate workloads connected by a Service.

    What interviewers listen for
    • Only for tightly coupled, co-scaled containers
    • Sidecar, ambassador and adapter patterns
    • Shared network: communicate over localhost
    • Shared volumes such as emptyDir for files
    • Don't pack an app and its database together
  33. 33.What problems do native sidecar containers solve, and how do you declare one?mid

    Sidecars used to be just extra entries in containers, which caused three problems:

    • No startup order: the app could start before its proxy was ready, so early requests failed.
    • No shutdown order: the sidecar could exit before the app finished draining.
    • Jobs never finished: the Job waited for every container, and a log shipper or proxy never exits.

    Native sidecars, stable since v1.33, are init containers with restartPolicy: Always:

    • They start in order with the other init containers, and the kubelet moves on as soon as the sidecar has started, gated by its startup probe if it has one. So they're running before the app containers.
    • They keep running for the pod's lifetime and are restarted if they exit, whatever the pod's restart policy.
    • On termination they're stopped after the main containers, in reverse order.
    • They don't block Job completion.
    • They support probes, and a sidecar's readiness counts toward the pod's readiness.

    Service-mesh proxies and log or metrics agents are the main users.

    spec:
      initContainers:
        - name: log-shipper
          image: example/log-shipper:1.0
          restartPolicy: Always        # this makes it a sidecar
          volumeMounts: [{ name: logs, mountPath: /var/log/app }]
      containers:
        - name: app
          image: example/app:2.0
          volumeMounts: [{ name: logs, mountPath: /var/log/app }]
      volumes:
        - { name: logs, emptyDir: {} }
    What interviewers listen for
    • Declared as init containers with restartPolicy: Always
    • Start before app containers and keep running
    • Stopped after main containers on shutdown
    • Don't block Job completion
    • Stable since v1.33; support probes

    Likely follow-up: How did people stop sidecars in Jobs before this feature existed?

  34. 34.How do pods discover other services in the cluster? What DNS names does Kubernetes create?easy

    Through cluster DNS, usually CoreDNS running in kube-system behind a Service called kube-dns. It watches Services and serves records for them:

    • my-svc.my-ns.svc.cluster.local resolves to the Service's ClusterIP. cluster.local is the default cluster domain, but it's configurable.
    • Named ports get SRV records such as _http._tcp.my-svc.my-ns.svc.cluster.local.
    • A headless Service resolves to the IPs of its ready pods, and StatefulSet pods get names like db-0.db.my-ns.svc.cluster.local.

    The kubelet writes each pod's /etc/resolv.conf to point at the DNS Service and adds search domains, starting with my-ns.svc.cluster.local. That's why a pod can just call http://my-svc in its own namespace, and my-svc.other-ns for another namespace.

    One gotcha: the default ndots:5 makes short external names like api.example.com try every search domain first, which multiplies DNS queries. A trailing dot or a tuned dnsConfig avoids that.

    There's also an older mechanism: environment variables such as MY_SVC_SERVICE_HOST, but only for Services that existed when the pod started.

    What interviewers listen for
    • CoreDNS serves records for every Service
    • service.namespace.svc.cluster.local resolves to the ClusterIP
    • Search domains allow the short name within a namespace
    • Headless Services resolve to pod IPs
    • ndots:5 can multiply lookups for external names

    Likely follow-up: How would you debug a pod that cannot resolve a Service name?

  35. 35.What is a headless Service, and when do you need one?mid

    A headless Service sets clusterIP: None. It gets no virtual IP, kube-proxy ignores it, and nothing load-balances it. Instead, a DNS lookup of the Service name returns the IPs of all its ready pods directly.

    You need one when clients must see or address individual pods:

    • StatefulSets: the governing Service in serviceName gives each pod a stable name like db-0.db.shop.svc.cluster.local, so clients can reach the primary specifically, and cluster members like Kafka brokers or etcd peers can find each other.
    • Client-side load balancing: a gRPC client can resolve every pod IP and balance per request, which works around the per-connection balancing of a normal ClusterIP.
    • Discovery for clustered software that wants the full member list.

    publishNotReadyAddresses: true publishes pods before they're ready, which helps members find each other while a cluster is bootstrapping.

    The trade-off is that clients take on the work: they must re-resolve DNS as pods come and go, and not cache old IPs for too long.

    What interviewers listen for
    • clusterIP: None: no virtual IP, no kube-proxy
    • DNS returns the ready pod IPs directly
    • Gives StatefulSet pods stable per-pod DNS names
    • Enables client-side load balancing, e.g. gRPC
    • Clients must handle DNS changes and caching
  36. 36.How do NetworkPolicies work? How would you lock down traffic to an API so only the frontend can reach it?mid

    By default the pod network is flat: every pod can reach every other pod, across namespaces. A NetworkPolicy is a layer 3/4 firewall for pods, selected by labels.

    • It's enforced by the network plugin, such as Calico or Cilium. If the plugin doesn't support NetworkPolicy, the objects are accepted but do nothing.
    • A pod becomes isolated for ingress or egress as soon as any policy of that type selects it. From then on, only traffic some policy allows gets through. Policies are additive: there's no deny rule and no ordering. Replies to allowed connections are allowed automatically.
    • Peers are chosen with podSelector, namespaceSelector (both in one entry means AND) or ipBlock, plus ports.

    The usual pattern is a default-deny policy per namespace (podSelector: {} with no allow rules), then narrow allow policies like the one shown. If you also deny egress, remember to allow DNS to kube-dns on port 53, or everything breaks.

    For layer-7 rules, explicit denies or cluster-wide policies you need CNI-specific resources, such as Cilium's or Calico's own policy types.

    apiVersion: networking.k8s.io/v1
    kind: NetworkPolicy
    metadata: { name: api-allow-frontend, namespace: shop }
    spec:
      podSelector: { matchLabels: { app: api } }
      policyTypes: [Ingress]
      ingress:
        - from:
            - podSelector: { matchLabels: { app: frontend } }
          ports:
            - { protocol: TCP, port: 8080 }
    What interviewers listen for
    • Default: all pod traffic allowed
    • Enforced by the CNI; unsupported plugins ignore it
    • Selected pods become isolated; allowed traffic is a union
    • Start with default deny, then allow specific flows
    • Allow DNS egress when restricting egress

    Likely follow-up: What is the difference between namespaceSelector and podSelector in the same from entry versus separate entries?

  37. 37.Compare rolling, blue-green and canary deployments. How would you implement each on Kubernetes?mid
    • Rolling update: the Deployment default. Pods are replaced gradually under maxSurge and maxUnavailable. It needs no extra infrastructure, but both versions serve traffic during the rollout, so changes must be backward compatible, and a rollback is just another rollout. (Recreate is the blunt alternative: stop everything, then start the new version, with downtime.)
    • Blue-green: run the complete new version (green) next to the old one (blue), test it, then switch all traffic at once, for example by changing a Service selector from version: blue to version: green, or repointing a route. Rollback is an instant switch back. The cost is double capacity during the switch.
    • Canary: send a small share of real traffic to the new version, watch error rates and latency, then increase step by step. The crude way is two Deployments behind one Service, where traffic roughly follows the replica ratio. The precise way is weighted routing: Gateway API backendRefs weights, a service mesh, or ingress controller features.

    Tools like Argo Rollouts and Flagger automate blue-green and canary, including metric analysis and automatic rollback.

    What interviewers listen for
    • Rolling: gradual, built in, versions overlap
    • Blue-green: full parallel stack, instant switch and rollback
    • Canary: small traffic share, promote on healthy metrics
    • Weighted routing via Gateway API or a mesh
    • Argo Rollouts or Flagger automate analysis and rollback

    Likely follow-up: How do database schema changes affect these strategies?

  38. 38.Why is etcd so critical, and how do you back it up and restore it?hard

    etcd stores all cluster state: every object, including Secrets. Running containers keep going without it for a while, but nothing can be created, changed or rescheduled. Losing it without a backup means rebuilding the cluster.

    Backup: take a snapshot from a live member with etcdctl snapshot save, using the etcd client TLS certificates, or snapshot the etcd data volume. Check it with etcdutl snapshot status. Run it on a schedule, store copies off the cluster and encrypted (a snapshot holds your Secrets, in plaintext unless encryption at rest is on), and rehearse restores.

    Restore: stop all API servers, restore the snapshot into a new data directory on each etcd member with etcdutl snapshot restore, point etcd at that directory (on kubeadm, in the etcd static pod manifest), start etcd, then the API servers. Restarting the scheduler, controller manager and kubelets is recommended so nothing acts on stale data.

    On managed services like EKS, GKE or AKS you can't reach etcd. You back up at the resource level instead, with Velero or GitOps, plus volume snapshots.

    # kubeadm certificate paths shown
    ETCDCTL_API=3 etcdctl --endpoints=https://127.0.0.1:2379 \
      --cacert=/etc/kubernetes/pki/etcd/ca.crt \
      --cert=/etc/kubernetes/pki/etcd/server.crt \
      --key=/etc/kubernetes/pki/etcd/server.key \
      snapshot save /backup/etcd-snapshot.db
    
    etcdutl --write-out=table snapshot status /backup/etcd-snapshot.db
    
    # restore into a NEW data dir, with all API servers stopped
    etcdutl --data-dir /var/lib/etcd-restored snapshot restore /backup/etcd-snapshot.db
    What interviewers listen for
    • etcd holds all cluster state, including Secrets
    • etcdctl snapshot save with TLS certificates
    • Store encrypted and off-cluster; test restores
    • Restore with etcdutl into a new data dir, API servers stopped
    • Managed clusters: Velero or GitOps instead

    Likely follow-up: Why should an etcd cluster have an odd number of members?

  39. 39.How does the kube-scheduler decide which node a pod runs on?mid

    The scheduler watches for pods with no spec.nodeName and handles each in two phases:

    • Filtering removes nodes that can't run the pod: not enough unreserved allocatable CPU or memory for its requests, a nodeSelector or required affinity that doesn't match, taints it doesn't tolerate, host port conflicts, volume zone constraints, required pod anti-affinity, unschedulable nodes.
    • Scoring ranks the remaining nodes with scoring plugins: spreading replicas, balancing resource usage, preferred affinity weights, topology spread, whether the image is already on the node.

    It picks the highest-scoring node, breaking ties randomly, and binds the pod by setting its node through the API server. The kubelet on that node sees the assignment and starts the containers.

    If no node passes filtering, the pod stays Pending with a FailedScheduling event. A higher-priority pod, set through a PriorityClass, can trigger preemption, evicting lower-priority pods to make room.

    The scheduler is built from plugins, supports profiles and can run alongside other schedulers (schedulerName). It only places pods once; it never moves running pods. The separate descheduler project does that.

    What interviewers listen for
    • Filtering removes infeasible nodes
    • Scoring ranks feasible nodes; the best one wins
    • Binding sets the node; the kubelet starts the pod
    • Decisions use requests, taints, affinity and topology
    • Preemption for priority; no rebalancing of running pods

    Likely follow-up: How does pod priority and preemption work?

  40. 40.What is a PodDisruptionBudget, and what does it protect against and not protect against?mid

    A PDB limits how many pods of an application can be taken down at once by voluntary disruptions: node drains, cluster upgrades, autoscaler scale-downs, anything that goes through the Eviction API. You set minAvailable or maxUnavailable, as a number or a percentage, and a selector.

    If an eviction would break the budget, the API server refuses it with 429 Too Many Requests, and kubectl drain keeps retrying until replacement pods are healthy somewhere else. That's how a three-replica service survives a rolling node upgrade.

    What it doesn't do:

    • It can't prevent involuntary disruptions like node crashes or OOM kills, though those count against the budget.
    • Deleting pods or a Deployment directly bypasses it.
    • Deployment and StatefulSet rolling updates aren't limited by PDBs; maxUnavailable in the workload controls those.

    Pitfalls: minAvailable equal to the replica count, or maxUnavailable: 0, blocks drains forever, and so does minAvailable: 1 on a single-replica app. Crash-looping pods can also block drains; unhealthyPodEvictionPolicy: AlwaysAllow lets them be evicted.

    apiVersion: policy/v1
    kind: PodDisruptionBudget
    metadata: { name: web-pdb }
    spec:
      minAvailable: 2        # or maxUnavailable: 1
      selector:
        matchLabels: { app: web }
    What interviewers listen for
    • Limits voluntary disruptions through the Eviction API
    • minAvailable or maxUnavailable, plus a selector
    • Blocked evictions return 429; drain retries
    • No protection from crashes or direct deletes
    • Too-strict budgets block node drains

    Likely follow-up: How would you drain a node that runs a single-replica app with a PDB?

  41. 41.Compare the Horizontal Pod Autoscaler, the Vertical Pod Autoscaler and the cluster autoscaler. How do they work together?mid

    They scale different things:

    • HPA changes the number of pods, based on CPU, memory or custom metrics. It's built in.
    • VPA changes each pod's resource requests (and limits proportionally) based on observed usage. It's an add-on with its own CRD. Modes include Off (recommendations only), Initial (set at pod creation), Recreate (evict pods to apply) and InPlaceOrRecreate, which uses in-place pod resize where it can.
    • Cluster autoscaler changes the number of nodes. It adds nodes to a node group when pods are Pending for lack of capacity, and removes underused nodes whose pods fit elsewhere, respecting PDBs. Karpenter is an alternative that provisions right-sized nodes directly.

    They chain naturally: load rises, the HPA adds pods, some can't be scheduled, the cluster autoscaler adds nodes. The VPA keeps requests honest so both decide well, since scheduling works on requests, not usage.

    The main rule: don't let the HPA and VPA act on the same CPU or memory metric for one workload, or they fight. Use VPA in Off mode, or scale the HPA on custom metrics.

    What interviewers listen for
    • HPA scales replicas; VPA scales requests; CA scales nodes
    • VPA is an add-on with Off, Initial and Recreate modes
    • Cluster autoscaler reacts to Pending pods
    • Pending pods from HPA trigger node scale-up
    • Don't combine HPA and VPA on the same metric

    Likely follow-up: How does Karpenter differ from the cluster autoscaler?

  42. 42.What are Custom Resource Definitions and Operators, and how does an operator work?hard

    A CustomResourceDefinition extends the Kubernetes API with a new resource type, like PostgresCluster or Certificate. It declares an OpenAPI v3 schema for validation, versions and optionally a status subresource. After that, the API server serves the new type like a built-in one: kubectl get, RBAC and watches all work. But a CRD on its own just stores data.

    An operator is CRDs plus a custom controller that encodes the operational knowledge for one piece of software. Its reconcile loop:

    • Watches its custom resources and the objects it creates.
    • Compares the desired spec with what actually exists.
    • Acts: creates or updates StatefulSets, Services and Secrets, runs backups, handles failover and version upgrades.
    • Reports progress in status.

    Good operators are level-triggered and idempotent. They set owner references so garbage collection removes child objects, and use finalizers to clean up external resources before deletion. Examples include cert-manager, the Prometheus Operator, CloudNativePG and Strimzi. They're usually built with Kubebuilder or the Operator SDK.

    The trade-offs: another privileged component to run and trust, and CRD version upgrades to manage.

    What interviewers listen for
    • CRD adds a new API type with a schema
    • A CRD alone only stores data
    • Operator = CRD + controller with domain knowledge
    • Reconcile loop: watch, compare, act, update status
    • Owner references and finalizers handle cleanup

    Likely follow-up: Why should reconcile logic be level-triggered instead of edge-triggered?

  43. 43.Which kubectl commands do you use day to day to inspect, change and debug workloads?easy

    Grouped by what I'm doing:

    • Inspect: kubectl get pods -n shop -o wide, filtered with -l app=web, across namespaces with -A, or live with -w. kubectl describe pod <name> is the first stop because of its events. kubectl get deploy web -o yaml shows the full object.
    • Logs and access: kubectl logs <pod> -c <container> -f, --previous after a crash, kubectl exec -it <pod> -- sh, kubectl port-forward svc/web 8080:80, and kubectl debug for images without a shell.
    • Change: kubectl apply -f, with kubectl diff -f first; kubectl rollout status, history, undo and restart; kubectl scale deploy/web --replicas=5.
    • Cluster health: kubectl top pods and kubectl top nodes (needs metrics-server), kubectl events, kubectl cordon and kubectl drain, kubectl auth can-i.
    • Discovery: kubectl explain deployment.spec.strategy, kubectl api-resources.
    • Scaffolding: kubectl create deployment web --image=nginx --dry-run=client -o yaml to generate a starting manifest.

    And always check kubectl config current-context before changing anything, so you don't apply to production by accident.

    What interviewers listen for
    • get, describe (events) and -o yaml to inspect
    • logs --previous, exec, port-forward, debug
    • apply, diff and rollout subcommands to change
    • top, events, drain, auth can-i for operations
    • Check the current context before changing anything
  44. 44.How does Kubernetes compare with Docker Compose and Amazon ECS, and when would you pick each?easy
    • Docker Compose describes a multi-container app in one YAML file and runs it on a single Docker host. It's ideal for local development and simple single-server setups, but there's no scheduling across machines, no rescheduling when a host dies, and no autoscaling.
    • Amazon ECS is AWS's managed orchestrator. You write task definitions and services instead of pods and Deployments, and run them on EC2 or serverless on Fargate. There's no control plane to operate, and it integrates tightly with AWS: IAM roles per task, load balancers, CloudWatch. The trade-offs are AWS lock-in and a smaller, less extensible ecosystem.
    • Kubernetes is the portable open standard across clouds and on-premises, with an extensible API and a huge ecosystem: Helm, operators, service meshes, GitOps. The cost is complexity: upgrades, networking, security and platform skills, even on EKS, GKE or AKS.

    My rule of thumb: Compose for local development; ECS with Fargate for smaller teams all-in on AWS who want minimal operations; Kubernetes when you need portability, hybrid or multi-cloud, many teams and services, or its ecosystem.

    What interviewers listen for
    • Compose: single host, great for local development
    • ECS: managed, AWS-native, Fargate for serverless
    • Kubernetes: portable, extensible, rich ecosystem
    • Kubernetes costs more operational complexity
    • Choose based on scale, team skills and cloud strategy

    Likely follow-up: What would a migration from Compose to Kubernetes involve?

  45. 45.What are the most important security best practices for workloads running on Kubernetes?hard

    I think in layers:

    • Pod hardening: run as a non-root numeric UID, block privilege escalation, drop all capabilities, use a read-only root filesystem and the RuntimeDefault seccomp profile. Avoid privileged, hostNetwork, hostPID and hostPath. Enforce it with Pod Security Admission at restricted.
    • Identity and access: least-privilege RBAC, one ServiceAccount per app, token automount off where unused, and no cluster-admin for CI pipelines.
    • Network: default-deny NetworkPolicies with explicit allows, and mTLS for sensitive traffic.
    • Supply chain: minimal or distroless images, CVE scanning, pinning by digest, signing images with cosign and verifying signatures at admission, and only trusted registries.
    • Secrets: encryption at rest, ideally with KMS, or an external secret store, and never secrets baked into images.
    • Cluster: stay on a supported, patched version, keep the API endpoint private, enable audit logging, lock down etcd and kubelet access, and set resource limits so one workload can't starve a node.
    securityContext:                  # per container
      runAsNonRoot: true
      runAsUser: 10001
      allowPrivilegeEscalation: false
      readOnlyRootFilesystem: true
      capabilities: { drop: [ALL] }
      seccompProfile: { type: RuntimeDefault }
    What interviewers listen for
    • Non-root, no privilege escalation, drop capabilities
    • Enforce with Pod Security Admission restricted
    • Least-privilege RBAC and dedicated ServiceAccounts
    • Default-deny NetworkPolicies
    • Scan, sign and verify images; protect Secrets

    Likely follow-up: How would you give a container a writable temp directory with a read-only root filesystem?

  46. 46.What are the Pod Security Standards, and how does Pod Security Admission enforce them?hard

    The Pod Security Standards define three profiles:

    • Privileged: unrestricted, for trusted system workloads such as CNI or storage agents.
    • Baseline: blocks known privilege escalations, such as privileged containers, host namespaces, hostPath volumes, host ports and dangerous capabilities.
    • Restricted: current hardening practice. On top of Baseline, pods must run as non-root, set allowPrivilegeEscalation: false, drop ALL capabilities (only NET_BIND_SERVICE may be added back), use a RuntimeDefault or Localhost seccomp profile, and use only a limited set of volume types.

    Pod Security Admission is the built-in admission controller that replaced PodSecurityPolicy, which was removed in v1.25. You apply it per namespace with labels such as pod-security.kubernetes.io/enforce: restricted. There are three modes: enforce rejects violating pods, audit records them in the audit log, and warn returns a warning to the user.

    A safe rollout is to enforce baseline with warn and audit set to restricted, fix the violations, then enforce restricted. Enforcement applies to pods, so a Deployment can be accepted while its pods are rejected; look at the ReplicaSet's events. For custom rules, use ValidatingAdmissionPolicy, Kyverno or Gatekeeper.

    What interviewers listen for
    • Three profiles: Privileged, Baseline, Restricted
    • Restricted: non-root, no escalation, drop ALL, seccomp
    • Applied per namespace with pod-security.kubernetes.io/* labels
    • Modes: enforce, audit, warn
    • Replaced PodSecurityPolicy, removed in v1.25

    Likely follow-up: Why might a Deployment be created successfully but have no pods?

  47. 47.Pods restart every few minutes with Liveness probe failed, or never become Ready. How do you debug probe failures?mid

    Start with kubectl describe pod: the events say exactly how the probe failed, and the message points to the cause.

    • "connection refused": nothing is listening on that port. The probe targets the wrong port, the app is still starting, or it's bound to 127.0.0.1 while the kubelet probes the pod IP.
    • Timeout ("context deadline exceeded"): timeoutSeconds defaults to just 1 second. A slow health endpoint, a GC pause or CPU throttling makes it fail. Make the endpoint cheap, or raise the timeout.
    • Bad status code: anything outside 200–399 fails, for example a 401 because the health route requires authentication.
    • Killed while starting: liveness runs before the app is up, causing a restart loop. Add a startupProbe whose failureThreshold × periodSeconds covers the worst-case startup.
    • Never Ready: check what the readiness endpoint depends on; a missing downstream service keeps it false.

    Then reproduce the probe yourself: kubectl exec into the pod and curl the endpoint, or port-forward to it. kubectl get endpointslices -l kubernetes.io/service-name=web shows which endpoints are ready.

    What interviewers listen for
    • Probe failure details are in the pod events
    • Connection refused: wrong port, bind address or slow start
    • Default timeoutSeconds is only 1 second
    • Use a startup probe for slow starters
    • Reproduce the probe with exec or port-forward

    Likely follow-up: Should a readiness probe check the database?

  48. 48.What is Gateway API, and how does it improve on Ingress?hard

    Gateway API is the successor to Ingress, developed by SIG Network. It's an add-on: you install its CRDs, and implementations such as Envoy Gateway, Istio, Cilium or cloud load balancers provide the controller.

    What it improves:

    • Role-oriented resources instead of one overloaded object. A GatewayClass (the infrastructure provider) says which controller to use. A Gateway (the cluster operator) defines listeners, ports, TLS and which namespaces may attach routes. Routes such as HTTPRoute and GRPCRoute (the app teams) attach to a Gateway through parentRefs.
    • Expressive, portable features in the spec: header and query matching, weighted traffic splitting, redirects, rewrites, header changes and mirroring. With Ingress those needed controller-specific annotations.
    • Safe cross-namespace sharing: a Gateway's allowedRoutes controls who can attach, and ReferenceGrant is required to reference backends in another namespace.
    • More protocols: HTTP and gRPC routes are stable; others exist at different maturity levels.

    GatewayClass, Gateway, HTTPRoute and GRPCRoute are stable. With the Ingress API frozen, it's the recommended path, and the ingress2gateway tool helps migrate.

    apiVersion: gateway.networking.k8s.io/v1
    kind: HTTPRoute
    metadata: { name: shop, namespace: shop }
    spec:
      parentRefs: [{ name: public-gw, namespace: infra }]
      hostnames: [shop.example.com]
      rules:
        - matches: [{ path: { type: PathPrefix, value: /api } }]
          backendRefs:
            - { name: api-v1, port: 80, weight: 90 }
            - { name: api-v2, port: 80, weight: 10 }   # 10% canary
    What interviewers listen for
    • Add-on CRDs; the successor to Ingress
    • GatewayClass, Gateway and Routes map to roles
    • Built-in weighted splits, header matching, rewrites
    • Cross-namespace attachment needs explicit permission
    • Recommended over the frozen Ingress API

    Likely follow-up: Who would own the Gateway versus the HTTPRoute in your organization?

  49. 49.What is a service mesh, what does it give you, and when is it worth the cost?hard

    A service mesh is an infrastructure layer that handles service-to-service traffic without changing application code. Its data plane is a set of proxies that intercept traffic: traditionally a sidecar in every pod (Envoy in Istio, Linkerd's own proxy), or per-node proxies as in Istio's ambient mode. Its control plane configures the proxies and issues workload certificates.

    What you get:

    • Security: automatic mTLS between services, with workload identities and authorization policies such as "only checkout may call payments".
    • Traffic control: retries, timeouts, circuit breaking, per-request load balancing (which fixes gRPC imbalance), weighted canaries, mirroring and fault injection.
    • Observability: consistent latency, error and throughput metrics for every hop, plus tracing spans, though apps must still propagate trace headers.

    The costs are real: another complex system to run and upgrade, extra latency and CPU and memory per hop, and one more layer to debug.

    It's worth it with many services, a zero-trust or mTLS requirement, or advanced traffic management needs. For a handful of services, Gateway API, NetworkPolicies and good client libraries are usually enough.

    What interviewers listen for
    • Proxies intercept traffic; control plane configures them
    • Automatic mTLS and service-level authorization
    • Retries, timeouts, traffic splitting, per-request balancing
    • Uniform metrics and tracing across services
    • Adds latency, resource cost and operational complexity

    Likely follow-up: How does a sidecar-less mesh differ from the sidecar model?

  50. 50.What happens when a pod is terminated, and how do you make rolling deployments truly zero-downtime?hard

    When a pod is deleted, by a rollout, a scale-down or a drain:

    • It's marked terminating and its grace period starts (30 seconds by default).
    • In parallel, the control plane marks its endpoints as not ready, so kube-proxy, ingress controllers and load balancers stop sending new traffic, while the kubelet runs any preStop hook and then sends SIGTERM to each container's PID 1.
    • Whatever is still running when the grace period ends gets SIGKILL.

    The catch is that race. SIGTERM can arrive before every node and proxy has stopped routing to the pod, so an app that exits instantly drops requests.

    Zero-downtime checklist:

    • Handle SIGTERM: stop accepting connections, finish in-flight requests, then exit. Make sure your process is PID 1 or runs under an init that forwards signals; a shell-form entrypoint often swallows SIGTERM.
    • Add a short preStop sleep so routing updates propagate first.
    • Set terminationGracePeriodSeconds longer than the preStop plus drain time, since preStop counts against it.
    • Use readiness probes, maxUnavailable: 0, and PDBs for node drains.
    spec:
      terminationGracePeriodSeconds: 45
      containers:
        - name: web
          image: example/web:3.2
          lifecycle:
            preStop:
              sleep: { seconds: 10 }   # let endpoint removal propagate first
          readinessProbe:
            httpGet: { path: /ready, port: 8080 }
    What interviewers listen for
    • Endpoint removal and SIGTERM happen in parallel
    • SIGKILL after the grace period (default 30s)
    • App must handle SIGTERM and drain in-flight work
    • preStop sleep covers routing propagation delay
    • Grace period must cover preStop plus drain time

    Likely follow-up: Why does CMD npm start sometimes break graceful shutdown?

  51. 51.Walk me through everything that happens after you run kubectl apply -f deployment.yaml, until the containers are serving traffic.hard
    • kubectl reads the file, uses the current kubeconfig context for the server and credentials, and sends the object to the API server, as a patch computed from the last-applied state.
    • The API server authenticates the request, authorizes it with RBAC, runs mutating admission (defaults, webhooks such as sidecar injection), validates the schema, runs validating admission (Pod Security, ValidatingAdmissionPolicy, webhooks), and persists the Deployment to etcd.
    • The Deployment controller sees it through a watch and creates a ReplicaSet. The ReplicaSet controller creates Pod objects with no node assigned.
    • The scheduler filters and scores nodes and binds each pod to one.
    • The kubelet on that node sees the pod, has the container runtime pull images and create the pod sandbox, the CNI plugin assigns an IP, volumes are mounted, init containers run, then app containers start and probes begin.
    • When readiness passes, the EndpointSlice controller adds the pod to its Services, and kube-proxy and ingress controllers route traffic to it.

    No controller calls another directly; they coordinate through watches on the API server.

    What interviewers listen for
    • Authentication, authorization, admission, then etcd
    • Deployment controller creates a ReplicaSet, which creates pods
    • Scheduler binds pods to nodes
    • Kubelet, runtime, CNI and volumes start the containers
    • Ready pods join EndpointSlices and receive traffic

    Likely follow-up: Where would a mutating webhook failure show up? · What is the difference between client-side and server-side apply?

  52. 52.What does externalTrafficPolicy: Local do on a LoadBalancer or NodePort Service, and what are the trade-offs versus Cluster?hard

    It controls where a node may send traffic that arrives from outside the cluster.

    • Cluster (the default): any node can forward to any ready pod, including pods on other nodes. Load spreads evenly and every node can accept traffic, but the extra hop is source-NATed so replies come back the same way. The pod sees a node's IP, not the client's.
    • Local: a node forwards only to pods on that same node. There's no second hop and no SNAT, so the client source IP is preserved (useful for allowlists, rate limiting and logs) and latency drops. Nodes with no local pod drop the traffic, so for LoadBalancer Services Kubernetes allocates a healthCheckNodePort that the cloud load balancer probes to find which nodes have endpoints.

    The trade-offs of Local: load can become imbalanced, because the load balancer spreads per node, not per pod, so a node with one pod gets as much traffic as a node with three. Pod moves also depend on the load balancer's health checks catching up. Spreading pods evenly helps.

    What interviewers listen for
    • Cluster: any node to any pod, SNAT hides client IP
    • Local: only node-local pods, client IP preserved
    • Local drops traffic on nodes without endpoints
    • healthCheckNodePort tells the LB which nodes have pods
    • Local risks imbalance: LB spreads per node, not per pod
  53. 53.How would you make sure only trusted, approved container images can run in your cluster?hard

    I'd combine controls at build time and at admission time:

    • Trusted registries only: reject pods whose images don't come from your approved registries. Enforce it at admission with ValidatingAdmissionPolicy (CEL rules, built in), Kyverno or OPA Gatekeeper.
    • Immutable references: tags can be moved, digests can't. Require image@sha256:…, or at least block :latest and untagged images.
    • Signing and provenance: sign images in CI with Sigstore cosign, attach SBOM and provenance attestations, and verify signatures at admission, with Kyverno's image verification, the Sigstore policy-controller, or Ratify with Gatekeeper.
    • Vulnerability scanning: scan in CI with a tool like Trivy or Grype, fail builds above a severity threshold, and keep rescanning running images, since new CVEs appear after deployment.
    • Minimal base images, distroless or scratch, to shrink the attack surface.
    • Pull behavior in multi-tenant clusters: the AlwaysPullImages admission plugin forces a registry check with the pod's own credentials, so one tenant can't reuse another's cached private image.

    Admission policies are the enforcement point; everything else feeds them trustworthy data.

    What interviewers listen for
    • Admission policy restricts allowed registries
    • Pin images by digest; block :latest
    • Sign with cosign and verify at admission
    • Scan in CI and continuously for CVEs
    • Minimal images reduce the attack surface

    Likely follow-up: What happens to already-running pods when a new policy is added?

  54. 54.A worker node suddenly dies. What happens to its pods, step by step, and how long does recovery take?hard
    • The kubelet stops renewing its Lease in kube-node-lease. After a grace period (tens of seconds by default) the node controller sets the node's Ready condition to Unknown, marks its pods not ready so they drop out of Service endpoints, and taints the node node.kubernetes.io/unreachable.
    • Pods get a default toleration for that taint of 300 seconds, so after about five minutes they're evicted.
    • Deployment and ReplicaSet pods then get replacements on healthy nodes, because terminating pods no longer count toward the replica total.
    • StatefulSet pods are not replaced automatically. Kubernetes can't tell a dead node from a network partition, and two db-0 pods could corrupt data. You confirm the node is really gone, then delete the Node, force-delete the pod, or add the node.kubernetes.io/out-of-service taint.
    • ReadWriteOnce volumes stay attached to the dead node until they're detached. Kubernetes force-detaches after about six minutes if the node is still unhealthy.
    • DaemonSet pods tolerate the taint and stay bound.

    Shorter tolerationSeconds on critical stateless apps speeds up failover.

    What interviewers listen for
    • Missed Lease renewals mark the node Unknown
    • Unreachable taint plus 300s default toleration
    • ReplicaSets replace pods after about five minutes
    • StatefulSets don't auto-replace; manual confirmation needed
    • RWO volumes must detach before pods move

    Likely follow-up: Why is force-deleting a StatefulSet pod dangerous?

  55. 55.How do you debug a running pod whose image is distroless and has no shell? What does kubectl debug do?mid

    kubectl exec needs a shell and tools inside the image, which distroless images deliberately leave out. kubectl debug brings your own tools, in three ways:

    • Ephemeral container: adds a temporary container with a debug image to the running pod, without restarting it. With --target, it shares the target container's process namespace, so you can see its processes, inspect /proc/<pid>/root and use tools like netstat or nslookup in the pod's network namespace. Ephemeral containers are stable since v1.25. They can't have ports, probes or resources, they're never restarted, and they can't be removed from the pod afterwards.
    • Pod copy (--copy-to): creates a copy of the pod, optionally with a debug container, a changed command or different images. That's useful for a container that crashes on startup, because you can start it with a different command.
    • Node debugging (node/<name>): runs a pod in the node's host namespaces with the host filesystem at /host, for checking the kubelet or container runtime without SSH.

    --profile selects preset security settings for the debug container, such as general, baseline, restricted or sysadmin.

    # ephemeral container attached to a running pod, sharing the app's process namespace
    kubectl debug -it api-6c9f --image=busybox:1.36 --target=api
    
    # copy of the pod plus a debug container (the original keeps serving)
    kubectl debug api-6c9f -it --image=busybox:1.36 --copy-to=api-debug --share-processes
    
    # pod in the node's host namespaces, host filesystem at /host
    kubectl debug node/worker-1 -it --image=busybox:1.36
    What interviewers listen for
    • Ephemeral containers add tools to a running pod
    • --target shares the app container’s process namespace
    • --copy-to debugs a modified copy of the pod
    • node/<name> debugs the node without SSH
    • Ephemeral containers: no ports, probes or restarts
  56. 56.How do you make the Kubernetes control plane highly available, and why does etcd need an odd number of members?hard

    Run several control-plane nodes, typically three, spread across failure zones. Each component handles redundancy differently:

    • kube-apiserver is stateless, so all replicas are active behind a load balancer. That load balancer's address is the cluster endpoint for kubectl and the kubelets.
    • kube-scheduler and kube-controller-manager run on every control-plane node, but only one of each is active at a time. They use leader election with Lease objects in kube-system, and a standby takes over when the leader stops renewing.
    • etcd uses the Raft consensus protocol: a write needs a majority of members. Three members tolerate one failure, five tolerate two. Four still tolerate only one, since the quorum is three, so an even count adds overhead without adding fault tolerance.

    You can run etcd stacked on the control-plane nodes (fewer machines, but failures are coupled) or as a separate external cluster (more machines, better isolation).

    If etcd loses quorum, the API stops accepting changes. Existing containers keep running on their nodes, but nothing is rescheduled or reconciled until quorum is restored. Managed services run all of this for you.

    What interviewers listen for
    • Multiple control-plane nodes across zones
    • API servers active-active behind a load balancer
    • Scheduler and controller manager use leader election
    • etcd quorum is a majority; odd sizes like 3 or 5
    • Lost quorum: no changes, but workloads keep running

    Likely follow-up: What are the trade-offs between stacked and external etcd?

  57. 57.How do emptyDir, hostPath and PersistentVolumeClaim volumes differ, and when would you use each?easy

    A container's own filesystem is lost when the container restarts. Volumes are declared at the pod level and mounted into containers, and they differ mainly in how long the data lives:

    • emptyDir: created empty when the pod lands on a node and shared by all its containers. It survives container crashes but is deleted with the pod. Use it for scratch space, caches and handing files between containers. medium: Memory makes it a tmpfs that counts against the container's memory limit, and sizeLimit caps it.
    • configMap, secret, downwardAPI and projected: inject configuration, credentials or pod metadata as files.
    • persistentVolumeClaim: durable storage backed by a PersistentVolume, such as a cloud disk or NFS, that outlives the pod and can follow it to another node.
    • hostPath: mounts a directory from the node itself. It ties the pod to that node's data and is a serious security risk, since it can expose host files. Pod Security Baseline forbids it; keep it for trusted system agents, like a log collector reading /var/log.
    What interviewers listen for
    • emptyDir lives and dies with the pod
    • configMap and secret volumes inject files
    • PVCs provide durable storage that outlives pods
    • hostPath exposes the node filesystem; avoid it
    • Memory-backed emptyDir counts against memory limits
  58. 58.How do you spread replicas evenly across zones and nodes? Explain topologySpreadConstraints.hard

    A topology spread constraint tells the scheduler how unevenly matching pods may be distributed across domains, defined by a node label in topologyKey: zones, nodes or racks.

    • maxSkew: the maximum allowed difference between the number of matching pods in a domain and the lowest count in any eligible domain. With maxSkew: 1 and three zones, six replicas land 2/2/2, and seven land 3/2/2.
    • whenUnsatisfiable: DoNotSchedule makes it a hard rule, so the pod stays Pending; ScheduleAnyway makes it a scoring preference.
    • labelSelector: which pods count, usually the app's own labels. matchLabelKeys can restrict counting to pods of the same revision during rollouts.

    Compared with pod anti-affinity, which is essentially all-or-nothing ("not on a node that already has one"), spread constraints scale with the replica count. A common combination is a hard rule across zones and a soft one across nodes, as shown.

    Caveat: they're only evaluated at scheduling time. Scale-downs or node failures can leave things skewed, and the descheduler project can rebalance. Without explicit constraints, the scheduler still applies soft built-in defaults for nodes and zones.

    topologySpreadConstraints:
      - maxSkew: 1
        topologyKey: topology.kubernetes.io/zone
        whenUnsatisfiable: DoNotSchedule     # hard rule
        labelSelector: { matchLabels: { app: web } }
      - maxSkew: 1
        topologyKey: kubernetes.io/hostname
        whenUnsatisfiable: ScheduleAnyway    # soft preference
        labelSelector: { matchLabels: { app: web } }
    What interviewers listen for
    • topologyKey defines domains such as zones or nodes
    • maxSkew bounds the imbalance between domains
    • DoNotSchedule is hard; ScheduleAnyway is soft
    • Scales better than anti-affinity for many replicas
    • Checked only at scheduling; no automatic rebalancing

    Likely follow-up: What happens to spreading when a Deployment scales down?

Prefer multiple choice? All 20 Kubernetes MCQs with answers →

esc