Study notes · 10% of the exam

Pods and Workload Controllers

Know how Pods work and which controller to use for each workload, and how restarts, probes, termination and Job/CronJob semantics behave in practice.

Key points

  1. 1

    A Pod is the smallest unit: its containers share one network namespace (localhost, one IP) and any mounted volumes. Pods are ephemeral and not self-healing; controllers replace them.

  2. 2

    Pick the controller by shape: Deployment (stateless replicas, rolling updates via ReplicaSets), StatefulSet (ordinal names, stable DNS via a headless Service, per-Pod storage), DaemonSet (one per eligible node), Job/CronJob (run to completion, on a schedule).

  3. 3

    Probes: readiness gates traffic, liveness restarts the container, and a startup probe suspends both until the app has started (budget = failureThreshold × periodSeconds). Keep liveness checks to the process itself.

  4. 4

    Restarts back off 10s, 20s, 40s… up to 300 s (CrashLoopBackOff) and reset after 10 minutes of healthy running. OOMKilled with exit code 137 means the memory limit was hit.

  5. 5

    Termination: deletion timestamp → (in parallel) endpoint removal and preStop → SIGTERM → SIGKILL after terminationGracePeriodSeconds (default 30). A short preStop sleep avoids dropped requests.

  6. 6

    Jobs: completions vs parallelism, backoffLimit (default 6), activeDeadlineSeconds takes precedence, Indexed mode gives each Pod JOB_COMPLETION_INDEX, and podFailurePolicy (first match wins, needs restartPolicy Never).

  7. 7

    Native sidecars (initContainers with restartPolicy: Always, stable since 1.33) start before the app, keep running, support probes, stop after the main containers, and do not block Job completion.

Common traps

  • A liveness probe that checks the database turns a database blip into restart storms across every replica.

  • Rolling back to revision 2 renumbers it as the newest revision; revision 2 disappears from rollout history.

  • Bare Pods whose labels match a controller's selector get adopted (or deleted as surplus), and overlapping selectors between controllers cause unpredictable fights.

Test yourself on Pods and Workload Controllers

Ten questions, with the answer and explanation after each one.