Study notes · 8% of the exam

Observability and Troubleshooting

Know where each signal comes from (logs, events, metrics, traces) and work through failures systematically: get pods, describe, logs (including --previous), then a debug container.

Key points

  1. 1

    Containers should log to stdout and stderr; the runtime writes those to node files that kubectl logs reads. Only the latest file is served (10Mi by default), and logs disappear with the pod.

  2. 2

    Events are API objects kept for about an hour by default; read them with kubectl describe or sort kubectl get events explicitly.

  3. 3

    metrics-server feeds kubectl top and autoscalers with recent CPU and memory only; Prometheus with kube-state-metrics, node-exporter and cAdvisor gives history and alerting.

  4. 4

    Read container state carefully: ImagePullBackOff (pull failing), CreateContainerConfigError (missing ConfigMap or Secret key), CrashLoopBackOff (back-off from 10 s to 5 min), OOMKilled with exit 137.

  5. 5

    Use kubectl debug ephemeral containers with --target for distroless images, and kubectl debug node/<name> (node filesystem at /host) when SSH is unavailable.

  6. 6

    Node problems show up as conditions (Ready, MemoryPressure, DiskPressure, PIDPressure); start with the kubelet and runtime logs on the node.

  7. 7

    Right-size requests from observed peaks; scheduling and autoscaling follow requests, not usage.

Common traps

  • kubectl logs -l app=x shows only 10 lines per container unless you set --tail.

  • --previous only covers restarts inside the same pod; a replacement pod has no previous logs.

  • Average CPU can look low while CFS throttling still causes tail latency under a CPU limit.

Test yourself on Observability and Troubleshooting

Ten questions, with the answer and explanation after each one.