Observability and Troubleshooting
Know where each signal comes from (logs, events, metrics, traces) and work through failures systematically: get pods, describe, logs (including --previous), then a debug container.
Key points
- 1
Containers should log to stdout and stderr; the runtime writes those to node files that
kubectl logsreads. Only the latest file is served (10Mi by default), and logs disappear with the pod. - 2
Events are API objects kept for about an hour by default; read them with
kubectl describeor sortkubectl get eventsexplicitly. - 3
metrics-server feeds
kubectl topand autoscalers with recent CPU and memory only; Prometheus with kube-state-metrics, node-exporter and cAdvisor gives history and alerting. - 4
Read container state carefully: ImagePullBackOff (pull failing), CreateContainerConfigError (missing ConfigMap or Secret key), CrashLoopBackOff (back-off from 10 s to 5 min), OOMKilled with exit 137.
- 5
Use
kubectl debugephemeral containers with--targetfor distroless images, andkubectl debug node/<name>(node filesystem at /host) when SSH is unavailable. - 6
Node problems show up as conditions (Ready, MemoryPressure, DiskPressure, PIDPressure); start with the kubelet and runtime logs on the node.
- 7
Right-size requests from observed peaks; scheduling and autoscaling follow requests, not usage.
Common traps
kubectl logs -l app=xshows only 10 lines per container unless you set--tail.--previousonly covers restarts inside the same pod; a replacement pod has no previous logs.Average CPU can look low while CFS throttling still causes tail latency under a CPU limit.
Test yourself on Observability and Troubleshooting
Ten questions, with the answer and explanation after each one.