Study notes · 13% of the exam

Reliability, Resilience and Observability

Design so partial failures stay partial: bounded retries, shrinking timeouts, isolation between dependencies, overload protection and SLOs that tell you when to act.

Key points

  1. 1

    Retry with capped exponential backoff and jitter, at one layer only, inside a retry budget. Retries at every layer multiply (3 attempts at 3 layers = 27×).

  2. 2

    Timeouts should shrink as calls go deeper and a layer's whole retry plan must fit in its caller's deadline. Propagate the remaining deadline downstream.

  3. 3

    Hard serial dependencies multiply availability (0.999³ ≈ 99.7%). Fan-out turns rare tail latency into the common case; hedged requests trim it cheaply.

  4. 4

    Circuit breakers fail fast on sick dependencies, bulkheads isolate resource pools, load shedding protects a server past capacity and rate limits enforce per-client fairness.

  5. 5

    An error budget is 1 − SLO over the window (99.9% over 30 days = 43.2 minutes). Alert on burn rate with a long and a short window, for example 14.4× over 1 hour and 5 minutes.

  6. 6

    Metastable failures outlive their trigger because retries and cold caches keep load above capacity. Recovery means cutting offered load, not adding callers.

Common traps

  • A timeout does not mean the request did nothing, so retried operations must be idempotent.

  • Adding app servers during an overload often worsens it, because they all hit the same bottleneck.

  • A token bucket's capacity caps the burst: a long idle period does not bank unlimited requests.

Test yourself on Reliability, Resilience and Observability

Ten questions, with the answer and explanation after each one.