Reliability, Resilience and Observability
Design so partial failures stay partial: bounded retries, shrinking timeouts, isolation between dependencies, overload protection and SLOs that tell you when to act.
Key points
- 1
Retry with capped exponential backoff and jitter, at one layer only, inside a retry budget. Retries at every layer multiply (3 attempts at 3 layers = 27×).
- 2
Timeouts should shrink as calls go deeper and a layer's whole retry plan must fit in its caller's deadline. Propagate the remaining deadline downstream.
- 3
Hard serial dependencies multiply availability (0.999³ ≈ 99.7%). Fan-out turns rare tail latency into the common case; hedged requests trim it cheaply.
- 4
Circuit breakers fail fast on sick dependencies, bulkheads isolate resource pools, load shedding protects a server past capacity and rate limits enforce per-client fairness.
- 5
An error budget is 1 − SLO over the window (99.9% over 30 days = 43.2 minutes). Alert on burn rate with a long and a short window, for example 14.4× over 1 hour and 5 minutes.
- 6
Metastable failures outlive their trigger because retries and cold caches keep load above capacity. Recovery means cutting offered load, not adding callers.
Common traps
A timeout does not mean the request did nothing, so retried operations must be idempotent.
Adding app servers during an overload often worsens it, because they all hit the same bottleneck.
A token bucket's capacity caps the burst: a long idle period does not bank unlimited requests.
Read the source
Test yourself on Reliability, Resilience and Observability
Ten questions, with the answer and explanation after each one.