Scaling, Load Balancing and Statelessness
Scale out by making instances interchangeable, routing work to them with the right balancing policy, and watching the real saturation signal. Know where adding machines stops helping.
Key points
- 1
Vertical scaling (a bigger box) is simple but has a ceiling and leaves a single point of failure. Horizontal scaling (more boxes) needs stateless instances or partitioned state.
- 2
Keep request-handling tiers stateless: put sessions in a shared store or in signed client tokens so any instance can serve any request.
- 3
A layer-4 balancer works per TCP connection. A layer-7 balancer works per request, can route by path or header, and must terminate TLS to do so.
- 4
Round robin assumes equal-cost requests. Least-outstanding-requests or power-of-two-choices handle variable costs, and power of two is robust to stale load information.
- 5
Modulo hashing remaps about N/(N+1) of keys when a node is added. Consistent hashing remaps only about 1/(N+1).
- 6
Scale on the resource that actually saturates. An I/O-bound tier needs concurrency or latency signals, not CPU, and scaling it out can overload the database behind it.
Common traps
Treating sticky sessions as a scaling strategy: state still dies with the instance.
Expecting more nodes to always add throughput. Cross-node coordination (the coherency cost) can make throughput fall.
Terminating instances without draining them first, which drops in-flight requests on every deploy.
Read the source
Test yourself on Scaling, Load Balancing and Statelessness
Ten questions, with the answer and explanation after each one.