Performance, Memory and Scaling
Node runs JavaScript on one thread, so scaling means keeping that thread free and adding processes or threads deliberately. Know where memory lives, how to profile, and which resource caps throughput.
Key points
- 1
CPU-bound JavaScript blocks every request. Move it to worker threads (with a pool), another service or native code; I/O-bound work does not need workers.
- 2
cluster forks separate processes that share a port but no memory, so sessions, caches and counters must live in an external store.
- 3
Read process.memoryUsage() by field: heapUsed growth means JS objects; arrayBuffers or external growth means Buffers; RSS-only growth means native memory or fragmentation.
- 4
Use heap snapshots (three-snapshot technique, --heapsnapshot-signal, --heapsnapshot-near-heap-limit) for leaks and --cpu-prof or Clinic.js for CPU hot spots.
- 5
Throughput is capped by the scarcest resource: CPU per request, pool connections (Little's Law), the 4-thread libuv pool, or ephemeral ports.
- 6
Reuse outbound connections (keep-alive agents, undici Pool) and set worker counts and heap limits explicitly in containers.
- 7
Benchmark from a separate machine, warm up first, and use fixed-rate load when you care about tail latency.
Common traps
os.cpus().length sees the whole machine, not your container's CPU quota.
Raising setMaxListeners hides a listener leak instead of fixing it.
Closed-loop benchmarks hide queuing during stalls (coordinated omission), so p99 looks far better than production.
Test yourself on Performance, Memory and Scaling
Ten questions, with the answer and explanation after each one.