Performance and CPython Internals
Know how CPython runs and frees your code, so you can predict what is slow or leaky and measure it with the right tool instead of guessing.
Key points
- 1
Imported modules are compiled to bytecode and cached in
__pycache__; the evaluation loop runs that bytecode. Since 3.11 the specializing adaptive interpreter (PEP 659) swaps in type-specific instructions such asBINARY_OP_ADD_INTafter a warm-up. The 3.13 JIT exists only in builds made with--enable-experimental-jit. - 2
Memory is freed by reference counting the moment a count hits zero; a separate generational collector frees reference cycles, and since 3.4 it also runs
__del__on cycles. None, True, False and small ints are immortal since 3.12, so their refcounts never change. - 3
sys.getsizeofis shallow and in 3.12+ doesn't include a plain instance's inline attribute values. Measure real footprints withtracemalloc: on 3.13 a two-attribute instance is about 96 bytes, about 56 with__slots__, and about 160 once its__dict__is materialised. - 4
Pick the right data structure first: set and dict lookups are O(1) on average,
list.pop(0)andinsert(0, …)are O(n) (usedeque),list.removein a loop is quadratic, and''.joinbeats repeated+=. - 5
Profile before optimising: cProfile for per-function time (with per-call overhead), py-spy to sample live processes, line_profiler for lines inside a known hot function, tracemalloc snapshots to diff memory growth, timeit for micro-benchmarks.
- 6
Move hot numeric loops out of Python: NumPy runs one C loop over an unboxed, contiguous buffer instead of dispatching bytecode and allocating an object per element. Reach for PyPy, Cython or Rust (PyO3) when the work can't be vectorised.
- 7
Caches are a common leak:
lru_cacheon a method keeps everyselfalive, saved exceptions keep their frames' locals alive, and bound methods in registries keep instances alive. Bound caches or hold values weakly.
Common traps
timeitdisables garbage collection while timing, so allocation-heavy code looks faster than in production; addgc.enable()to the setup when GC cost matters.iscomparisons on ints and strings depend on caching and constant merging (int('256') is int('256')is True, 257 is False; one script versus REPL lines), so always compare values with==.CPython's in-place
s += tspeed-up only applies when nothing else referencess; keep another reference and every step copies, so it turns quadratic.
Read the source
Test yourself on Performance and CPython Internals
Ten questions, with the answer and explanation after each one.