Containers, Namespaces, cgroups and OCI
A container is an ordinary process on the host kernel, isolated with namespaces, limited with cgroups and given a layered root filesystem. Know where that isolation ends.
Key points
- 1
Containers share the host kernel; VMs boot their own. The image brings userspace only, so kernel ABI and CPU architecture must match (or be emulated).
- 2
Namespaces isolate what a process sees: PID, network, mount, UTS, IPC, user and cgroup. Docker creates no time namespace, so the clock and kernel are shared.
- 3
cgroups v2 limit what it uses: memory.max (hard limit, then OOM kill, exit 137), cpu.max (quota/period from --cpus), cpu.weight (relative shares) and pids.max.
- 4
overlay2 stacks read-only image layers under one writable container layer. Reads search upper then lower layers; writes copy the whole file up; deletions become 0/0 whiteouts.
- 5
The stack is CLI → dockerd → containerd → shim → runc. runc (the OCI runtime) creates the process and exits; the shim keeps it alive across daemon restarts.
- 6
The OCI defines the runtime, image and distribution specs, which is why images move between Docker, Podman, containerd and CRI-O.
- 7
Docker Desktop runs Linux containers in a VM, so bind mounts cross a file-sharing boundary while named volumes stay fast.
Common traps
--cpusis a time quota, not pinning:nprocandfreeinside the container still report host-wide values.With --memory set and --memory-swap unset, a host with swap lets the container use the same amount again as swap.
Root in a default container is host UID 0 with fewer capabilities, not an unprivileged user, unless user namespaces are enabled.
Read the source
Test yourself on Containers, Namespaces, cgroups and OCI
Ten questions, with the answer and explanation after each one.