Containers before Kubernetes
Namespaces, cgroups and image layers by hand, so that pod behaviour stops looking like magic and starts looking like Linux.
Almost every confusing thing Kubernetes does to a container is something Linux was already doing. If you learn the orchestrator first, you end up memorising behaviour. If you spend an afternoon below it, the same behaviour becomes obvious.
A container is three kernel features and a chroot
There is no container object in the Linux kernel. What you get instead is a process with a restricted view of the system, assembled from namespaces, cgroups and a mounted filesystem. You can build one by hand:
# a new PID, mount, UTS and network namespace, with a shell inside it
sudo unshare --pid --mount --uts --net --fork --mount-proc bash
# inside: you are process 1, and you can see almost nothing
ps aux
hostname container-by-hand
ip addr # just lo - no route out, because nothing was plugged in yet
That last line is the whole of pod networking in miniature. A fresh network namespace has a loopback interface and nothing else. Something outside has to create a virtual ethernet pair, move one end in, and add routes. On a Kubernetes node, that something is the CNI plugin.
cgroups are why your pod was killed
Namespaces control what a process can see. Control groups control what it can use. On cgroup v2 the interface is a filesystem:
sudo mkdir /sys/fs/cgroup/demo
echo "100M" | sudo tee /sys/fs/cgroup/demo/memory.max
echo $$ | sudo tee /sys/fs/cgroup/demo/cgroup.procs
# now allocate more than that in this shell and watch what happens
cat /sys/fs/cgroup/demo/memory.events # look at oom_kill
memory.max on a cgroup. When the process crosses it, the kernel OOM killer acts — not the kubelet, not the scheduler. That is why the container exits with code 137 and your application logs show nothing: it was not asked to stop.CPU behaves differently, and the difference matters. A CPU limit becomes a quota per period, so exceeding it does not kill anything — it throttles. A pod that is slow but alive is almost always CPU throttling; a pod that dies abruptly is almost always memory.
Images are layers, and layers are a tax
An image is an ordered stack of tarballs plus a JSON manifest. Each instruction in a Dockerfile that changes the filesystem adds a layer, and layers are immutable — so deleting a file in a later layer hides it without reclaiming the space.
# where the size actually went
docker history --no-trunc --format '{{.Size}}\t{{.CreatedBy}}' your-image:tag | head -20
This stops being an aesthetic concern the moment you work on anything with CUDA or a model in it. A 12 GB image is 12 GB pulled onto every node that has to run it, before your process starts. It is the single most common reason a GPU pod appears to "hang" on first schedule — it is not hanging, it is pulling.
What to actually do with this
- Run the
unsharecommand above once. Ten minutes, and network policy later will make sense. - Put a
memory.maxon a shell and get OOM-killed deliberately, so you recognise it in production. - Run
docker historyon the largest image you own, and find the one layer that is most of it. - Multi-stage builds: compile in one stage, copy only the artefact into a slim final stage. Usually the single biggest win available.
Everything above is a property of Linux, not of Kubernetes. That is the point — it is all still true three abstractions up.
Something wrong or out of date? Open an issue — corrections are welcome and get credited.