Pod networking from first principles
Follow a packet from one pod to another across nodes, then do it again with the CNI removed so you can see what it was doing for you.
Kubernetes networking has exactly three rules. Everything else is an implementation detail of satisfying them:
- Every pod gets its own IP address.
- Pods can reach each other at those addresses, without NAT.
- Agents on a node can reach all pods on that node.
That is the whole contract. It feels complicated because nothing in Linux provides it, so a CNI plugin has to build it on every node, and the plugins do it differently.
Inside a node: veth pairs
A pod is a network namespace. To give it connectivity the plugin creates a virtual ethernet pair — a cable with two ends — puts one end inside the namespace as eth0, and leaves the other on the host.
# find the pod's interface index from inside the pod
kubectl exec my-pod -- cat /sys/class/net/eth0/iflink
# then match that index on the node
ip link | grep "^<index>:"
# the host side of every pod on this node
ip -d link show type veth
The pod routing table is deliberately almost empty: a default route pointing at the host side of the pair. Every real decision is made on the node, not in the pod.
Between nodes: overlay or routed
- Overlay (VXLAN, Geneve) wraps the pod packet inside a node-to-node packet. Works on any network, costs you MTU and a little CPU.
- Routed (BGP, or a cloud route table) tells the underlying network where each node pod CIDR lives, so pod packets travel unmodified. Faster and easier to debug, but the network has to cooperate.
kubectl exec pod -- ip link show eth0 and expect 1450, not 1500.Services are not processes
A ClusterIP does not exist anywhere. No process listens on it. It is a rule on every node that rewrites the destination of outgoing packets to one of the backing pod IPs:
# iptables mode - the classic
sudo iptables -t nat -L KUBE-SERVICES -n | grep my-svc
# IPVS mode - a real load balancer table
sudo ipvsadm -Ln
# what the Service actually resolves to
kubectl get endpointslices -l kubernetes.io/service-name=my-svc -o yaml
The modes differ only in how that rewriting is implemented. iptables builds a chain per service and scales linearly, which starts hurting in the thousands. IPVS uses a kernel hash table and stays flat. eBPF replaces the chains with programs attached to the socket and the interface, and can skip much of the network stack for pods on the same node.
DNS is the usual suspect
kubectl exec -it my-pod -- cat /etc/resolv.conf
# nameserver 10.96.0.10
# search team-a.svc.cluster.local svc.cluster.local cluster.local
# options ndots:5
ndots:5 means any name with fewer than five dots is tried against every search domain first. A lookup of api.example.com — two dots — generates four failing queries before the correct one. Harmless until DNS is under load, at which point it is four times the load you think you have. A trailing dot (api.example.com.) skips the search list entirely.
What to actually do with this
- Find the veth pair for one of your pods, on the node. Two commands, and the model stops being abstract.
- Check your pod MTU. If it is 1500 on an overlay, you have a latent bug.
- Dump the iptables NAT table or run
ipvsadm -Lnonce, so a Service stops being magic.
Something wrong or out of date? Open an issue โ corrections are welcome and get credited.