One article for every checkpoint

The roadmap says what to learn and in what order. These say how the thing actually behaves, including the parts that only show up once you run it.

20articles
200minutes
7stages

Foundations

Enough of the object model that everything later has somewhere to attach.

Containers before Kubernetes

Namespaces, cgroups and image layers by hand, so that pod behaviour stops looking like magic and starts looking like Linux.

KCNA8 min

The API object model

Why everything is a resource, what the control loop actually does, and how to read any CRD you meet later without waiting for someone to document it.

KCNACKAD9 min

Workloads and configuration

Getting your own software to run, survive a restart and be configurable without a rebuild.

Deployments, probes and rollouts

Rolling updates that do not drop traffic, and the three probes people keep confusing.

CKADCKA10 min

Config, secrets and the twelve-factor edge cases

ConfigMap reload behaviour, projected volumes, and why your Secret is base64 and not encrypted.

CKADSecurity8 min

Declarative delivery

Kustomize overlays and a pull-based reconciler, so environments stop drifting silently and a cluster can be rebuilt from Git alone.

GitOpsCKAD9 min

Running a cluster

The part where it is your pager. Control plane internals, upgrades, and recovering from your own mistakes.

Control plane components and etcd

What each component owns, plus a real backup and restore of etcd on a cluster you can afford to destroy.

CKA11 min

RBAC and multi-tenancy

Roles that are actually least-privilege, and the three places tenancy leaks anyway.

CKASecurityKCSA10 min

Troubleshooting under time pressure

A fixed order of operations for a broken cluster, so you stop guessing when it matters most.

CKAObservability9 min

Networking

Where most senior interviews go, and where the CKNE lives. Packets, policy and the mesh question.

Pod networking from first principles

Follow a packet from one pod to another across nodes, then do it again with the CNI removed so you can see what it was doing for you.

CKNENetworkingCKA12 min

Network policy that someone can debug

Default-deny without taking production down, and proving the policy does what you claimed it does.

CKNENetworkingCKSSecurity10 min

Gateway API and the mesh decision

Ingress to Gateway API route by route, and an honest list of what a service mesh costs you.

CKNENetworking11 min

Security

Supply chain, admission and runtime. The CKS track, plus the parts no exam covers.

Admission control and policy as code

Pod Security Admission first, then a policy engine for the rules PSA cannot express and a webhook only when nothing else will do.

CKSSecurityKCSA10 min

Supply chain and image provenance

Signing, verifying, and actually failing closed when verification fails — which is the only part that matters.

CKSSecurity9 min

Runtime detection and response

Seeing a container do something it should not, and having a plan for the next ten minutes rather than inventing one live.

CKSSecurityObservability9 min

AI inference and GPUs

No certification covers this yet, which is exactly why it is worth writing down.

Getting a GPU into a pod

The device plugin model, drivers, and the difference between MIG and time-slicing when cost is the constraint.

InferenceScheduling11 min

Serving a model like a production service

vLLM behind a Gateway, with readiness that reflects model load rather than process start — which is where most first attempts go wrong.

InferenceNetworking12 min

Scaling inference on real traffic

Queue-depth autoscaling, cold starts measured in minutes, and the bill as a design constraint rather than an afterthought.

InferenceSchedulingObservability11 min

Platform and scale

What changes when other teams depend on you and nobody reads your docs.

Capacity, bin-packing and autoscaling

Requests and limits chosen from data, plus node autoscaling that does not thrash.

SchedulingCKA10 min

SLOs people actually use

Three signals per service, one dashboard, and alerts that correspond to someone being paged.

Observability9 min

Extending Kubernetes yourself

A CRD and controller with controller-runtime — which is also the shortest honest route to your first upstream contribution.

SchedulingGitOps12 min