Admission control and policy as code
Pod Security Admission first, then a policy engine for the rules PSA cannot express and a webhook only when nothing else will do.
Admission control is the last point at which you can say no. It runs after authentication and authorisation, after the object has been validated, and before anything is written to etcd — so a rejection means the bad object never existed.
Three tiers, in order of how much they cost you
- Pod Security Admission — built in, three labels, no components to run. Covers the dangerous pod fields.
- A policy engine (Kyverno, Gatekeeper) — declarative rules for everything else, including mutation and cross-object checks.
- A custom webhook — arbitrary code, maximum power, and now you own an availability problem.
Most teams reach straight for tier two and skip tier one, then write forty policies reimplementing what three labels would have done.
Pod Security Admission, which is free
apiVersion: v1
kind: Namespace
metadata:
name: team-a
labels:
pod-security.kubernetes.io/enforce: baseline # reject the worst
pod-security.kubernetes.io/enforce-version: v1.31
pod-security.kubernetes.io/audit: restricted # log what would fail
pod-security.kubernetes.io/warn: restricted # warn the applier
warn and audit to restricted while enforce stays at baseline. You get the full list of what would break, in warnings and the audit log, without breaking it. Tighten enforce once that list is empty.baseline blocks privileged containers, host namespaces, hostPath and dangerous capabilities. restricted additionally requires non-root, dropping all capabilities and a seccomp profile. The gap between them is where almost every real application argument happens.
Policy engines: validate, then mutate, then generate
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-requests
spec:
validationFailureAction: Audit # start here, never Enforce
rules:
- name: resources-set
match:
any:
- resources: { kinds: [Pod] }
validate:
message: "every container needs cpu and memory requests"
pattern:
spec:
containers:
- resources:
requests:
memory: "?*"
cpu: "?*"
The underrated half is mutation and generation: default a security context rather than rejecting pods that lack one, or generate a default-deny NetworkPolicy into every new namespace automatically. A policy that fixes the problem produces far less friction than one that lectures about it.
The failure mode to design for
failurePolicy: Fail that becomes unavailable will block every matching API write in the cluster. If its scope includes namespaces, and the webhook pod needs a namespace, you have a deadlock that survives restarts. This has taken down real clusters. Always exclude kube-system and your policy engine’s own namespace, and think hard before Fail.# the exclusion that saves you
namespaceSelector:
matchExpressions:
- key: kubernetes.io/metadata.name
operator: NotIn
values: [kube-system, kyverno]
Policy that people do not route around
- Audit before enforce, always. Ship the rule in audit mode, read what it would have blocked, then decide.
- Error messages with the fix in them. “policy violation” gets you a ticket; “add resources.requests.memory, see <link>” gets you a fixed manifest.
- Exceptions as objects, with expiry. A documented, time-limited exception beats a policy quietly weakened for everyone.
- Test in CI. Both engines can evaluate policies against manifests without a cluster, so a developer finds out in their pull request, not at deploy.
What to actually do with this
- Add PSA labels in warn mode to every namespace today. It costs nothing and breaks nothing.
- Read the warnings for a week before enforcing anything.
- Check the
failurePolicyand namespace exclusions on every webhook already in your cluster.
Something wrong or out of date? Open an issue — corrections are welcome and get credited.