Security

Admission control and policy as code

Pod Security Admission first, then a policy engine for the rules PSA cannot express and a webhook only when nothing else will do.

CKSSecurityKCSA 10 min read

Admission control is the last point at which you can say no. It runs after authentication and authorisation, after the object has been validated, and before anything is written to etcd — so a rejection means the bad object never existed.

Three tiers, in order of how much they cost you

  1. Pod Security Admission — built in, three labels, no components to run. Covers the dangerous pod fields.
  2. A policy engine (Kyverno, Gatekeeper) — declarative rules for everything else, including mutation and cross-object checks.
  3. A custom webhook — arbitrary code, maximum power, and now you own an availability problem.

Most teams reach straight for tier two and skip tier one, then write forty policies reimplementing what three labels would have done.

Pod Security Admission, which is free

apiVersion: v1
kind: Namespace
metadata:
  name: team-a
  labels:
    pod-security.kubernetes.io/enforce: baseline     # reject the worst
    pod-security.kubernetes.io/enforce-version: v1.31
    pod-security.kubernetes.io/audit: restricted     # log what would fail
    pod-security.kubernetes.io/warn: restricted      # warn the applier
The sequence that works: set warn and audit to restricted while enforce stays at baseline. You get the full list of what would break, in warnings and the audit log, without breaking it. Tighten enforce once that list is empty.

baseline blocks privileged containers, host namespaces, hostPath and dangerous capabilities. restricted additionally requires non-root, dropping all capabilities and a seccomp profile. The gap between them is where almost every real application argument happens.

Policy engines: validate, then mutate, then generate

apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
  name: require-requests
spec:
  validationFailureAction: Audit      # start here, never Enforce
  rules:
    - name: resources-set
      match:
        any:
          - resources: { kinds: [Pod] }
      validate:
        message: "every container needs cpu and memory requests"
        pattern:
          spec:
            containers:
              - resources:
                  requests:
                    memory: "?*"
                    cpu: "?*"

The underrated half is mutation and generation: default a security context rather than rejecting pods that lack one, or generate a default-deny NetworkPolicy into every new namespace automatically. A policy that fixes the problem produces far less friction than one that lectures about it.

The failure mode to design for

A validating webhook with failurePolicy: Fail that becomes unavailable will block every matching API write in the cluster. If its scope includes namespaces, and the webhook pod needs a namespace, you have a deadlock that survives restarts. This has taken down real clusters. Always exclude kube-system and your policy engine’s own namespace, and think hard before Fail.
# the exclusion that saves you
namespaceSelector:
  matchExpressions:
    - key: kubernetes.io/metadata.name
      operator: NotIn
      values: [kube-system, kyverno]

Policy that people do not route around

  • Audit before enforce, always. Ship the rule in audit mode, read what it would have blocked, then decide.
  • Error messages with the fix in them. “policy violation” gets you a ticket; “add resources.requests.memory, see <link>” gets you a fixed manifest.
  • Exceptions as objects, with expiry. A documented, time-limited exception beats a policy quietly weakened for everyone.
  • Test in CI. Both engines can evaluate policies against manifests without a cluster, so a developer finds out in their pull request, not at deploy.

What to actually do with this

  • Add PSA labels in warn mode to every namespace today. It costs nothing and breaks nothing.
  • Read the warnings for a week before enforcing anything.
  • Check the failurePolicy and namespace exclusions on every webhook already in your cluster.
This article covers one checkpoint on the roadmap. Open Admission control and policy as code on the roadmap → — it lists what this depends on and everything else written about it.

Something wrong or out of date? Open an issue — corrections are welcome and get credited.