Platform and scale

Capacity, bin-packing and autoscaling

Requests and limits chosen from data, plus node autoscaling that does not thrash.

SchedulingCKA 10 min read

Requests and limits are the two most consequential numbers in a pod spec and the two most often copied from an example. They do completely different jobs, and confusing them produces either a cluster that is 80% idle or one that evicts things at random.

Requests schedule. Limits constrain.

  • Requests are what the scheduler reserves. They decide where a pod goes and how full a node is considered. They are not enforced at runtime.
  • Limits are enforced by the kernel. Over a CPU limit, you are throttled; over a memory limit, you are killed.

So a pod requesting 100m CPU and limited to 2 cores may be scheduled onto a busy node and then try to use twenty times what it reserved. That works until the node is contended, at which point everything on it gets slow together.

QoS classes, and who dies first

  • Guaranteed — requests equal limits for every container. Evicted last.
  • Burstable — requests set, limits higher or absent. Evicted in the middle, worst-offender first.
  • BestEffort — nothing set. Evicted first, always.
This is a decision you make by accident if you do not make it deliberately. A production database with no resource fields is BestEffort, which means it is the first thing the kubelet kills under node pressure. Set requests equal to limits on anything you genuinely cannot lose.

The CPU-limit argument

There is a real case for setting no CPU limit at all. A CPU limit throttles in fixed periods, and a latency-sensitive service with a modest limit can be throttled while the node is mostly idle — you have taken a latency hit to prevent a problem that was not occurring. The counter-argument is predictability and noisy neighbours.

# the signal that this is happening to you
rate(container_cpu_cfs_throttled_seconds_total[5m]) > 0

A defensible position: always set CPU requests accurately, set memory requests and limits equal, and leave CPU limits off for latency-sensitive services while keeping them for batch. Memory is different because it is incompressible — there is no graceful degradation, only the OOM killer.

Get the numbers from the cluster, not from a guess

kubectl top pods -n prod --sort-by=memory

# p95 actual usage over a week, which is what a request should be based on
quantile_over_time(0.95,
  container_memory_working_set_bytes{namespace="prod"}[7d])

# requested versus used, per namespace - your idle percentage
sum(kube_pod_container_resource_requests{resource="cpu"}) by (namespace)
  / sum(rate(container_cpu_usage_seconds_total[1h])) by (namespace)

VPA in recommendation mode is worth running for this even if you never let it act: it watches real usage and tells you what it would set, which is a better starting point than anyone’s intuition.

Node autoscaling, and consolidation

The Cluster Autoscaler adds nodes from predefined groups when pods are unschedulable. Karpenter instead picks instance types to fit the pending pods, and continuously consolidates — replacing several underused nodes with one cheaper node.

Consolidation moves running pods, and it will respect only what you have told it to respect. Without a PodDisruptionBudget, consolidating a node can take down every replica of a service at once, legitimately, because nothing said otherwise. Every workload that matters needs a PDB before you enable consolidation.
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: { name: api }
spec:
  minAvailable: 2            # or maxUnavailable: 1
  selector:
    matchLabels: { app: api }

Also mind the interaction with topology spread: a pod spread across three zones with whenUnsatisfiable: DoNotSchedule constrains which nodes the autoscaler can usefully add. Unschedulable pods with an autoscaler that is doing nothing is almost always a constraint it cannot satisfy, not a broken autoscaler.

What to actually do with this

  • Find every pod with no resource fields. Those are BestEffort and will die first.
  • Set memory requests equal to limits on anything stateful.
  • Check for CFS throttling on your latency-sensitive services.
  • Write PDBs before enabling consolidation, not after the first incident.
This article covers one checkpoint on the roadmap. Open Capacity, bin-packing and autoscaling on the roadmap → — it lists what this depends on and everything else written about it.

Something wrong or out of date? Open an issue — corrections are welcome and get credited.