Capacity, bin-packing and autoscaling
Requests and limits chosen from data, plus node autoscaling that does not thrash.
Requests and limits are the two most consequential numbers in a pod spec and the two most often copied from an example. They do completely different jobs, and confusing them produces either a cluster that is 80% idle or one that evicts things at random.
Requests schedule. Limits constrain.
- Requests are what the scheduler reserves. They decide where a pod goes and how full a node is considered. They are not enforced at runtime.
- Limits are enforced by the kernel. Over a CPU limit, you are throttled; over a memory limit, you are killed.
So a pod requesting 100m CPU and limited to 2 cores may be scheduled onto a busy node and then try to use twenty times what it reserved. That works until the node is contended, at which point everything on it gets slow together.
QoS classes, and who dies first
- Guaranteed — requests equal limits for every container. Evicted last.
- Burstable — requests set, limits higher or absent. Evicted in the middle, worst-offender first.
- BestEffort — nothing set. Evicted first, always.
The CPU-limit argument
There is a real case for setting no CPU limit at all. A CPU limit throttles in fixed periods, and a latency-sensitive service with a modest limit can be throttled while the node is mostly idle — you have taken a latency hit to prevent a problem that was not occurring. The counter-argument is predictability and noisy neighbours.
# the signal that this is happening to you
rate(container_cpu_cfs_throttled_seconds_total[5m]) > 0
A defensible position: always set CPU requests accurately, set memory requests and limits equal, and leave CPU limits off for latency-sensitive services while keeping them for batch. Memory is different because it is incompressible — there is no graceful degradation, only the OOM killer.
Get the numbers from the cluster, not from a guess
kubectl top pods -n prod --sort-by=memory
# p95 actual usage over a week, which is what a request should be based on
quantile_over_time(0.95,
container_memory_working_set_bytes{namespace="prod"}[7d])
# requested versus used, per namespace - your idle percentage
sum(kube_pod_container_resource_requests{resource="cpu"}) by (namespace)
/ sum(rate(container_cpu_usage_seconds_total[1h])) by (namespace)
VPA in recommendation mode is worth running for this even if you never let it act: it watches real usage and tells you what it would set, which is a better starting point than anyone’s intuition.
Node autoscaling, and consolidation
The Cluster Autoscaler adds nodes from predefined groups when pods are unschedulable. Karpenter instead picks instance types to fit the pending pods, and continuously consolidates — replacing several underused nodes with one cheaper node.
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: { name: api }
spec:
minAvailable: 2 # or maxUnavailable: 1
selector:
matchLabels: { app: api }
Also mind the interaction with topology spread: a pod spread across three zones with whenUnsatisfiable: DoNotSchedule constrains which nodes the autoscaler can usefully add. Unschedulable pods with an autoscaler that is doing nothing is almost always a constraint it cannot satisfy, not a broken autoscaler.
What to actually do with this
- Find every pod with no resource fields. Those are BestEffort and will die first.
- Set memory requests equal to limits on anything stateful.
- Check for CFS throttling on your latency-sensitive services.
- Write PDBs before enabling consolidation, not after the first incident.
Something wrong or out of date? Open an issue — corrections are welcome and get credited.