Scaling inference on real traffic
Queue-depth autoscaling, cold starts measured in minutes, and the bill as a design constraint rather than an afterthought.
Autoscaling a stateless web service is a solved problem with good defaults. Autoscaling inference breaks all of those defaults at once, and the reason is simple: a saturated GPU does not look busy to anything Kubernetes measures by default.
Why CPU-based HPA fails
A model server at 100% GPU utilisation with a queue thirty requests deep is typically using 15% CPU. The HPA sees an idle pod and scales down. Meanwhile latency is climbing and nothing in the control loop knows.
- GPU utilisation is better, but it is a poor saturation signal — a GPU reads near 100% while still having headroom for more concurrent sequences.
- Queue depth (
num_requests_waiting) is the honest signal. A non-zero queue means demand exceeds capacity, right now. - Time to first token is what users feel, and it is the right thing to alert on even if you scale on queue depth.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: { name: llm }
spec:
scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: llm }
minReplicas: 1
maxReplicas: 8
metrics:
- type: Pods
pods:
metric: { name: vllm_num_requests_waiting }
target: { type: AverageValue, averageValue: "4" }
behavior:
scaleUp:
stabilizationWindowSeconds: 0 # react immediately
policies: [{ type: Pods, value: 2, periodSeconds: 60 }]
scaleDown:
stabilizationWindowSeconds: 600 # 10 minutes of calm before shrinking
policies: [{ type: Pods, value: 1, periodSeconds: 300 }]
behavior is the important part and it is not the default. Scaling up late costs you a latency spike; scaling down early costs you a multi-minute cold start on the next request. When those are the two errors available, you want to be slow to shrink and quick to grow.Cold start is the whole design constraint
Add it up honestly: node provisioning 60–180 seconds, image pull 30–120, model load 40–300, warm-up 10–30. Three to ten minutes from "we need another replica" to "it can serve". No autoscaler can hide that.
- Keep a warm pool.
minReplicasabove zero, sized to your trough, is the only reliable answer for interactive traffic. - Pre-pull images with a DaemonSet or a node image that already has them. Removes the largest variable chunk.
- Over-provision deliberately with low-priority placeholder pods, so a node is already warm and a real pod can preempt a placeholder instantly.
- Separate interactive from batch. Batch work tolerates cold starts and can use spot capacity; user-facing traffic cannot.
Scale to zero, honestly
Scale to zero is attractive and only workable when something can hold the first request while a replica starts, and the caller tolerates minutes. That is true for internal tools and batch endpoints. It is not true for anything a person is waiting on. Decide which you have before adopting it.
The bill as a design constraint
An idle A100 costs roughly the same as a busy one. Utilisation is therefore the only cost lever that matters, and the arithmetic is unforgiving: a GPU at 20% utilisation is a GPU you are mostly paying to keep warm.
# the number that should be on a dashboard: tokens per GPU-hour
sum(rate(vllm_generation_tokens_total[1h]))
/ count(DCGM_FI_DEV_GPU_UTIL)
# and the one that explains it: how much of the time are GPUs doing nothing?
avg_over_time(DCGM_FI_DEV_GPU_UTIL[24h])
- Batch aggressively. Continuous batching in vLLM is the single largest throughput win available, and it is a flag.
- Right-size the model. An 8B model that fits one GPU often beats a 70B model across four, for tasks where quality is comparable.
- Quantise when accuracy allows — fewer GPUs for the same throughput.
- Spot for batch. With checkpointing, interruptions are a cost, not a failure.
maxSurge on a GPU Deployment. A rolling update with surge needs a spare GPU for the new pod before the old one goes away. On a cluster with no idle GPU, the new pod sits Pending, the old one never terminates, and the rollout hangs until progressDeadlineSeconds expires. Set maxSurge: 0 and accept brief unavailability, or keep a spare GPU for deploys.What to actually do with this
- Stop scaling on CPU. Export queue depth and scale on that.
- Measure your real cold start, end to end, and set
minReplicasfrom it. - Put tokens per GPU-hour on a dashboard. It changes conversations about cost.
- Check
maxSurgeon every GPU Deployment you own.
Something wrong or out of date? Open an issue โ corrections are welcome and get credited.