SLOs people actually use
Three signals per service, one dashboard, and alerts that correspond to someone being paged.
Most Kubernetes monitoring setups collect everything and tell you nothing. The failure is not technical — it is that nobody decided what "working" means before building the dashboards.
Delete half your alerts first
Before adding anything, go through what fires today and ask one question of each: when this fired, did a human do something? If the honest answer is no, it is not an alert. It is a dashboard panel at best.
An alert nobody acts on is worse than no alert, because it trains the team to ignore the channel where the real one will arrive. The strongest predictor of whether monitoring works is not coverage, it is how many pages people trust.
Three signals is enough
- Availability — the fraction of requests that did not fail.
- Latency — the fraction served faster than a threshold you chose deliberately.
- Saturation — how close the thing is to a limit, so you have warning before the first two move.
# availability, as a ratio - not a count of errors
sum(rate(http_requests_total{status!~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
# latency, as a ratio of requests inside the target
sum(rate(http_request_duration_seconds_bucket{le="0.5"}[5m]))
/ sum(rate(http_request_duration_seconds_count[5m]))
Alert on burn rate, not on breach
A 99.9% monthly objective gives you about 43 minutes of error budget. Alerting the moment you dip below 99.9% in a five-minute window is pure noise. Alerting when you are consuming the budget fast enough to exhaust it is the signal.
# fast burn: 2% of a 30-day budget in an hour -> page
- alert: ErrorBudgetBurningFast
expr: |
(1 - (sum(rate(http_requests_total{status!~"5.."}[1h]))
/ sum(rate(http_requests_total[1h])))) > 14.4 * 0.001
for: 2m
labels: { severity: page }
# slow burn: on course to exhaust it this week -> ticket, not a page
- alert: ErrorBudgetBurningSlow
expr: |
(1 - (sum(rate(http_requests_total{status!~"5.."}[6h]))
/ sum(rate(http_requests_total[6h])))) > 6 * 0.001
for: 30m
labels: { severity: ticket }
Two severities, two response paths. The fast burn wakes someone; the slow burn becomes work on Monday. That distinction is what makes an on-call rotation survivable.
Cluster-level alerts worth keeping
- A node
NotReadyfor more than five minutes. - Pods in
CrashLoopBackOfffor more than fifteen. - PersistentVolume above 85% full — with hours of warning, not minutes.
- Certificates expiring within fourteen days.
- A Deployment with fewer ready replicas than its PDB requires.
Notably absent: anything about individual pod restarts, CPU above a threshold, or memory above a threshold. Those are symptoms, they fire constantly, and they are what dashboards are for.
Cardinality is how this gets expensive
user_id, a request path with an id in it, or a pod name on a high-churn Deployment creates a new time series per value, forever. Normalise paths to route templates before they reach the metric, and keep ids in logs and traces where they belong.What to actually do with this
- List every alert that fired last month and delete the ones nobody acted on.
- Pick one service and write down its three signals and one objective.
- Replace one threshold alert with a burn-rate alert and compare the noise.
- Find your highest-cardinality metric before your bill does.
Something wrong or out of date? Open an issue โ corrections are welcome and get credited.