Serving a model like a production service
vLLM behind a Gateway, with readiness that reflects model load rather than process start — which is where most first attempts go wrong.
A model server is an ordinary HTTP service with three unusual properties: it takes minutes to become useful, it holds a lot of state in GPU memory, and its requests vary in cost by two orders of magnitude. Every default in Kubernetes is tuned for services with none of those properties.
A minimal vLLM deployment, annotated
apiVersion: apps/v1
kind: Deployment
metadata: { name: llm }
spec:
replicas: 1
template:
spec:
containers:
- name: vllm
image: vllm/vllm-openai:latest
args:
- --model=/models/llama-3.1-8b
- --gpu-memory-utilization=0.90
- --max-model-len=8192
ports: [{ containerPort: 8000 }]
resources:
limits: { nvidia.com/gpu: 1, memory: 32Gi }
volumeMounts:
- { name: models, mountPath: /models, readOnly: true }
- { name: shm, mountPath: /dev/shm }
volumes:
- name: models
persistentVolumeClaim: { claimName: model-weights }
- name: shm # NOT optional
emptyDir: { medium: Memory, sizeLimit: 8Gi }
/dev/shm volume is the single most common missing line. The default shared-memory segment in a container is 64 MB. PyTorch uses shared memory for inter-process communication, so anything with tensor parallelism or multiple workers crashes with an opaque "bus error" or a NCCL failure that reads like a GPU problem. It is not — it is 64 MB of /dev/shm.Model load time breaks every probe default
Loading 16 GB of weights from a network volume into GPU memory takes anywhere from 40 seconds to 6 minutes. With default probes, the container is killed and restarted before it finishes, forever, and the logs show a half-finished load each time.
startupProbe:
httpGet: { path: /health, port: 8000 }
failureThreshold: 60
periodSeconds: 10 # up to 10 minutes to load
readinessProbe:
httpGet: { path: /health, port: 8000 }
periodSeconds: 5
livenessProbe:
httpGet: { path: /health, port: 8000 }
periodSeconds: 30
failureThreshold: 3 # 90s of failure before a restart that costs minutes
Be conservative with liveness specifically. Restarting a model server is not cheap — you pay the full load time again — so a liveness probe that trips on a transient stall makes an incident worse rather than better.
Where the weights live
- Baked into the image — simplest, and gives you a 20 GB image to pull on every new node. Fine for one fixed model.
- A ReadOnlyMany volume — pulled once, shared by replicas on the same node. Usually the right default.
- Downloaded by an init container — flexible, and every pod start depends on an external registry being up.
- An OCI artefact — weights as a separate layer, cached by the node like any image. Increasingly the tidiest answer.
Whichever you choose, the first pod on a new node pays the cold cost. That number is the floor on your scale-up latency, and it is the number to measure before you design autoscaling.
Routing: why a plain Service is wrong
A ClusterIP distributes connections roughly evenly and knows nothing about what it is balancing. For LLM traffic that is actively harmful: requests differ enormously in cost, and a replica with a warm KV cache for a given prefix is dramatically cheaper to use than a cold one.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata: { name: llm }
spec:
parentRefs: [{ name: ai-gateway }]
rules:
- matches: [{ path: { type: PathPrefix, value: /v1 } }]
timeouts:
request: 300s # generation is slow; the default will cut it off
backendRefs:
- { name: llm, port: 8000 }
The timeout is the immediate fix — most ingress defaults are 30 or 60 seconds and will truncate long generations. Beyond that, the Gateway API Inference Extension adds model-aware routing: queue depth, KV-cache awareness, and routing by model name to different backends behind one endpoint. If you are serving more than one model, that is the thing to look at rather than building it yourself.
Streaming changes the whole request path
Token streaming is server-sent events over a long-lived connection. Anything in the path that buffers responses, or closes idle connections, breaks it in a way that looks like a model problem: the client waits, then gets everything at once, or nothing. Check buffering and idle timeouts on every proxy between the client and the pod before debugging the server.
What to actually do with this
- Add the
/dev/shmvolume now, whether or not you have hit the bug. - Measure your actual cold start from pod creation to first token, and write it down.
- Set a
startupProbewith a failure threshold based on that measurement. - Raise the request timeout on every proxy in front of the model.
Something wrong or out of date? Open an issue โ corrections are welcome and get credited.