AI inference and GPUs

Serving a model like a production service

vLLM behind a Gateway, with readiness that reflects model load rather than process start — which is where most first attempts go wrong.

InferenceNetworking 12 min read

A model server is an ordinary HTTP service with three unusual properties: it takes minutes to become useful, it holds a lot of state in GPU memory, and its requests vary in cost by two orders of magnitude. Every default in Kubernetes is tuned for services with none of those properties.

A minimal vLLM deployment, annotated

apiVersion: apps/v1
kind: Deployment
metadata: { name: llm }
spec:
  replicas: 1
  template:
    spec:
      containers:
        - name: vllm
          image: vllm/vllm-openai:latest
          args:
            - --model=/models/llama-3.1-8b
            - --gpu-memory-utilization=0.90
            - --max-model-len=8192
          ports: [{ containerPort: 8000 }]
          resources:
            limits: { nvidia.com/gpu: 1, memory: 32Gi }
          volumeMounts:
            - { name: models, mountPath: /models, readOnly: true }
            - { name: shm, mountPath: /dev/shm }
      volumes:
        - name: models
          persistentVolumeClaim: { claimName: model-weights }
        - name: shm                      # NOT optional
          emptyDir: { medium: Memory, sizeLimit: 8Gi }
That /dev/shm volume is the single most common missing line. The default shared-memory segment in a container is 64 MB. PyTorch uses shared memory for inter-process communication, so anything with tensor parallelism or multiple workers crashes with an opaque "bus error" or a NCCL failure that reads like a GPU problem. It is not — it is 64 MB of /dev/shm.

Model load time breaks every probe default

Loading 16 GB of weights from a network volume into GPU memory takes anywhere from 40 seconds to 6 minutes. With default probes, the container is killed and restarted before it finishes, forever, and the logs show a half-finished load each time.

startupProbe:
  httpGet: { path: /health, port: 8000 }
  failureThreshold: 60
  periodSeconds: 10          # up to 10 minutes to load

readinessProbe:
  httpGet: { path: /health, port: 8000 }
  periodSeconds: 5

livenessProbe:
  httpGet: { path: /health, port: 8000 }
  periodSeconds: 30
  failureThreshold: 3        # 90s of failure before a restart that costs minutes

Be conservative with liveness specifically. Restarting a model server is not cheap — you pay the full load time again — so a liveness probe that trips on a transient stall makes an incident worse rather than better.

Where the weights live

  • Baked into the image — simplest, and gives you a 20 GB image to pull on every new node. Fine for one fixed model.
  • A ReadOnlyMany volume — pulled once, shared by replicas on the same node. Usually the right default.
  • Downloaded by an init container — flexible, and every pod start depends on an external registry being up.
  • An OCI artefact — weights as a separate layer, cached by the node like any image. Increasingly the tidiest answer.

Whichever you choose, the first pod on a new node pays the cold cost. That number is the floor on your scale-up latency, and it is the number to measure before you design autoscaling.

Routing: why a plain Service is wrong

A ClusterIP distributes connections roughly evenly and knows nothing about what it is balancing. For LLM traffic that is actively harmful: requests differ enormously in cost, and a replica with a warm KV cache for a given prefix is dramatically cheaper to use than a cold one.

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata: { name: llm }
spec:
  parentRefs: [{ name: ai-gateway }]
  rules:
    - matches: [{ path: { type: PathPrefix, value: /v1 } }]
      timeouts:
        request: 300s            # generation is slow; the default will cut it off
      backendRefs:
        - { name: llm, port: 8000 }

The timeout is the immediate fix — most ingress defaults are 30 or 60 seconds and will truncate long generations. Beyond that, the Gateway API Inference Extension adds model-aware routing: queue depth, KV-cache awareness, and routing by model name to different backends behind one endpoint. If you are serving more than one model, that is the thing to look at rather than building it yourself.

Streaming changes the whole request path

Token streaming is server-sent events over a long-lived connection. Anything in the path that buffers responses, or closes idle connections, breaks it in a way that looks like a model problem: the client waits, then gets everything at once, or nothing. Check buffering and idle timeouts on every proxy between the client and the pod before debugging the server.

What to actually do with this

  • Add the /dev/shm volume now, whether or not you have hit the bug.
  • Measure your actual cold start from pod creation to first token, and write it down.
  • Set a startupProbe with a failure threshold based on that measurement.
  • Raise the request timeout on every proxy in front of the model.
This article covers one checkpoint on the roadmap. Open Serving a model like a production service on the roadmap → โ€” it lists what this depends on and everything else written about it.

Something wrong or out of date? Open an issue โ€” corrections are welcome and get credited.