AI inference and GPUs

Getting a GPU into a pod

The device plugin model, drivers, and the difference between MIG and time-slicing when cost is the constraint.

InferenceScheduling 11 min read

No certification covers this, which is exactly why it is worth writing down. The mechanics are not hard; the failure modes are just unfamiliar, and every layer fails silently in its own way.

GPUs are not like CPU and memory

CPU and memory are compressible, divisible, and known to the kubelet natively. A GPU is an extended resource: an opaque integer advertised by a device plugin, which cannot be fractional and cannot be overcommitted.

resources:
  limits:
    nvidia.com/gpu: 1        # requests is implied and must equal limits

# what the node claims to have
kubectl get nodes -o custom-columns=\
  'NODE:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu'

If that column is empty or zero, no pod requesting a GPU will ever schedule, and the only symptom is Pending. Nothing tells you the plugin is unhealthy; the resource simply does not exist.

Six layers, and the error is always in one of them

  1. Hardware — the card is present. lspci | grep -i nvidia.
  2. Kernel driver — loaded and matching. nvidia-smi on the node.
  3. Container runtime — configured to inject devices, usually the NVIDIA container toolkit.
  4. Device plugin — a DaemonSet advertising the resource to the kubelet.
  5. Node labels — so you can target specific hardware.
  6. The pod spec — requesting the resource at all.
# the debugging order, top down
kubectl get ds -n gpu-operator                        # plugin running?
kubectl logs -n gpu-operator ds/nvidia-device-plugin-daemonset --tail=50
kubectl describe node gpu-node-1 | grep -A5 Allocatable
kubectl run smi --rm -it --restart=Never \
  --image=nvidia/cuda:12.4.0-base-ubuntu22.04 \
  --overrides='{"spec":{"containers":[{"name":"smi","image":"nvidia/cuda:12.4.0-base-ubuntu22.04","command":["nvidia-smi"],"resources":{"limits":{"nvidia.com/gpu":1}}}]}}'
Use the GPU Operator rather than assembling layers two to five by hand. It manages drivers, the toolkit, the device plugin, Node Feature Discovery and DCGM metrics as one unit, and it handles the thing that is genuinely painful manually: driver upgrades across a fleet without a cordon-and-reboot script of your own.

Sharing one GPU: two mechanisms, different guarantees

  • Time-slicing — the GPU context-switches between processes. Memory is not partitioned: every process sees the whole card and they can collectively exhaust it. No isolation, works on any GPU, zero reconfiguration.
  • MIG (Multi-Instance GPU) — the hardware is partitioned into instances with dedicated memory, cache and compute. Real isolation, enforced by the card. Only on A100/H100-class hardware, and the partition layout is fixed until you reconfigure the device.
Time-slicing plus an untuned inference server is an OOM factory. Frameworks like vLLM pre-allocate a large fraction of visible GPU memory by default. Two time-sliced replicas each try to reserve 90% of the same card, and the second one dies — or worse, both survive startup and one dies later under load. If you time-slice, you must cap memory per process in the application (--gpu-memory-utilization for vLLM) because Kubernetes cannot do it for you.
# time-slicing: four virtual GPUs from one physical card
apiVersion: v1
kind: ConfigMap
metadata:
  name: device-plugin-config
data:
  config.yaml: |
    sharing:
      timeSlicing:
        resources:
          - name: nvidia.com/gpu
            replicas: 4

# MIG: request a specific profile instead
# resources:
#   limits:
#     nvidia.com/mig-1g.10gb: 1

Choosing between them

  • Many small, bursty, trusted workloads — notebooks, batch inference, development — time-slicing, with application memory caps.
  • Multiple tenants, or anything with a latency target — MIG. A noisy neighbour on a time-sliced GPU can make your p99 unrecognisable.
  • Training — neither. Give the job whole GPUs.

What to actually do with this

  • Run the six-layer check on a working GPU node so you know what healthy looks like.
  • Confirm whether your cards support MIG before designing around it.
  • If you time-slice, set a memory cap in the application today.
This article covers one checkpoint on the roadmap. Open Getting a GPU into a pod on the roadmap → โ€” it lists what this depends on and everything else written about it.

Something wrong or out of date? Open an issue โ€” corrections are welcome and get credited.