Getting a GPU into a pod
The device plugin model, drivers, and the difference between MIG and time-slicing when cost is the constraint.
No certification covers this, which is exactly why it is worth writing down. The mechanics are not hard; the failure modes are just unfamiliar, and every layer fails silently in its own way.
GPUs are not like CPU and memory
CPU and memory are compressible, divisible, and known to the kubelet natively. A GPU is an extended resource: an opaque integer advertised by a device plugin, which cannot be fractional and cannot be overcommitted.
resources:
limits:
nvidia.com/gpu: 1 # requests is implied and must equal limits
# what the node claims to have
kubectl get nodes -o custom-columns=\
'NODE:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu'
If that column is empty or zero, no pod requesting a GPU will ever schedule, and the only symptom is Pending. Nothing tells you the plugin is unhealthy; the resource simply does not exist.
Six layers, and the error is always in one of them
- Hardware — the card is present.
lspci | grep -i nvidia. - Kernel driver — loaded and matching.
nvidia-smion the node. - Container runtime — configured to inject devices, usually the NVIDIA container toolkit.
- Device plugin — a DaemonSet advertising the resource to the kubelet.
- Node labels — so you can target specific hardware.
- The pod spec — requesting the resource at all.
# the debugging order, top down
kubectl get ds -n gpu-operator # plugin running?
kubectl logs -n gpu-operator ds/nvidia-device-plugin-daemonset --tail=50
kubectl describe node gpu-node-1 | grep -A5 Allocatable
kubectl run smi --rm -it --restart=Never \
--image=nvidia/cuda:12.4.0-base-ubuntu22.04 \
--overrides='{"spec":{"containers":[{"name":"smi","image":"nvidia/cuda:12.4.0-base-ubuntu22.04","command":["nvidia-smi"],"resources":{"limits":{"nvidia.com/gpu":1}}}]}}'
Sharing one GPU: two mechanisms, different guarantees
- Time-slicing — the GPU context-switches between processes. Memory is not partitioned: every process sees the whole card and they can collectively exhaust it. No isolation, works on any GPU, zero reconfiguration.
- MIG (Multi-Instance GPU) — the hardware is partitioned into instances with dedicated memory, cache and compute. Real isolation, enforced by the card. Only on A100/H100-class hardware, and the partition layout is fixed until you reconfigure the device.
--gpu-memory-utilization for vLLM) because Kubernetes cannot do it for you.# time-slicing: four virtual GPUs from one physical card
apiVersion: v1
kind: ConfigMap
metadata:
name: device-plugin-config
data:
config.yaml: |
sharing:
timeSlicing:
resources:
- name: nvidia.com/gpu
replicas: 4
# MIG: request a specific profile instead
# resources:
# limits:
# nvidia.com/mig-1g.10gb: 1
Choosing between them
- Many small, bursty, trusted workloads — notebooks, batch inference, development — time-slicing, with application memory caps.
- Multiple tenants, or anything with a latency target — MIG. A noisy neighbour on a time-sliced GPU can make your p99 unrecognisable.
- Training — neither. Give the job whole GPUs.
What to actually do with this
- Run the six-layer check on a working GPU node so you know what healthy looks like.
- Confirm whether your cards support MIG before designing around it.
- If you time-slice, set a memory cap in the application today.
Something wrong or out of date? Open an issue โ corrections are welcome and get credited.