Requesting a GPU¶
Attaching a GPU (or another PCI device) to a worker in your cluster.
Nothing about the device is yours to configure: your platform team decides which devices exist and what they are called. Your part is naming one on your Cluster object, in your Environment.
Requesting one¶
Add a gpus entry to your Cluster's topology variables:
apiVersion: cluster.x-k8s.io/v1beta2
kind: Cluster
spec:
topology:
variables:
- name: gpus
value:
- name: gpu0
deviceName: nvidia.com/A10
deviceName has to match a resourceName your platform team has already whitelisted. Ask them what is available rather than guessing; an unrecognized deviceName does not fail your kubectl apply, it leaves the affected worker Pending with nothing pointing at why.
Requesting more than one, or a mixed pool¶
gpus can be overridden per MachineDeployment, so only the pool that needs the device carries the cost of one:
workers:
machineDeployments:
- class: default-worker
name: workers
replicas: 3
- class: default-worker
name: gpu-workers
replicas: 1
variables:
overrides:
- name: gpus
value:
- name: gpu0
deviceName: nvidia.com/A10
What to expect once it is attached¶
- No live migration. A worker holding a passed-through device cannot move to another node while running. Losing that node means the Machine is recreated, not migrated.
- No virtual display. The platform disables it by default, since most GPUs (real or virtual) do not support one. Nothing about the workload running inside the VM changes because of this.
- A real PCI device, not a proxy. Once attached, the device shows up in the guest exactly as a real PCI device would, addressable with ordinary tools like
lspci.
GPUs count against a quota¶
The platform team can cap how many devices a tenant holds at once, the same way clusters or storage are capped. Adding a gpus entry does not fail the kubectl apply when the tenant is over that limit: the Cluster is accepted, and the affected worker's Machine (or its VirtualMachineInstance) is the thing that fails, with an event naming the quota rather than the deviceName. When a worker never comes up on a device name that has worked before, check with the platform team whether the tenant is at its device quota before assuming the device itself is broken.
A GPU that never attaches¶
| Symptom | Cause |
|---|---|
Machine stuck Pending, no useful event |
deviceName does not match anything whitelisted on the platform. Confirm the exact string with the platform team. |
| Machine or VM fails with a quota-related event | The tenant is at its device quota for this resourceName. Raise it with the platform team, or free one up elsewhere in the tenant. |
| Machine schedules, but on the wrong hardware | More than one node carries a device under that same name. Ask whether the pool needs pinning to a specific node. |
VM never reaches Running |
The device itself may be unhealthy on that node. This is a platform-side issue to report, not something visible from inside the cluster. |
See Also: Deploying Applications ยท GPU & Device Passthrough for how the platform whitelists devices in the first place