Skip to content

Requesting a GPU

Attaching a GPU (or another PCI device) to a worker in your cluster.

Nothing about the device is yours to configure: your platform team decides which devices exist and what they are called. Your part is naming one on your Cluster object, in your Environment.

Requesting one

Add a gpus entry to your Cluster's topology variables:

apiVersion: cluster.x-k8s.io/v1beta2
kind: Cluster
spec:
  topology:
    variables:
      - name: gpus
        value:
          - name: gpu0
            deviceName: nvidia.com/A10

deviceName has to match a resourceName your platform team has already whitelisted. Ask them what is available rather than guessing; an unrecognized deviceName does not fail your kubectl apply, it leaves the affected worker Pending with nothing pointing at why.

Requesting more than one, or a mixed pool

gpus can be overridden per MachineDeployment, so only the pool that needs the device carries the cost of one:

workers:
  machineDeployments:
    - class: default-worker
      name: workers
      replicas: 3
    - class: default-worker
      name: gpu-workers
      replicas: 1
      variables:
        overrides:
          - name: gpus
            value:
              - name: gpu0
                deviceName: nvidia.com/A10

What to expect once it is attached

  • No live migration. A worker holding a passed-through device cannot move to another node while running. Losing that node means the Machine is recreated, not migrated.
  • No virtual display. The platform disables it by default, since most GPUs (real or virtual) do not support one. Nothing about the workload running inside the VM changes because of this.
  • A real PCI device, not a proxy. Once attached, the device shows up in the guest exactly as a real PCI device would, addressable with ordinary tools like lspci.

GPUs count against a quota

The platform team can cap how many devices a tenant holds at once, the same way clusters or storage are capped. Adding a gpus entry does not fail the kubectl apply when the tenant is over that limit: the Cluster is accepted, and the affected worker's Machine (or its VirtualMachineInstance) is the thing that fails, with an event naming the quota rather than the deviceName. When a worker never comes up on a device name that has worked before, check with the platform team whether the tenant is at its device quota before assuming the device itself is broken.

A GPU that never attaches

Symptom Cause
Machine stuck Pending, no useful event deviceName does not match anything whitelisted on the platform. Confirm the exact string with the platform team.
Machine or VM fails with a quota-related event The tenant is at its device quota for this resourceName. Raise it with the platform team, or free one up elsewhere in the tenant.
Machine schedules, but on the wrong hardware More than one node carries a device under that same name. Ask whether the pool needs pinning to a specific node.
VM never reaches Running The device itself may be unhealthy on that node. This is a platform-side issue to report, not something visible from inside the cluster.

See Also: Deploying Applications ยท GPU & Device Passthrough for how the platform whitelists devices in the first place