GPU & Device Passthrough¶
Tenant VMs are KubeVirt VMs, so they can request a real PCI device, or a mediated device (mdev/vGPU), the same way any KubeVirt VM does. kMetal's part is exposing that whitelist on the KMetal CR and threading a matching deviceName through the ClusterClass, instead of every operator hand-patching the KubeVirt CR themselves.
This page covers enabling and whitelisting devices on the under cluster. For a tenant requesting one on their own cluster, see Requesting a GPU on a Tenant Cluster.
KubeVirt's own reference docs are the source of truth for the underlying mechanism:
Node prerequisites¶
Passthrough is a bare-metal feature before it is a Kubernetes one. Every compute node carrying a device you intend to pass through needs:
- IOMMU enabled in firmware (Intel VT-d / AMD-Vi).
- A kernel boot parameter turning it on:
intel_iommu=onoramd_iommu=on. - The device unbound from its default driver and bound to
vfio-pci.
Whitelisting devices¶
Both kinds of device are whitelisted the same way: set spec.kubevirt on the KMetal CR. kMetal's operator renders whatever is set there onto the KubeVirt CR's own permittedHostDevices allowlist. Nothing here needs applying by hand.
PCI passthrough¶
A real, physical PCI device (a GPU, an NVMe controller, anything with a vendor:device ID) is whitelisted under permittedHostDevices:
apiVersion: clastix.io/v1
kind: KMetal
metadata:
name: dev-env
spec:
kubevirt:
permittedHostDevices:
- pciVendorSelector: "10de:2236"
resourceName: "nvidia.com/A10"
pciVendorSelector: the device'svendor:devicehex pair. Read it off the node withlspci -nn.resourceName: what the tenant-facing ClusterClassgpusvariable'sdeviceNamehas to match.
mdev passthrough¶
A mediated device (a vGPU instance carved out of a physical GPU by its vendor driver) is whitelisted under mediatedDevices:
apiVersion: clastix.io/v1
kind: KMetal
metadata:
name: dev-env
spec:
kubevirt:
mediatedDevices:
- mdevNameSelector: "nvidia-222"
resourceName: "nvidia.com/GRID_T4-2Q"
mdevNameSelector: the mdev type's exact name, as reported undermdev_supported_types/*/nameon the node, not necessarily the same string as the type's sysfs directory name.resourceName: what the tenant-facing ClusterClassgpusvariable'sdeviceNamehas to match.
See Platform Configuration Reference: spec.kubevirt for the full field list.
Capping tenant consumption¶
Whitelisting a device only decides that it can be requested at all; it does not bound how many a given tenant may hold at once. A device is scarce and physically tied to a node, so cap it the same way cluster count or external addresses are capped: with a resourceQuotas entry on the tenant's Tenant object, targeting requests.<resourceName> for the exact resourceName set above.
See Cap GPUs and other devices for the quota YAML and how it aggregates across a tenant's Environments.