VM Images¶
Tenant worker nodes are KubeVirt VMs, so every tenant cluster needs a worker disk image. kMetal supports two boot modes and a golden-image pipeline for producing them.
Boot modes¶
ContainerDisk — ephemeral, fast first boot¶
A ContainerDisk is a KubeVirt boot mode where the OS image is shipped as a container image. The VM mounts the image from the container registry at boot time and runs against a writable overlay; the underlying image is read-only and shared across worker VMs.
- First-boot time: ~50 seconds.
- Persistence: ephemeral. The overlay is discarded when the VM is destroyed.
- Best for: production worker nodes that are routinely rebuilt, where the cluster expects worker churn (e.g., the autoscaler creates and tears down workers). No state survives a VM rebuild, which matches what kubeadm-managed workers want anyway.
- Used by: the
kubevirt-kubeadmClusterClass tier (Ubuntu ContainerDisk + kubeadm bootstrap, the default production tier).
DataVolume — persistent, snapshot-cloned¶
A DataVolume is a KubeVirt boot mode where the disk is a PersistentVolume cloned from a snapshot of a golden image. The VM writes to its own copy; the snapshot is unchanged.
- First-boot time: ~3 minutes from
Machinecreated toNodeReady, measured on Ceph RBD. Only ~9 seconds of that is the clone — the rest is the VM booting and joining. Subsequent restarts are fast (the disk already exists). - Persistence: yes. The VM's disk survives shutdown / restart; deleting the VM deletes the disk.
- Best for: tenant workloads that benefit from a fast image with the entire kubelet, container runtime, and add-ons pre-installed in the disk (Kairos golden image). Long-running worker pools where a 3-minute one-time cost amortizes well.
- Used by: the
standardandexpertClusterClass tiers (Kairos), which are the ones enabled by default.
The 9 seconds depend on a CSI clone
That number is the backend snapshotting internally. If CDI cannot take that path it falls back to a host-assisted copy — a pod-to-pod transfer of the whole boot disk, per machine — and first boot goes from minutes to tens of minutes.
The condition is that the golden volume and the machine's root disk sit on the same StorageClass, which is why kMetal computes both from storage.underclusterClassName and does not offer them as separate knobs.
Check which path a machine took rather than assuming.
The boot mode is part of the ClusterClass; tenants pick a tier, the tier picks the mode. Operators don't configure boot mode per cluster — they pick the ClusterClass tier per cluster.
Golden-image pipeline¶
The persistent DataVolume mode needs an image to clone from.
kMetal builds those images in-cluster via the image-builder Helm chart, in kmetal-image-builder on the under cluster.
The whole flow runs inside the under cluster — no external image-builder service, no manual upload, no out-of-band image registry. Operators can rebuild the golden image at will (kernel CVE, Kubernetes version bump, add-on update), and the next time a tenant cluster's MachineDeployment scales up, the new workers come up on the new image.
For what to put in an image and the names it has to land on, see Nodes Management. This section is what the pipeline does once you apply one, and what it costs.
What actually happens, with timings¶
Every object is named after the artifact, which makes the whole run greppable by one name.
The run below is a real one: ubuntu-2404-30gi-v1358-rc0 — Ubuntu 24.04 on Kairos, the 30Gi boot tier, Kubernetes v1.35.8 — on Ceph RBD.
Three artifacts built in parallel; the timings are per artifact.
| Step | What appears | Elapsed |
|---|---|---|
1. The OSArtifact is applied by the ClusterClass chart |
the OSArtifact, and PVC <artifact>-artifacts (10Gi) |
— |
| 2. Kairos builds the OS image and a bootable ISO into that PVC | a build Pod, then status.phase: Ready on the artifact |
~11 min |
| 3. A Job boots that ISO in a VM, which installs Kairos onto a fresh disk | Job <artifact>, VM <artifact>, PVC <artifact>-rootdisk (30Gi) |
~7 min, VM Succeeded when done |
| 4. The installed disk is snapshotted | VolumeSnapshot <artifact>-golden |
~1 s |
| 5. The snapshot is restored into the volume machines clone from | DataVolume + PVC <artifact>-golden (30Gi) |
seconds |
Roughly 18 minutes per artifact, and it is the two build steps that cost. Steps 4 and 5 are metadata operations on the backend — which is also the reason a machine creation later is seconds rather than minutes.
$ kubectl get osartifacts -n kmetal-image-builder
NAME PHASE AGE
ubuntu-2404-30gi-v13411-rc0 Ready 32m
ubuntu-2404-30gi-v1358-rc0 Ready 32m
ubuntu-2404-30gi-v1364-rc0 Ready 32m
$ kubectl get pvc,volumesnapshot -n kmetal-image-builder -o name | grep v1358
persistentvolumeclaim/ubuntu-2404-30gi-v1358-rc0-artifacts
persistentvolumeclaim/ubuntu-2404-30gi-v1358-rc0-golden
persistentvolumeclaim/ubuntu-2404-30gi-v1358-rc0-rootdisk
volumesnapshot/ubuntu-2404-30gi-v1358-rc0-golden
Ready on the OSArtifact means step 2 finished, not that the golden volume exists.
The thing a machine actually clones is the -golden PVC, so that is what to wait for before telling tenants a version is available.
Budget 70 GiB per artifact, not 30¶
The two intermediate volumes are not cleaned up when the build finishes.
Long after the run above completed, each artifact still holds three PVCs — -artifacts (10Gi), -rootdisk (30Gi) and -golden (30Gi) — and its installer VM is still there in Stopped.
So the cost of offering n Kubernetes versions at m boot sizes is n × m artifacts at (tier size × 2 + 10 GiB) each, not one golden volume each.
Three versions at one 30Gi tier is 210 GiB on underclusterClassName before a single tenant machine exists.
That is worth knowing in both directions: it is capacity to plan for, and it is also the first place to look for space to reclaim.
The -rootdisk PVC and the installer VM are dead once -golden is ready — but they are part of what the chart rendered, so deleting them by hand puts you outside what the release will reconcile back. Plan the space instead.
Check which clone path a machine took¶
CDI records the decision on the machine's own DataVolume, in the tenant's namespace:
$ kubectl get dv -n dynamo-dev fission-md-0-qpljw-swcn7-l4tn6-rootdisk \
-o jsonpath='{.metadata.annotations.cdi\.kubevirt\.io/cloneType}{"\n"}'
csi-clone
csi-clone is the good state, and the whole reason machine creation is fast.
host-assisted means CDI could not use the backend: the source and target classes differ, or the class has no usable snapshot capability.
copy is the same story with no snapshot involved at all.
The source is a golden PVC in another namespace, which is normal here:
$ kubectl get dv -n dynamo-dev fission-md-0-qpljw-swcn7-l4tn6-rootdisk \
-o jsonpath='{.spec.source.pvc}{"\n"}'
{"name":"ubuntu-2404-30gi-v1358-rc0-golden","namespace":"kmetal-image-builder"}
Cross-namespace cloning needs a RoleBinding the image builder maintains for you — see Why cross-namespace cloning works when a clone fails forbidden rather than stalling.
What CDI believes about the class itself is on the StorageProfile.
Ask it about whatever storage.underclusterClassName names — on the cluster below that is still tenant-storage-class, shared with tenant workloads, which is the arrangement the quota note argues against:
$ kubectl get storageprofile tenant-storage-class \
-o jsonpath='{.status.cloneStrategy}{"\t"}{.status.snapshotClass}{"\n"}'
csi-clone tenant-storage-class-snapshot
An empty snapshotClass there is the root cause of most stalled builds: no snapshot class the driver serves, so no golden volume, so nothing to clone.
Caveats when building a custom OS¶
These are the ones that bite in practice, in the order they bite:
- The class must snapshot, and it is the platform class. The golden image and every machine root disk are provisioned on
storage.underclusterClassName. A node-local provisioner cannot snapshot, so an artifact on one stalls with no-goldenvolume and every machine waits on a clone source that will never exist. UseetcdClassNameif what you wanted node-local was etcd. dataVolume.volumeSnapshotClasshas no chart default. From ClusterClass chart v1.10.0 the chart refuses to render without it, so this is a first-install decision rather than something to fix after a failed build. It must be served by the same driver asunderclusterClassName— same driver, not necessarily the same name.- Step 3 needs a compute node. The installer is a real VM, so it only runs where
virt-handlerruns, which kMetal pins to nodes labellednode-role.kubernetes.io/compute. With no such node the artifact reachesReadyand then nothing further happens — no Job pod, no rootdisk PVC, no error on the artifact. - The build pod runs
buildahwith AppArmor unconfined. TheOSArtifactcarriescontainer.apparmor.security.beta.kubernetes.io/buildah-build: unconfinedinspec.podAnnotations, because building an OS image needs it. A Pod Security or AppArmor policy onkmetal-image-builderthat forbids that annotation rejects the build pod at step 2 — and the failure is the pod's, not the artifact's, so check the pod rather than the CR. - Artifacts survive the chart. Each carries
helm.sh/resource-policy: keep, so a chart upgrade or an uninstall leaves them — deliberately, because a machine that reboots still needs its clone source. The corollary is that old revisions accumulate until someone prunes them, and an artifact in use must never be rebuilt in place. - The CDI pins do not cover the platform class.
tenantClaimPropertySetsand the filesystem-overhead pin both nametenantClassName, so golden and machine volumes take CDI's derived capabilities and its default 6% overhead — see the warning in Storage Configuration. - Machine disks land in the tenant's namespace. They are the platform's volumes on the platform's class, but the namespace is the tenant's, so a
ResourceQuotaon that class counts them against the tenant — see Keep the platform class and the tenant class apart.
Picking a boot mode¶
| Concern | ContainerDisk | DataVolume |
|---|---|---|
| First-boot time | ~50 s | ~3 min, of which ~9 s is the clone |
| Restart time (existing VM) | ~50 s (re-clone from image) | seconds |
| Persistence | Ephemeral | Persistent |
| Storage cost | Negligible (shared read-only) | One PVC per VM, plus ~70 GiB per artifact on the platform class |
| Build cost | None — the image is pulled | ~18 min per artifact, in-cluster |
| Best fit | High-churn worker pools | Long-lived workers, baked-in add-ons |
| ClusterClass tier | kubevirt-kubeadm (off by default) |
kubevirt-standard and kubevirt-expert (Kairos, on by default) |
When in doubt, start with kubevirt-kubeadm (ContainerDisk): nothing has to be built, and the only thing to keep in step is the image tag.
The Kairos tiers are the right pick when you need a more comprehensive image (e.g., add-ons pre-installed for an air-gapped tenant, or a specific kernel build) — at the cost of the pipeline above, and of a platform storage class that can snapshot.