Skip to content

VM Images

Tenant worker nodes are KubeVirt VMs, so every tenant cluster needs a worker disk image. kMetal supports two boot modes and a golden-image pipeline for producing them.

Boot modes

ContainerDisk — ephemeral, fast first boot

A ContainerDisk is a KubeVirt boot mode where the OS image is shipped as a container image. The VM mounts the image from the container registry at boot time and runs against a writable overlay; the underlying image is read-only and shared across worker VMs.

  • First-boot time: ~50 seconds.
  • Persistence: ephemeral. The overlay is discarded when the VM is destroyed.
  • Best for: production worker nodes that are routinely rebuilt, where the cluster expects worker churn (e.g., the autoscaler creates and tears down workers). No state survives a VM rebuild, which matches what kubeadm-managed workers want anyway.
  • Used by: the kubevirt-kubeadm ClusterClass tier (Ubuntu ContainerDisk + kubeadm bootstrap, the default production tier).

DataVolume — persistent, snapshot-cloned

A DataVolume is a KubeVirt boot mode where the disk is a PersistentVolume cloned from a snapshot of a golden image. The VM writes to its own copy; the snapshot is unchanged.

  • First-boot time: ~3 minutes from Machine created to NodeReady, measured on Ceph RBD. Only ~9 seconds of that is the clone — the rest is the VM booting and joining. Subsequent restarts are fast (the disk already exists).
  • Persistence: yes. The VM's disk survives shutdown / restart; deleting the VM deletes the disk.
  • Best for: tenant workloads that benefit from a fast image with the entire kubelet, container runtime, and add-ons pre-installed in the disk (Kairos golden image). Long-running worker pools where a 3-minute one-time cost amortizes well.
  • Used by: the standard and expert ClusterClass tiers (Kairos), which are the ones enabled by default.

The 9 seconds depend on a CSI clone

That number is the backend snapshotting internally. If CDI cannot take that path it falls back to a host-assisted copy — a pod-to-pod transfer of the whole boot disk, per machine — and first boot goes from minutes to tens of minutes.

The condition is that the golden volume and the machine's root disk sit on the same StorageClass, which is why kMetal computes both from storage.underclusterClassName and does not offer them as separate knobs. Check which path a machine took rather than assuming.

The boot mode is part of the ClusterClass; tenants pick a tier, the tier picks the mode. Operators don't configure boot mode per cluster — they pick the ClusterClass tier per cluster.

Golden-image pipeline

The persistent DataVolume mode needs an image to clone from. kMetal builds those images in-cluster via the image-builder Helm chart, in kmetal-image-builder on the under cluster.

The whole flow runs inside the under cluster — no external image-builder service, no manual upload, no out-of-band image registry. Operators can rebuild the golden image at will (kernel CVE, Kubernetes version bump, add-on update), and the next time a tenant cluster's MachineDeployment scales up, the new workers come up on the new image.

For what to put in an image and the names it has to land on, see Nodes Management. This section is what the pipeline does once you apply one, and what it costs.

What actually happens, with timings

Every object is named after the artifact, which makes the whole run greppable by one name. The run below is a real one: ubuntu-2404-30gi-v1358-rc0 — Ubuntu 24.04 on Kairos, the 30Gi boot tier, Kubernetes v1.35.8 — on Ceph RBD. Three artifacts built in parallel; the timings are per artifact.

Step What appears Elapsed
1. The OSArtifact is applied by the ClusterClass chart the OSArtifact, and PVC <artifact>-artifacts (10Gi) —
2. Kairos builds the OS image and a bootable ISO into that PVC a build Pod, then status.phase: Ready on the artifact ~11 min
3. A Job boots that ISO in a VM, which installs Kairos onto a fresh disk Job <artifact>, VM <artifact>, PVC <artifact>-rootdisk (30Gi) ~7 min, VM Succeeded when done
4. The installed disk is snapshotted VolumeSnapshot <artifact>-golden ~1 s
5. The snapshot is restored into the volume machines clone from DataVolume + PVC <artifact>-golden (30Gi) seconds

Roughly 18 minutes per artifact, and it is the two build steps that cost. Steps 4 and 5 are metadata operations on the backend — which is also the reason a machine creation later is seconds rather than minutes.

$ kubectl get osartifacts -n kmetal-image-builder
NAME                          PHASE   AGE
ubuntu-2404-30gi-v13411-rc0   Ready   32m
ubuntu-2404-30gi-v1358-rc0    Ready   32m
ubuntu-2404-30gi-v1364-rc0    Ready   32m

$ kubectl get pvc,volumesnapshot -n kmetal-image-builder -o name | grep v1358
persistentvolumeclaim/ubuntu-2404-30gi-v1358-rc0-artifacts
persistentvolumeclaim/ubuntu-2404-30gi-v1358-rc0-golden
persistentvolumeclaim/ubuntu-2404-30gi-v1358-rc0-rootdisk
volumesnapshot/ubuntu-2404-30gi-v1358-rc0-golden

Ready on the OSArtifact means step 2 finished, not that the golden volume exists. The thing a machine actually clones is the -golden PVC, so that is what to wait for before telling tenants a version is available.

Budget 70 GiB per artifact, not 30

The two intermediate volumes are not cleaned up when the build finishes. Long after the run above completed, each artifact still holds three PVCs — -artifacts (10Gi), -rootdisk (30Gi) and -golden (30Gi) — and its installer VM is still there in Stopped.

So the cost of offering n Kubernetes versions at m boot sizes is n × m artifacts at (tier size × 2 + 10 GiB) each, not one golden volume each. Three versions at one 30Gi tier is 210 GiB on underclusterClassName before a single tenant machine exists.

That is worth knowing in both directions: it is capacity to plan for, and it is also the first place to look for space to reclaim. The -rootdisk PVC and the installer VM are dead once -golden is ready — but they are part of what the chart rendered, so deleting them by hand puts you outside what the release will reconcile back. Plan the space instead.

Check which clone path a machine took

CDI records the decision on the machine's own DataVolume, in the tenant's namespace:

$ kubectl get dv -n dynamo-dev fission-md-0-qpljw-swcn7-l4tn6-rootdisk \
    -o jsonpath='{.metadata.annotations.cdi\.kubevirt\.io/cloneType}{"\n"}'
csi-clone

csi-clone is the good state, and the whole reason machine creation is fast. host-assisted means CDI could not use the backend: the source and target classes differ, or the class has no usable snapshot capability. copy is the same story with no snapshot involved at all.

The source is a golden PVC in another namespace, which is normal here:

$ kubectl get dv -n dynamo-dev fission-md-0-qpljw-swcn7-l4tn6-rootdisk \
    -o jsonpath='{.spec.source.pvc}{"\n"}'
{"name":"ubuntu-2404-30gi-v1358-rc0-golden","namespace":"kmetal-image-builder"}

Cross-namespace cloning needs a RoleBinding the image builder maintains for you — see Why cross-namespace cloning works when a clone fails forbidden rather than stalling.

What CDI believes about the class itself is on the StorageProfile. Ask it about whatever storage.underclusterClassName names — on the cluster below that is still tenant-storage-class, shared with tenant workloads, which is the arrangement the quota note argues against:

$ kubectl get storageprofile tenant-storage-class \
    -o jsonpath='{.status.cloneStrategy}{"\t"}{.status.snapshotClass}{"\n"}'
csi-clone   tenant-storage-class-snapshot

An empty snapshotClass there is the root cause of most stalled builds: no snapshot class the driver serves, so no golden volume, so nothing to clone.

Caveats when building a custom OS

These are the ones that bite in practice, in the order they bite:

  1. The class must snapshot, and it is the platform class. The golden image and every machine root disk are provisioned on storage.underclusterClassName. A node-local provisioner cannot snapshot, so an artifact on one stalls with no -golden volume and every machine waits on a clone source that will never exist. Use etcdClassName if what you wanted node-local was etcd.
  2. dataVolume.volumeSnapshotClass has no chart default. From ClusterClass chart v1.10.0 the chart refuses to render without it, so this is a first-install decision rather than something to fix after a failed build. It must be served by the same driver as underclusterClassName — same driver, not necessarily the same name.
  3. Step 3 needs a compute node. The installer is a real VM, so it only runs where virt-handler runs, which kMetal pins to nodes labelled node-role.kubernetes.io/compute. With no such node the artifact reaches Ready and then nothing further happens — no Job pod, no rootdisk PVC, no error on the artifact.
  4. The build pod runs buildah with AppArmor unconfined. The OSArtifact carries container.apparmor.security.beta.kubernetes.io/buildah-build: unconfined in spec.podAnnotations, because building an OS image needs it. A Pod Security or AppArmor policy on kmetal-image-builder that forbids that annotation rejects the build pod at step 2 — and the failure is the pod's, not the artifact's, so check the pod rather than the CR.
  5. Artifacts survive the chart. Each carries helm.sh/resource-policy: keep, so a chart upgrade or an uninstall leaves them — deliberately, because a machine that reboots still needs its clone source. The corollary is that old revisions accumulate until someone prunes them, and an artifact in use must never be rebuilt in place.
  6. The CDI pins do not cover the platform class. tenantClaimPropertySets and the filesystem-overhead pin both name tenantClassName, so golden and machine volumes take CDI's derived capabilities and its default 6% overhead — see the warning in Storage Configuration.
  7. Machine disks land in the tenant's namespace. They are the platform's volumes on the platform's class, but the namespace is the tenant's, so a ResourceQuota on that class counts them against the tenant — see Keep the platform class and the tenant class apart.

Picking a boot mode

Concern ContainerDisk DataVolume
First-boot time ~50 s ~3 min, of which ~9 s is the clone
Restart time (existing VM) ~50 s (re-clone from image) seconds
Persistence Ephemeral Persistent
Storage cost Negligible (shared read-only) One PVC per VM, plus ~70 GiB per artifact on the platform class
Build cost None — the image is pulled ~18 min per artifact, in-cluster
Best fit High-churn worker pools Long-lived workers, baked-in add-ons
ClusterClass tier kubevirt-kubeadm (off by default) kubevirt-standard and kubevirt-expert (Kairos, on by default)

When in doubt, start with kubevirt-kubeadm (ContainerDisk): nothing has to be built, and the only thing to keep in step is the image tag. The Kairos tiers are the right pick when you need a more comprehensive image (e.g., add-ons pre-installed for an air-gapped tenant, or a specific kernel build) — at the cost of the pipeline above, and of a platform storage class that can snapshot.