Nodes Management Tasks¶
Building the disks tenant worker nodes boot from.
Tenant worker nodes are KubeVirt virtual machines, and where their root disk comes from depends on the tier:
| Tier | Root disk |
|---|---|
standard, expert |
A golden image built on the platform: an OSArtifact goes through the image builder and comes out as either a DataVolume to clone or a containerDisk to pull. |
kubeadm |
A prebuilt containerDisk pulled from a registry, named by kubeadm.containerDisk.image and tag. No OSArtifact involved. |
So this page is about the Kairos tiers.
If you only offer kubeadm, the whole subject reduces to keeping that image's tag in step with the Kubernetes version.
For the boot modes themselves, see VM Images.
What the chart builds for you¶
Enabling standard or expert makes the ClusterClass chart render the build inputs, from two values it multiplies together:
- one Secret per entry in
osArtifact.tags, holding thecloud-config.yamland theDockerfilethe image is built from, - one
OSArtifactper (tag × entry indataVolume.storageTiers), annotated for the image builder,
all in osArtifact.imageBuilderNamespace.
Three Kubernetes versions and two boot volume sizes is six builds, so the multiplication is worth doing before you add either.
spec:
clusterClass:
values:
osArtifact:
name: ubuntu-2404 # the family name every artifact is prefixed with
latestVersion: rc0 # your OS revision, independent of Kubernetes
tags: # Kubernetes versions to build for
- v1.35.4
- v1.34.7
dataVolume:
storageTiers:
- 30Gi
strategy: dataVolume # dataVolume | oci
| Value | Notes |
|---|---|
osArtifact.name |
Prefix of every generated object. Changing it orphans the artifacts already built under the old name. |
osArtifact.latestVersion |
Your own OS revision — CVE rebuilds, tool updates, a changed Dockerfile. It is not the Kubernetes version, and it is what you bump when the content changes but the version list does not. |
osArtifact.tags |
The Kubernetes versions worker images exist for. A tenant asking for a version not in this list gets workers with no disk to boot. |
osArtifact.storageClassName |
Computed by kMetal from spec.storage.underclusterClassName — the class the golden image is built on. |
dataVolume.storageClassName |
Computed by kMetal from spec.storage.underclusterClassName too, and it is the default of the machineStorage.storageClass variable — so the machine root disk lands on the same class as the golden image it is cloned from. That is what keeps machine creation a CSI clone rather than a full copy; see Check which clone path a machine took. |
osArtifact.imageBuilderNamespace |
Computed by kMetal as kmetal-image-builder, where the image builder actually runs. The chart's own default points elsewhere; do not set it. |
dataVolume.volumeSnapshotClass |
The class the golden snapshot is taken with. Must be served by the same driver as underclusterClassName, or the snapshot never becomes ready. The chart ships no default for it from v1.10.0 and refuses to render without one. |
strategy |
dataVolume clones a PVC per machine; oci pushes a containerDisk to oci.registry and machines pull it. Chart-wide, and it decides which of the two names below the machine template asks for. |
The naming contract¶
This is the part that decides whether a machine boots, and it is worth reading before building anything by hand.
Nothing searches for a matching artifact. The worker machine template computes a name from four inputs and asks for exactly that:
strategy |
What the machine asks for |
|---|---|
dataVolume |
PVC <name>-<size>-<version>-<latestVersion>-golden in <imageBuilderNamespace>, cloned by CDI |
oci |
Image <oci.registry>/<name>-<size>-<version>-<latestVersion>:latest |
where <size> is the boot volume size lowercased and <version> is the cluster's Kubernetes version with the dots removed.
For osArtifact.name: ubuntu-2404, a 30Gi tier, a cluster on v1.34.7 and latestVersion: rc0:
| Object | Name |
|---|---|
| Secret (cloud-config + Dockerfile) | ubuntu-2404-v1347-rc0 |
OSArtifact |
ubuntu-2404-30gi-v1347-rc0 |
Golden VolumeSnapshot and DataVolume |
ubuntu-2404-30gi-v1347-rc0-golden |
| What a worker clones | that -golden PVC, in kmetal-image-builder |
An artifact you build yourself has to land on that name.
A machine that waits forever on a volume is almost always a name that does not match — a version whose dots were kept, a latestVersion bumped on one side only, a size that is not one of the tiers.
Build your own artifact¶
The chart's templates/osartifact.yaml is written to be copied: it is a reference, not a fixed pipeline.
Copy it, override the version inputs, and apply the result yourself when you need something the chart does not render — a different base OS, extra packages, a preloaded add-on, an air-gapped image.
Two objects per artifact. First a Secret carrying the build:
apiVersion: v1
kind: Secret
metadata:
name: ubuntu-2404-v1347-rc0
namespace: kmetal-image-builder
type: Opaque
stringData:
cloud-config.yaml: |
#cloud-config
install:
auto: true
device: "auto"
Dockerfile: |
FROM quay.io/kairos/kairos-init:v0.7.0 AS kairos-init
FROM ubuntu:24.04
# ... your build ...
COPY --from=kairos-init /kairos-init /kairos-init
RUN /kairos-init -l debug -s install -m generic --version v3.6.0 \
&& /kairos-init -l debug -s init -m generic --version v3.6.0 \
&& rm -f /kairos-init
Then the OSArtifact pointing at it, once per boot volume size:
apiVersion: build.kairos.io/v1alpha2
kind: OSArtifact
metadata:
name: ubuntu-2404-30gi-v1347-rc0
namespace: kmetal-image-builder
annotations:
image-builder.clastix.io/requests.storage: 30Gi
image-builder.clastix.io/storageClassName: platform-storage # underclusterClassName
image-builder.clastix.io/volumeSnapshotClassName: platform-storage-snapshot
spec:
artifacts:
arch: amd64
cloudConfigRef:
key: cloud-config.yaml
name: ubuntu-2404-v1347-rc0
grubConfig: |
set default=0
set timeout=0
set timeout_style=hidden
iso: true
image:
ociSpec:
ref:
key: Dockerfile
name: ubuntu-2404-v1347-rc0
push: false
| Annotation | Required | Sets |
|---|---|---|
image-builder.clastix.io/requests.storage |
yes | The resulting disk size. This is the <size> in the name, and it must equal one of your storageTiers. |
image-builder.clastix.io/storageClassName |
yes | The class the image is built on. Match spec.storage.underclusterClassName: a machine cloning from a golden volume on some other class gets a host-assisted copy of the whole disk instead of a snapshot. |
image-builder.clastix.io/volumeSnapshotClassName |
yes in practice | The class the golden snapshot is taken with. |
image-builder.clastix.io/strategy |
no | oci to push a containerDisk instead of producing a DataVolume. |
image-builder.clastix.io/registry |
with oci |
Registry path; the push lands at <registry>/<artifact name>:latest. |
image-builder.clastix.io/registrySecret |
no | A dockerconfigjson Secret for that registry. |
image-builder.clastix.io/googleApplicationCredentials |
no | A Secret holding a config.json for Google Artifact Registry, used instead of the above. |
Without storageClassName or requests.storage the builder logs MissingAnnotations and stops, rather than failing the object — so an artifact that never starts building is worth a controller-log check.
What the Dockerfile has to end up with¶
The build is an ordinary container image that kairos-init then converts into a bootable Kairos system, so anything a worker node needs has to be baked in:
kubelet,kubeadmandkubectlpinned to the Kubernetes version the artifact is named for, and held so an unattended upgrade cannot move them,- a container runtime —
containerdandrunc, withSystemdCgroup = truein/etc/containerd/config.toml, socat,conntrackandqemu-guest-agent,- the Kairos kubeadm provider under
/system/providers, with its scripts in/opt/kubeadm/scripts, - optionally the control-plane and CNI images pre-pulled into
/opt/content/images, numbered in load order — this is what makes an air-gapped or slow-registry first boot survivable, kairos-init -s installthen-s init.
The chart's rendered Dockerfile does all of this and is the honest starting point; diff yours against it after a chart bump.
What happens after you apply it¶
The image builder runs four steps, and each leaves an object named after the artifact:
| Step | Object | Typical |
|---|---|---|
Kairos builds the artifacts and reports status.phase: Ready |
the OSArtifact itself, PVC <artifact>-artifacts |
~11 min |
| A Job converts them into a bootable disk | Job <artifact> |
seconds |
| A VirtualMachine boots that disk once and installs the OS | VM <artifact>, PVC <artifact>-rootdisk |
~7 min |
| The installed disk is snapshotted and turned into the golden volume | VolumeSnapshot and DataVolume <artifact>-golden |
seconds |
On the oci strategy the last step is replaced by a Kaniko Job, <artifact>-oci-push, which wraps the installed disk as a FROM scratch containerDisk and pushes it.
Step three is a real VM, so it runs only where virt-handler runs — on nodes labelled node-role.kubernetes.io/compute.
On a cluster with no such node the artifact reaches Ready and the pipeline simply stops there, with nothing on the artifact to say why.
Ready is therefore not the signal to act on.
The object a machine clones is the -golden PVC, and that is what to wait for.
For the timings, the disk this holds and the rest of the caveats, see VM Images.
kubectl get osartifacts -n kmetal-image-builder
kubectl get jobs,vm,volumesnapshots,datavolumes -n kmetal-image-builder
kubectl describe osartifact <name> -n kmetal-image-builder
Never rebuild an artifact in use¶
Artifacts are immutable in practice, whatever the API allows.
Deleting or replacing an OSArtifact takes its golden volume with it, and a machine created from that volume cannot start again once virt-handler restarts — the clone source is gone.
Which means an artifact stays for as long as any machine built from it might reboot.
So a rebuild is always an addition:
- bump
osArtifact.latestVersion(andtags, if the version list changed), - let the new artifacts build to
-golden, - move clusters onto them, which rolls their workers,
- remove the old artifacts only once nothing references them.
Keeping two revisions costs one golden volume per size per version, which is the price of being able to roll back.
Publish a new Kubernetes version¶
Both tiers need the same three moves, in this order:
- Add the version to
osArtifact.tags— or, for thekubeadmtier, pointkubeadm.containerDisk.tagat an image for it. - Wait for the artifacts to reach
-golden. Nothing validates this for a tenant; a cluster created early gets a control plane that comes up and workers that never join. - Tell tenants the version is available.
Add a boot volume size¶
A size is a tier in dataVolume.storageTiers and an artifact built at that size, for every Kubernetes version you offer.
Add the tier alone and the tenants who pick it wait on a clone source that does not exist.
See Set the boot volume sizes and snapshot class.
Why cross-namespace cloning works¶
Golden volumes live in kmetal-image-builder; the machines that clone them live in tenant namespaces, and CDI refuses a cross-namespace clone unless the consuming namespace's ServiceAccount holds datavolume-clone-source where the source PVC is.
The image builder maintains that itself: a controller watches every CAPI Cluster and keeps a single kmetal-clone-source RoleBinding in its own namespace, whose subjects are the default ServiceAccount of each namespace holding a live Cluster.
Namespaces whose clusters are gone are pruned.
Nothing to configure, but it is where to look when clones fail with a forbidden error rather than a missing source:
Diagnose a machine that will not boot¶
| Symptom | Usually |
|---|---|
The DataVolume on the machine stays pending, no clone starts |
The -golden volume for that name does not exist — see the naming contract. |
| Clone fails with a forbidden error from CDI | The clone-source RoleBinding does not list that namespace yet, which normally means the Cluster was only just created. |
The OSArtifact never leaves its initial phase |
Missing storageClassName or requests.storage, or Kairos itself failed the build — check the artifact's own status first, then the builder's logs. |
The golden VolumeSnapshot never becomes ready |
The snapshot class is not served by the same driver as the storage class, or the class cannot snapshot at all — kubectl get storageprofile <class> -o jsonpath='{.status.snapshotClass}' is empty. |
The artifact is Ready but no Job or installer VM ever appears |
No node is labelled node-role.kubernetes.io/compute, so there is nowhere for the installer VM to run. |
| Machines boot, but each one takes tens of minutes | CDI fell back to a host-assisted copy: the golden volume and the machine root disk are not on the same class. Check cdi.kubevirt.io/cloneType on the machine's DataVolume. |
| Workers boot but never join | The image's kubelet/kubeadm do not match the cluster's Kubernetes version — an artifact built for one version and named for another. |