MultiCluster Management Tasks¶
Deciding which cluster shapes, machine sizes and Kubernetes versions tenants may ask for.
All of it is spec.clusterClass.values, a passthrough merged over what kMetal computes — and winning on conflict.
See Component Configuration for what that passthrough is and is not.
Read the surface before changing it¶
The ClusterClass chart carries far more knobs than the KMetal API models, and they move between chart versions.
Read the accepted values off the pinned chart rather than guessing:
The pinned version is the clusterclass-kmetal component's version:
kubectl get km kmetal -o jsonpath='{range .status.components[*]}{.name}{"\t"}{.installedVersion}{"\n"}{end}' | grep clusterclass
Choose which tiers tenants can use¶
Tiers are the cluster shapes on offer. Each is enabled independently:
spec:
clusterClass:
values:
tiers:
kubeadm:
enabled: true
standard:
enabled: false
expert:
enabled: false
| Tier | ClusterClass | Bootstrap | Enabled by default | Surface |
|---|---|---|---|---|
standard |
kubevirt-standard |
Kairos | yes | The narrow set: datastore, control-plane endpoint, machine size, boot volume. |
expert |
kubevirt-expert |
Kairos | yes | Everything standard has, plus control-plane pod resources, a registry override, custom machine specs and worker scheduling. |
kubeadm |
kubevirt-kubeadm |
kubeadm, from a containerDisk image | no | Like expert, minus boot-volume selection, plus extra Linux users on workers. |
kubeadm is off by default for two reasons: it needs the CAPI kubeadm bootstrap provider on the management cluster, and its containerDisk image forces BIOS firmware on the worker VMs.
Disabling a tier removes it from what tenants can reference. Existing clusters on a disabled tier keep running, but the ClusterClass they depend on may stop being reconciled — treat disabling a tier in use as a migration, not a toggle.
What a tenant writes¶
A Cluster names a tier and fills in variables; everything else is the ClusterClass's.
This is a working one on the kubeadm tier — the shape a tenant actually applies:
apiVersion: cluster.x-k8s.io/v1beta2
kind: Cluster
metadata:
name: solar
namespace: photon-prod
spec:
topology:
classRef:
name: kubevirt-kubeadm # kubevirt-<tier>
namespace: kmetal-capi-providers # where the ClusterClasses are installed
version: v1.34.1 # a version the worker image exists for
variables:
- name: network
value:
subnet: production # the tenant's SubnetClaim, by claim name
- name: controlPlane
value:
dataStoreName: default
network:
serviceType: LoadBalancer
serviceAddress: 172.21.208.100 # from tenant-cp-pool
certSANs:
- 198.51.100.10 # the public address, if there is one
- name: machineSize
value: small
workers:
machineDeployments:
- class: default-worker
name: md-0
replicas: 2
Four things have to line up outside the manifest, and none of them is validated for the tenant:
- the tier is enabled, or
classRefnames a ClusterClass that does not exist, - the version has a matching worker image — see below,
network.subnetnames aReadySubnetClaimin the same namespace,serviceAddressis insidetenant-cp-pool, and is not already taken.
The variables, and which tier has them¶
| Variable | standard |
expert |
kubeadm |
Sets |
|---|---|---|---|---|
network.subnet |
✓ | ✓ | ✓ | Which of the tenant's subnets the worker VMs attach to. The only variable the schema marks required. |
controlPlane.dataStoreName |
✓ | ✓ | ✓ | The Kamaji datastore holding this control plane's etcd data. |
controlPlane.network.serviceType |
✓ | ✓ | ✓ | LoadBalancer (default), NodePort or ClusterIP. |
controlPlane.network.serviceAddress |
✓ | ✓ | ✓ | The VIP the Service is published on. |
controlPlane.network.certSANs |
✓ | ✓ | ✓ | Extra names in the api-server certificate — the public address, an FQDN. |
controlPlane.apiServer / controllerManager / scheduler |
— | ✓ | ✓ | Requests and limits on the control-plane pods. |
controlPlane.registry |
— | ✓ | ✓ | A registry override for the control-plane images. |
machineSize |
✓ | ✓ | ✓ | small, medium, large, and on expert and kubeadm also custom. |
machineSpecs.cpuCores / memoryGiB |
— | ✓ | ✓ | The shape behind machineSize: custom. |
machineStorage.size / storageClass / storageAccessMode |
✓ | ✓ | — | The worker boot volume, chosen from the tiers you configured. storageClass defaults to the class kMetal computes from storage.underclusterClassName; a tenant who overrides it to anything else loses the CSI clone and gets a full image copy per machine. |
infrastructure.nodeSelector / tolerations / terminationGracePeriodSeconds |
— | ✓ | ✓ | Where the worker VMs land on the under cluster. It narrows within the compute nodes rather than escaping them: a selector resolving to a node without virt-handler — anything not labelled node-role.kubernetes.io/compute — leaves the VM unable to start, with nothing on the Cluster to say why. |
bootstrapConfig.serverAddress |
✓ | ✓ | ✓ | An in-subnet control-plane address for the workers to bootstrap against, instead of the platform-wide VIP. |
bootstrapConfig.users |
— | — | ✓ | Extra Linux users on the worker nodes, with SSH keys and sudo rules. |
Every variable is optional as far as admission is concerned, network aside — an omitted one takes the ClusterClass's default, which is why a tenant can get a cluster from a very short manifest and also why a wrong assumption about a default is not caught for them.
Example: a custom-sized cluster with its own boot volume¶
On expert, where the full surface is exposed:
variables:
- name: network
value:
subnet: production
- name: controlPlane
value:
dataStoreName: default
network:
serviceAddress: 172.21.208.111
apiServer:
resources:
requests: { cpu: "500m", memory: 1Gi }
limits: { cpu: "2", memory: 4Gi }
- name: machineSize
value: custom
- name: machineSpecs
value:
cpuCores: 12
memoryGiB: 48
- name: machineStorage
value:
size: 30Gi # one of the storageTiers you configured
storageAccessMode: ReadWriteOnce
- name: infrastructure
value:
nodeSelector:
node-role.kubernetes.io/compute: ""
- name: bootstrapConfig
value:
serverAddress: 10.100.0.254 # in-subnet, required in practice on Kairos tiers
machineStorage.size is validated against the storageTiers you set, so a tenant asking for a size you did not build a golden image for is rejected rather than left waiting on a clone that cannot happen.
Set the machine sizes tenants can pick¶
A tenant chooses a shape with the machineSize variable on their Cluster:
machineSize |
Worker shape |
|---|---|
small (default) |
2 CPU, 4 GiB RAM |
medium |
4 CPU, 8 GiB RAM |
large |
8 CPU, 16 GiB RAM |
custom |
Whatever the tenant asks for, within bounds — see below. Only on the expert and kubeadm tiers, from chart v1.7.0. |
The three fixed sizes are ClusterClass patches, so what medium means is a chart-level decision rather than a per-cluster one.
Each patch replaces the worker VM's CPU cores and guest memory on the default-worker machine deployment, and nothing else — control-plane sizing is unrelated, since a hosted control plane is pods rather than VMs.
Custom shapes¶
machineSize: custom hands the sizing to the tenant, through a second variable.
It is not offered on standard, whose enum stops at large:
spec:
topology:
variables:
- name: machineSize
value: custom
- name: machineSpecs
value:
cpuCores: 12
memoryGiB: 48
| Field | Accepted |
|---|---|
cpuCores |
1–64 |
memoryGiB |
2–128, rendered as <n>Gi of guest memory |
Both bounds come from the ClusterClass's own variable schema, so a request outside them is rejected when the Cluster is applied rather than becoming a VM that never schedules.
The two patches fire only when machineSize is custom, and they write the same two fields the fixed sizes do.
What this does not do is bound the total. Per-worker limits are not a fleet limit: a tenant may ask for the largest shape on every worker they are allowed to run, so the ceiling belongs in their quota — see Cap what a tenant can consume.
Whether to offer custom at all is a platform decision: fixed sizes make capacity planning arithmetic, and custom shapes make it a question of what tenants happen to have asked for.
Match worker images to the Kubernetes version¶
Two values must agree, and nothing checks them for you:
spec:
clusterClass:
values:
kubeadm:
containerDisk:
image: quay.io/capk/ubuntu-2404-container-disk
tag: v1.34.1 # must match spec.topology.version on tenant Clusters
A tenant asking for a version whose worker image does not exist gets a control plane that comes up and workers that never join. When you raise the version tenants may use, raise this tag in the same change.
Set the boot volume sizes and snapshot class¶
spec:
clusterClass:
values:
dataVolume:
volumeSnapshotClass: platform-storage-snapshot
storageTiers:
- 30Gi
accessMode:
- ReadWriteOnce
volumeSnapshotClass has no chart default, and the driver has to match
From chart v1.10.0 the chart ships no default and refuses to render without this value, so it is part of the first install rather than something to discover from a machine that never boots.
It must be served by the same driver as storage.underclusterClassName — the class the golden image is snapshotted on, not the tenant class.
A snapshot class on any other driver never becomes ready, and the artifact stops before it has a golden volume.
dataVolume.storageClassName is computed by kMetal from storage.underclusterClassName, and so is osArtifact.storageClassName.
Both deliberately: CDI takes the fast CSI clone path only when the clone's source and target classes match, and these two are that source and that target.
Overriding either here — the passthrough wins on conflict — turns every machine creation into a host-assisted copy of the whole boot disk.
storageTiers are the boot volume sizes machines can be created at.
Each tier is a full artifact build per Kubernetes version, holding roughly twice the tier size plus 10 GiB on the platform class — so add sizes because tenants need them, not defensively.
See Golden-image pipeline for the arithmetic.
Verify a change landed¶
# The ClusterClass component reconciled at the new values
kubectl get km <name> -o jsonpath='{range .status.components[*]}{.name}{"\t"}{.phase}{"\t"}{.message}{"\n"}{end}' | grep clusterclass
# What tenants can now reference
kubectl get clusterclass -A
What a change does to running clusters¶
Existing clusters are not left behind: Cluster API's topology controller watches the ClusterClass and re-reconciles every Cluster that references it, so an edit reaches them without anyone re-applying anything.
What that reconcile actually does depends on what you changed:
| Changed | Effect on a running cluster |
|---|---|
A patch's value — what medium means, the containerDisk tag, a boot volume default |
The worker template is rotated: Cluster API creates a new KubevirtMachineTemplate, repoints the MachineDeployment, deletes the old one — and the workers roll. Machines are replaced, never resized in place. |
| Something on the control-plane template | The KamajiControlPlane is patched and the control-plane pods roll. No worker churn. |
| A variable's default | Nothing. Defaults are materialised into spec.topology.variables on the Cluster when it is admitted, so a cluster keeps the value it was created with even after you change the default. |
| A variable's schema — new enum value, different bounds | Nothing now. It governs the next write to that Cluster, so a value that is no longer allowed keeps running and fails the next time the topology is edited. |
So the one case that needs a human is a variable value: to move an existing cluster onto a new size, or onto a default you have since changed, the Cluster's own topology has to say so.
# What a cluster is actually asking for
kubectl get cluster solar -n photon-prod -o jsonpath='{.spec.topology.variables}'
# Change it — the tenant's manifest re-applied, or in place
kubectl edit cluster solar -n photon-prod
Do not merge-patch spec.topology.variables
It is a list, and kubectl patch --type=merge replaces a list wholesale rather than merging by name.
Patching one variable that way drops every other variable the cluster had — its subnet, its datastore, its control-plane address — and the topology controller will happily reconcile the cluster it is left with.
Changing a variable that feeds a worker patch rolls that cluster's workers, so it is a maintenance window, not a tweak. Whether the tenant does it or you do it on their behalf is a policy question; the object is theirs.