Platform Components¶
kMetal installs a curated set of components on the under cluster. Together they turn an existing Kubernetes cluster into a multi-tenant platform that runs tenant clusters.
The set is not assembled per deployment. A kMetal release pins which components are installed, what version each one is, and the order they go in — so upgrading kMetal means moving to a new release, not re-deriving a dependency graph.
What kMetal does not install¶
Three things are deliberately yours to bring, because kMetal has no way to be right about them:
- The primary CNI. Kube-OVN runs as a non-primary CNI and never takes over pod
eth0. Pod networking must already work before the operator starts. - Storage. No CSI driver, no provisioner, no
StorageClass, noVolumeSnapshotClass. You name existing classes in theKMetalspec. - Credentials. Registry credentials and the backup destination are referenced as Secrets you create out of band, never carried in the spec.
Capabilities¶
kMetal is best read as eight capabilities. Each is a job the platform does for you, and each is delivered by components that are pinned, installed and upgraded together.
| Capability | Delivered by |
|---|---|
| Hosted Control Planes | Kamaji |
| Multi-Tenancy | Capsule |
| MultiCluster Management | Cluster API, with the Kamaji and KubeVirt providers and the kMetal ClusterClass |
| Networking | Kube-OVN, MetalLB, kMetal Networking |
| Nodes Management | KubeVirt with KVM, CDI, Kairos, the image builder |
| Data Protection | kmetal-backup, kubevirt-csi-driver-operator |
| Fleet Management | Sveltos, Flux |
| Under Cluster Node Lifecycle | YAKI Operator |
Hosted Control Planes¶
Kamaji.
A tenant cluster's control plane does not run on the tenant's own machines.
It runs as pods on the under cluster — api-server, controller-manager, scheduler — with its etcd data on platform storage.
That is what makes a cluster cheap: no three dedicated control-plane machines per tenant, no idle capacity waiting for a failover that rarely comes. It is also what makes a cluster fast to create, because standing one up is scheduling pods rather than provisioning hardware.
Multi-Tenancy¶
Capsule.
A Tenant owns Environments (translated into Namespace in the Kubernetes context) — and everything created inside them.
Capsule enforces that ownership: who may write where, what quotas apply across a tenant's Environments,
how many Environments a Tenant can provision, or which storage classes their workloads may ask for.
This is the boundary on the under cluster itself. It sits alongside the network and compute boundaries below, and the three together are what hard multi-tenancy means here.
Who that boundary applies to is the one part of it kMetal does not decide.
Capsule acts only on requests from subjects the platform has declared, and the platform declares none by default — who owns a tenant is your identity provider's answer, so spec.multiTenancy on the KMetal object is where it gets stated.
Until it is, Capsule is installed and enforcing nothing.
See Make Capsule aware of the owner.
kMetal admission webhook
A tenant's quotas live where they are enforced: on the under cluster, next to the Tenant. But a tenant does not work there — they work inside their own cluster, where storage simply appears, provisioned transparently by the platform's CSI driver. Nothing in that cluster knows a quota exists.
Left alone, that gap is the worst kind: a PersistentVolumeClaim is accepted, then sits Pending forever,
with nothing in the tenant cluster to explain why.
The tenant has no way to tell an exhausted quota from a broken storage backend,
and finding out means asking the platform team.
The kMetal webhook closes it: it runs on the under cluster (one process serving every tenant cluster)
and installs itself into each tenant cluster as an admission webhook.
A PVC that would take a tenant past their storage quota is rejected at kubectl apply, with a message saying so.
The tenant learns about the limit at the moment they hit it, in the terminal they hit it from,
and can act on it — request more, or clean up.
The same mechanism guards external IPs: a Service asking for an address claim that is already bound elsewhere is refused up front,
rather than silently never getting one.
MultiCluster Management¶
Cluster API, with the Kamaji control-plane provider, the KubeVirt infrastructure provider, and the kMetal ClusterClass.
A tenant cluster is a Cluster object.
Creating, scaling, upgrading and deleting one is a change to that object,
and the providers reconcile the rest — control plane, machines, bootstrap, networking.
The ClusterClass is what keeps that honest across a fleet:
tenants pick a version and a machine size rather than assembling a cluster's parts,
so every cluster on the platform is built the same way.
Customisation occurs using the ClusterClass variables, strictly-typed values Tenant can opt in.
Networking¶
Kube-OVN and MetalLB
Each tenant gets a VPC (a routing domain with no path to any other tenant's) and one or more subnets inside it, carried on a Geneve overlay. MetalLB provides the addresses tenant control planes and platform services are reached on.
kMetal Networking
The resources a tenant is really asking for — a VPC, a subnet, a public address — are cluster-scoped in Kube-OVN. That is the awkward shape for multi-tenancy, and it is awkward twice over.
Cluster-scoped means no namespace, and a namespace is what every per-tenant boundary in Kubernetes is built on.
kMetal Networking is the claims layer: VpcClaim, SubnetClaim and EipClaim,
which let a tenant declare the network they need without being handed cluster-scoped Kube-OVN objects.
There is no way to say "this tenant may create VPCs, but only their own", and no way to stop a tenant reading,
editing or deleting another tenant's.
It also means no quota: ResourceQuota counts objects in a namespace,
so an object that lives outside every namespace cannot be counted at all.
And these are exactly the resources that need counting. Routable IPv4 is finite, shared, and the one thing on the platform that runs out — one tenant claiming addresses freely is one tenant able to exhaust the pool everybody else draws from. Subnets and VPCs are cheaper, but they still consume address space and OVN state that no single tenant should be able to take without limit.
A claim is a namespaced object in the tenant's Environment, and that changes both problems into ones the platform already knows how to solve. Capsule owns the namespace, so ownership and RBAC come for free — a tenant can only claim within their own Environments. And because claims are ordinary namespaced objects, they can be counted: a Tenant can be capped at so many external addresses, so many subnets, the same way it is capped on CPU or storage. A tenant that has used their allocation is told so when they claim, not when something quietly fails to get an address.
The indirection buys one more thing. A claim carries intent ("I want external access", "I want this CIDR") while the controller writes the object. The parts that must be right for isolation to hold, like the policy routes that keep one VPC off the provider segment and away from its neighbours, are injected by the platform and are never the tenant's to state or to omit.
Egress leaves through a provider network, and a tenant that needs its traffic separated on the wire can be given one of its own.
Nodes Management¶
KubeVirt with KVM, CDI, Kairos, and the image builder.
Tenant worker nodes are virtual machines on the under cluster's bare metal.
KVM gives each tenant's nodes real hardware-level separation (a kernel boundary, not a namespace one)
while the nodes remain Kubernetes objects that Cluster API can create and replace.
CDI provides the volumes those VMs boot from, cloning each machine's root disk from a golden image on the platform's own StorageClass.
Both halves are pinned to the nodes labelled node-role.kubernetes.io/compute: virt-handler, which is what allows a VM to start on a node at all, and CDI's importer pods, which fill the disk it will boot from.
So tenant machines and their volumes resolve to the same nodes by construction, and neither lands beside the platform that is supposed to survive them.
Kairos brings immutable-OS nodes, and the image builder lets you build the OS artifacts declaratively,
as Kubernetes objects (no Packer, no Ansible, no manual importing).
Data Protection¶
kmetal-backup.
A backup is an object next to the cluster it protects: a ClusterBackup in the same Environment, naming the Cluster by name.
That placement is the whole isolation story — Capsule already owns the namespace, so a tenant can back up their own clusters and reach nobody else's.
The controller pulls the cluster's kubeconfig, enumerates its API objects, archives them, encrypts the archive client-side, and uploads it to S3 or any S3-compatible endpoint.
Encryption is optional and effectively mandatory in practice: the archive carries full Secret contents.
What it captures is objects, not bytes.
Anything a controller will recreate is skipped, and a PersistentVolume only survives the round trip if its reclaim policy is Retain — a Delete-policy volume dies with the cluster, so the archive keeps the claim and lets the new cluster provision fresh storage for it.
Deciding which of your volumes are Retain is therefore a decision to make before you need the backup, not after.
Restore is not a separate object. Create a new cluster, annotate it with the backup to replay, and the workloads land in it once its control plane is up. The new cluster gets its own CA and identity — certificates are never archived — so a restore produces a genuinely new cluster carrying the old one's workloads, rather than a resurrection of the old one. The annotation is only honoured on a cluster that has not been provisioned yet, which keeps a stray annotation from replaying an archive over a running cluster.
See Cluster Backup & Restore for the walkthrough.
Volume snapshots, via the kubevirt-csi-driver-operator.
A ClusterBackup captures objects, not bytes.
The bytes are covered by the other half of data protection: the storage a tenant's workloads actually sit on, and their ability to snapshot it.
The kubevirt-csi-driver-operator is what puts that in a tenant's hands.
It runs once on the under cluster and drives the CSI driver for every tenant cluster, rather than a controller per tenant — and into each tenant cluster it installs only the node-side pieces: the node DaemonSet, a StorageClass, a VolumeSnapshotClass, and the snapshot controller with its CRDs.
The effect inside a tenant cluster is that storage is entirely ordinary.
A tenant creates a PersistentVolumeClaim against a StorageClass that is simply there,
and a VolumeSnapshot against a VolumeSnapshotClass that is simply there.
Provisioning, attaching, snapshotting and resizing all resolve to volumes on the under cluster's storage,
hot-plugged into that tenant's worker VMs — none of which the tenant configures, holds credentials for, or can see.
Snapshots are also what makes the platform itself work:
tenant machines boot from a golden OS artifact cloned through a snapshot,
which is why kMetal asks the class backing its own volumes to be served by a driver shipping a VolumeSnapshotClass.
See Storage for the model, and Tenant Storage (CSI) for how tenant volumes are provisioned.
Fleet Management¶
Sveltos and Flux.
Flux reconciles the undercloud platform itself: every component kMetal installs is a Flux object, continuously converged and drift-corrected. Sveltos does the same job in the other direction — delivering add-ons into tenant clusters by label, so a fleet of clusters carries a consistent set of them without being configured one at a time.
kMetal installs Sveltos and declares nothing through it. Which add-ons a tenant cluster gets, starting with its CNI, is the platform's decision — see Fleet Management Tasks.
Under Cluster Node Lifecycle¶
YAKI Operator.
Every capability above is about tenant clusters. This one is about the machines kMetal itself runs on, which nothing else in the platform would otherwise move: the under cluster is self-managed, and no node-pool API can upgrade its kubelets.
A node upgrade is a KubernetesNodeUpgrade resource.
The operator plans which nodes move and in what order, upgrades the control plane before the workers so the version skew policy holds by construction, bounds how many workers go at once, and waits for each node to report Ready at the target version before continuing.
The resource then freezes, so kubectl get knu reads as the under cluster's upgrade history.
The upgrade itself is YAKI, the CLASTIX bootstrap and upgrade script, run in the host namespace of each node — the same script that is the recommended way to build the under cluster in the first place. That is the whole argument for it: no Terraform run, no configuration-management pass and no bare-metal provisioning stack to keep the platform's own nodes aligned.
See Node Upgrades.
Supporting components¶
Three more sit underneath, and are worth knowing about only when something goes wrong:
cert-manager, and the kmetal-ca issuer |
Certificates for the platform's webhooks, controllers and console. Most of the platform waits on it during an install. |
The CDI StorageProfile, and its filesystem overhead |
How CDI provisions volumes on the class you named for tenants. Pinned by the platform because the defaults are wrong on some backends — see Storage. |
| Headlamp, with the kMetal plugin | The web console. Nothing else depends on it, and it can be left out. |
How they are delivered¶
Flux does the installing; the kMetal operator owns the sequence.
Every component above is expressed as a Flux object: a HelmRelease, an OCIRepository, a Kustomization.
Flux itself is component zero: the operator applies it directly to the API server from manifests embedded at build time,
so nothing has to be bootstrapped first.
The operator then walks the install order front to back, applies each component, reads back what Flux reports, and halts at the first component that is not ready. It publishes one status entry per component, so how far the platform got is readable off a single object: see Component Reference for the phases each component moves through and how to read them.
Nothing is rolled back. A failed upgrade leaves the older working version serving, and the components after it running whatever they already ran.
What is configurable¶
Very little, on purpose.
The KMetal spec carries the values that have no sane default — the overlay interface, the provider networks,
the address pools, the StorageClass names — plus a few explicit escape hatches.
Versions, feature gates, and the install order are absent from the API because a site that could change them would no longer be running a tested release.
See Platform Configuration for the full surface, and Component Configuration for the escape hatches that do exist.
For the component-by-component view — install order, dependencies, and the categories the operator reports on a stalled install — see Component Reference.