Skip to content

Storage

kMetal is storage agnostic. It installs no CSI driver, no provisioner, no StorageClass and no VolumeSnapshotClass, and it manages no storage cluster.

You bring storage the same way you bring the primary CNI: it exists before kMetal is installed, and you name it in the KMetal spec.

That is a deliberate position. Storage is the part of a bare-metal platform that organisations already have opinions, contracts and hardware for — an array, a Ceph cluster, a vendor CSI — and a platform that shipped its own would either be ignored or be in the way.

What kMetal expects of it

One requirement beyond "a working CSI driver": volume snapshots.

The class backing the platform's own volumes must be served by a driver shipping a VolumeSnapshotClass, because snapshotting and cloning are how tenant machines come into existence. The image builder snapshots the disk it just installed the OS onto, that snapshot becomes the golden volume, and CDI clones a fresh root disk from it for every machine. A class that cannot snapshot leaves the artifact with no golden volume and every tenant machine with no disk to boot. A clone the backend cannot take becomes a full image copy per machine, which makes scaling a cluster slow enough to be a different product.

That requirement is also why the platform's own storage is split in two below: it rules out most node-local provisioners, which are otherwise a reasonable place for etcd.

Everything else is ordinary: dynamic provisioning, and enough capacity for the three jobs below.

Three jobs, three classes

The KMetal spec names up to three StorageClass references, because the consumers pull in different directions. Pointing everything at one class silently gets one of them wrong.

Backs Wants
underclusterClassName The platform's own volumes: the golden OS images, and the root disk of every tenant machine cloned from them. Snapshots, before anything else — without them nothing boots. Capacity for one volume per machine, plus one golden volume per image.
etcdClassName The Kamaji datastore holding every tenant control plane's etcd data. Optional; defaults to underclusterClassName. Low latency. Small. Never touched by tenants. A node-local class is defensible, but it ties the durability of every hosted control plane to one node.
tenantClassName The volumes tenant workloads claim, through the KubeVirt CSI driver. Network-attached, and able to survive the loss of any single node — the opposite of what the class above may reasonably be.

etcdClassName exists because the first two wants are incompatible on most backends. The low-latency node-local classes generally cannot snapshot, and etcd has no use for a snapshot even where the class offers one. Leave it unset on a cluster with a single class that does both; set it where the class the golden images need is the wrong class for etcd.

The platform class and the tenant class stay two names regardless, and the reason is quota rather than performance — see below.

Persistent machines

Tenant worker nodes are virtual machines, and a virtual machine is only as durable as the disk under it.

Their disks are volumes on underclusterClassName, provisioned by CDI as DataVolumes cloned from the golden image, in the tenant's own namespace on the under cluster. Because those volumes are real persistent storage rather than node-local scratch, a worker VM survives the loss of the host it was running on: it is rescheduled onto another under-cluster worker node and comes back with its disk.

This is also why neither the platform class nor the tenant class should be node-local. A node-local class makes every tenant worker node a machine that dies permanently with its host — which a tenant will experience as their cluster losing nodes.

How tenant clusters consume it

A tenant cluster has no storage hardware of its own. It gets storage through the KubeVirt CSI driver, split across the two layers:

  • The CSI controller runs on the under cluster, where the storage actually is.
  • The CSI node component runs inside the tenant cluster, where the workload is.

A PersistentVolumeClaim created inside a tenant cluster therefore becomes a volume on the under cluster's storage, hot-plugged into that tenant's worker VM and mounted into the pod. From inside the tenant cluster it is an entirely ordinary PVC against an ordinary StorageClass — the tenant does not configure a backend, hold credentials for one, or know what it is.

That transparency is the point, and it has a security consequence worth stating plainly: the credentials for the storage backend live on the under cluster with the CSI controller, and never inside a tenant cluster. A tenant consumes the storage without ever holding a key to it — see Tenant Layer for what that is worth when a tenant cluster is compromised.

The transparency is also what creates the problem the next section solves.

Quota, and why it is enforced twice

Storage is finite and shared, so a tenant's consumption has to be bounded. The bound lives where tenants are defined — on the under cluster, as a quota on the Tenant, counted against tenantClassName.

That quota is keyed on the class name, not on who the volume belongs to. Machine root disks are provisioned in the tenant's own namespace, so naming one class under both underclusterClassName and tenantClassName has the platform's own volumes counted against the budget the tenant reads as theirs — which is why the two stay separate names even when one backend serves both. See Keep the platform class and the tenant class apart.

But a tenant does not work on the under cluster. They work inside their own cluster, where storage simply appears, and where nothing natively knows that a quota exists somewhere else.

Left alone, that gap produces the worst failure mode a platform can offer: a PVC accepted, then Pending forever, with no way for the tenant to tell an exhausted quota from a broken storage backend.

So kMetal enforces it in both places. The quota is held on the under cluster, and an admission webhook installed into each tenant cluster rejects an over-quota PersistentVolumeClaim when it is created, with a message saying so. The tenant finds out at kubectl apply, in the terminal they applied from, and can act: request more, or clean up.

See Tenant Layer for the same principle applied to the rest of what a tenant claims.

Where to go next