Skip to content

Tenant Storage (CSI)

Tenant clusters get persistent storage from the under cluster's own storage, through the KubeVirt CSI driver — but not by running that driver per tenant.

kMetal runs kubevirt-csi-driver-operator: one Deployment on the under cluster that drives the CSI controller for every tenant cluster in-process, and provisions only the node half into each tenant. There are no per-tenant controller pods and no sidecar sets to multiply.

Topology

Under cluster                                Tenant cluster
──────────────────────────────────────       ─────────────────────────────
[ kubevirt-csi-driver-operator ]             [ kubevirt-csi-node ]
  one Deployment, all tenants                  DaemonSet, kube-system
  · ControllerService in-process             [ csi.kubevirt.io ] CSIDriver
  · provision / attach / snapshot / resize   [ kubevirt ] StorageClass (default)
    as reconcilers over tenant objects       [ kubevirt-csi-snapclass ]
                                             [ snapshot-controller + CRDs ]
[ DataVolume in the Environment ]  ◀── hot-plugged as a disk to the worker VM
[ ResourceQuota on the tenant class ]
[ kmetal-webhook ] ──── pushes a ValidatingWebhookConfiguration ───▶ (PVC admission)

What the operator does per engaged cluster:

  • Runs the CSI ControllerService as a library, talking to the under cluster to create DataVolumes, hot-plug them onto the tenant's worker VMs, snapshot and resize them.
  • Replaces the four sig-storage sidecars with reconcilers over the tenant's own PersistentVolumeClaim, PersistentVolume, VolumeAttachment and VolumeSnapshotContent objects. Provisioning uses sig-storage-lib-external-provisioner; attach, snapshot and resize are handled directly.
  • Applies the node bundle into the tenant cluster with server-side apply: the kubevirt-csi-node DaemonSet, the csi.kubevirt.io CSIDriver, the default kubevirt StorageClass, the kubevirt-csi-snapclass VolumeSnapshotClass, the snapshot controller and its CRDs, and scoped RBAC — all in kube-system.

Clusters are discovered through Cluster API: every Cluster in every namespace, unless a label selector gates the rollout.

The identity it uses

The tenant's admin kubeconfig is used only to bootstrap. The operator creates a scoped kubevirt-csi-operator-agent ServiceAccount in the tenant's kube-system, then mints auto-refreshed TokenRequest tokens for all data-plane traffic — so the long-lived admin credential is not what the steady state runs on.

Where a volume lives

Every tenant PVC becomes a DataVolume in the tenant cluster's own namespace on the under cluster — the Environment — on the class named by storage.tenantClassName. It carries csi.kubevirt.io/cluster: <cluster-name>, which is how the operator knows whose it is, and what a restore relabels when volumes move to a new cluster.

The infra namespace is resolved by convention

The operator takes the tenant's infra namespace from the CAPI Cluster's own namespace. That is correct for the layout kMetal creates, where the worker VMs live beside the Cluster. A KubevirtCluster configured to place node VMs in a different namespace is not supported — every controller operation would target the wrong one.

Follow a volume end to end

A tenant applies an ordinary claim in their own cluster. Everything below is the same volume, seen from the under cluster — here the photon Environment, which holds two clusters:

$ kubectl get datavolume -n photon \
    -o custom-columns='NAME:.metadata.name,CLUSTER:.metadata.labels.csi\.kubevirt\.io/cluster,PHASE:.status.phase'
NAME                                       CLUSTER   PHASE
pvc-0e71dc33-7620-4a6a-b2ad-53ce9f440687   silicon   Succeeded
pvc-47a0148a-f861-440d-9ee9-147e30246d35   quantum   Succeeded

Two things to read off that. One Environment can hold volumes for several clusters — quantum and silicon both live in photon — and csi.kubevirt.io/cluster is the only thing separating them, which is why a restore has to relabel it.

The DataVolume itself is unremarkable, which is the point: a blank volume of the requested size on storage.tenantClassName.

apiVersion: cdi.kubevirt.io/v1beta1
kind: DataVolume
metadata:
  name: pvc-0e71dc33-7620-4a6a-b2ad-53ce9f440687
  namespace: photon
  labels:
    csi.kubevirt.io/cluster: silicon            # set by the operator
    capsule.clastix.io/managed-by: photon       # stamped by Capsule's webhook
    projectcapsule.dev/tenant: photon
spec:
  source:
    blank: {}
  storage:
    resources:
      requests:
        storage: "1073741824"                   # 1Gi, as the tenant asked
    storageClassName: tenant-storage-class
status:
  phase: Succeeded

CDI turns it into a PVC of the same name in the same namespace, bound to a volume from your backend:

$ kubectl get pvc -n photon
NAME                                       STATUS   VOLUME                                     CAPACITY   STORAGECLASS
pvc-0e71dc33-7620-4a6a-b2ad-53ce9f440687   Bound    pvc-4e04885b-56c1-42e6-85fe-1625896f6632   1Gi        tenant-storage-class

$ kubectl describe pvc pvc-0e71dc33-7620-4a6a-b2ad-53ce9f440687 -n photon | grep -E 'Annotations|Used By'
Annotations:   cdi.kubevirt.io/createdForDataVolume: a0f14e2a-1177-4e59-8ccc-8b452c413682
Used By:       <none>

Used By: <none> is the normal steady state for a volume nobody is mounting: provisioning and attachment are separate steps.

The names, in both directions

In the tenant cluster On the under cluster
PersistentVolumeClaim the tenant wrote, any name, any namespace —
PersistentVolume, which Kubernetes names pvc-<tenant PVC UID> DataVolume and its PVC, both carrying that same name, in the Environment
— a second PersistentVolume, provisioned by your own CSI driver on tenantClassName

So the volume handle is the join key: given a tenant PV, the object on the under cluster has the same name; given a DataVolume, the label says which cluster it belongs to.

# All volumes for one tenant cluster
kubectl get datavolume,pvc -n photon -l csi.kubevirt.io/cluster=silicon

Attachment is a hot-plug

When a pod in the tenant cluster mounts the claim, the operator hot-plugs that DataVolume onto the worker VM the pod was scheduled to, and the disk appears in the VM's volume list:

$ kubectl get vmi -n photon -o custom-columns='NAME:.metadata.name,VOLUMES:.status.volumeStatus[*].name'
NAME                             VOLUMES
quantum-md-0-rqxks-jcwb7-7pcms   cloudinitvolume,rootdisk
silicon-md-0-z5m6f-wjcwr-n2fbm   cloudinitvolume,rootdisk

Those two VMs are running with nothing attached — rootdisk and cloudinitvolume are all a worker starts with. A mounted volume shows up as a third entry named after the DataVolume, and disappears again when the pod moves, because attachment follows the pod rather than the cluster.

Per-tenant quota

The cap is a ResourceQuota on <tenant class>.storageclass.storage.k8s.io/requests.storage, and on this platform it is set through Capsule so it aggregates across all of a tenant's Environments:

apiVersion: capsule.clastix.io/v1beta2
kind: Tenant
metadata:
  name: dynamo
spec:
  resourceQuotas:
    scope: Tenant
    items:
      - hard:
          tenant-storage-class.storageclass.storage.k8s.io/requests.storage: 20Gi

Two things enforce it, at different moments:

  1. The under cluster's api-server, when the operator creates the DataVolume. This is the hard ceiling — nothing provisions past it.
  2. kmetal-webhook, at PVC admission inside the tenant cluster. One Deployment in kmetal-system serves every tenant, routing requests by the cluster UID embedded in the webhook URL; it reads the matching ResourceQuota and rejects the PVC up front.

Both enforce the same number. The second exists so the tenant gets a clear rejection on their own PVC instead of a claim that stays Pending while a DataVolume fails somewhere they cannot see.

Machine root disks count against it if the classes are the same

Worker VMs boot from volumes on storage.underclusterClassName, and those volumes are provisioned in this same namespace. The quota key is a class name, so pointing underclusterClassName and tenantClassName at one class has the platform's own volumes consume the tenant's workload budget:

$ kubectl get resourcequota -n dynamo-dev \
    -o jsonpath='{.items[0].status.used}{"\n"}'
{"tenant-storage-class.storageclass.storage.k8s.io/requests.storage":"33285996544"}

31 GiB, for a tenant with one 1 GiB volume and one worker on a 30Gi boot tier. Two StorageClass objects over the same driver keep the two apart — see Keep the platform class and the tenant class apart.

Capsule maintains the object itself, one per Environment, named capsule-<tenant>-N after the item's index in the Tenant:

$ kubectl get resourcequota capsule-photon-0 -n photon \
    -o custom-columns='NAME:.metadata.name,HARD:.status.hard.tenant-storage-class\.storageclass\.storage\.k8s\.io/requests\.storage,USED:.status.used.tenant-storage-class\.storageclass\.storage\.k8s\.io/requests\.storage'
NAME               HARD    USED
capsule-photon-0   100Gi   2147483648

That is the tenant's whole budget in the used column — 2 GiB here, one volume in quantum and one in silicon — because the scope is the tenant rather than the namespace. A tenant cannot get more storage by asking for another Environment, and the number to watch is this one rather than the sum of what any single cluster claimed.

The webhook's ValidatingWebhookConfiguration is pushed into each tenant cluster and maintained there: its controller re-applies on tenant-side edits and deliberately reverts attempts to narrow namespaceSelector, objectSelector or matchPolicy — narrowing it is not a way around the quota. It is registered failurePolicy: Fail, which makes its reachability a hard dependency:

The tenant control plane must reach kmetal-webhook.kmetal-system.svc:443

A tenant api-server runs as pods in the Environment, and that is what calls the webhook. Block that egress and PVC creation in the tenant cluster is denied, not merely unmetered — see the warning in Multi-Tenancy.

What this means for the tenant

Inside the tenant cluster, the integration is invisible — one default StorageClass, ordinary claims:

apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: my-data
spec:
  accessModes: [ "ReadWriteOnce" ]
  storageClassName: kubevirt        # the default; omitting it works too
  resources:
    requests:
      storage: 5Gi

The data lives on whatever the under cluster runs — Ceph, an enterprise array, a vendor CSI driver — and the tenant never sees KubeVirt, the DataVolume, or the class behind it.

Snapshots work the same way, through kubevirt-csi-snapclass. The platform's own snapshot requirement is on underclusterClassName, where the golden images live; this is the second, tenant-visible half of it — a tenant VolumeSnapshot becomes a snapshot of a DataVolume on tenantClassName, so that driver has to ship a snapshot class too, or snapshots are a feature tenants can see and not use. See Storage.

The kubevirt StorageClass is created with reclaimPolicy: Delete, and a StorageClass's reclaim policy is immutable. Offering Retain — which is what makes a volume survive into a restored cluster — means an additional class in the tenant's cluster: see Volumes move with a restore.

Verify

# The operator, and which clusters it engaged
kubectl get deploy kubevirt-csi-driver-operator -n kmetal-system
kubectl logs -n kmetal-system deploy/kubevirt-csi-driver-operator | grep -i engage

# Volumes for one tenant cluster, on the under cluster
kubectl get datavolumes,pvc -n photon -l csi.kubevirt.io/cluster=silicon

# The quota that bounds every cluster in the tenant
kubectl get resourcequota -n photon -o yaml | grep -A3 storageclass

# Anything stuck: a DataVolume that never reaches Succeeded is a backend problem,
# a tenant PVC with no DataVolume at all is an operator or webhook problem
kubectl get datavolume -n photon -o custom-columns='NAME:.metadata.name,PHASE:.status.phase,PROGRESS:.status.progress'

# The node bundle, from inside the tenant cluster
kubectl --kubeconfig silicon.kubeconfig get csidrivers,storageclasses,volumesnapshotclasses
kubectl --kubeconfig silicon.kubeconfig get ds kubevirt-csi-node -n kube-system