Skip to content

Storage Configuration

kMetal installs no storage: no CSI driver, no provisioner, no StorageClass, no VolumeSnapshotClass. You bring storage the way you bring the primary CNI — it exists before kMetal is installed — and name it in the KMetal spec.

This page is about what the platform requires of that storage, and what you configure once it exists. For the model, see Storage Concepts. For the tenant-facing wiring, see Tenant Storage (CSI).

What the platform requires

Requirement Why
Dynamic provisioning Every tenant volume and every tenant worker disk is provisioned on demand.
A VolumeSnapshotClass on the platform class's driver The image builder snapshots the disk it installed the OS onto, and CDI clones every machine's root disk from that snapshot. A class that cannot snapshot leaves the artifact with no golden volume and machines with no disk to boot; a clone the backend cannot take becomes a full image copy per machine.
allowVolumeExpansion Tenants resize volumes; the CSI layer passes the request through.

The snapshot requirement moved

It used to sit on tenantClassName, because that is where the golden images were built. They are on underclusterClassName now, together with the machine root disks cloned from them — so a node-local platform class, which used to be a defensible choice, no longer boots a single tenant machine. That is what etcdClassName below exists to give back.

Anything meeting those can back kMetal — a vendor array, Rook-Ceph, Longhorn, an NFS CSI driver in a lab. The examples further down this page are illustrative starting points, not kMetal-specific configuration.

Naming the classes

spec:
  storage:
    underclusterClassName: platform-storage      # golden images + machine root disks
    etcdClassName: local-path                    # optional; defaults to the class above
    tenantClassName: tenant-storage-class        # handed to tenant workloads

Three fields, because the consumers pull in different directions. Pointing everything at one class silently gets one of them wrong.

Field Required Backs Wants
underclusterClassName yes The golden OS images in kmetal-image-builder, and the root disk of every tenant machine cloned from them. Snapshots. Capacity for one volume per machine plus one golden volume per image.
etcdClassName no The Kamaji datastore holding every hosted control plane's etcd data. Defaults to underclusterClassName. Low latency, small. Node-local is defensible, but ties the durability of every tenant control plane to one node.
tenantClassName yes The volumes tenant workloads claim, through the KubeVirt CSI driver. Network-attached, surviving the loss of any single node.

A node-local class under underclusterClassName or tenantClassName makes every tenant worker node a machine that dies permanently with its host — which tenants experience as their cluster losing nodes. Under etcdClassName it is a deliberate trade instead: latency and simplicity, paid for with the etcd data of every hosted control plane if the node holding it is lost.

Set etcdClassName only where one class cannot serve both platform jobs. On a cluster whose single class snapshots and is fast enough for etcd, leave it unset and say so once in a comment.

All three are hard to change once tenants exist, because their data is on them. Decide before the first tenant — see Configuration Best Practices.

Keep the platform class and the tenant class apart

Give them two StorageClass objects even where the same driver and the same pool back both. The reason is quota, not performance.

A tenant's storage quota is a ResourceQuota keyed on <class>.storageclass.storage.k8s.io/requests.storage, so it counts by class name rather than by who the volume belongs to. Machine root disks are provisioned in the tenant's own namespace on underclusterClassName, which means one shared name has the platform's own volumes billed to the tenant:

$ kubectl get resourcequota -n dynamo-dev -o yaml | grep -A2 'hard:'
    hard:
      tenant-storage-class.storageclass.storage.k8s.io/requests.storage: 59Gi
    used:
      tenant-storage-class.storageclass.storage.k8s.io/requests.storage: "33285996544"

31 GiB used on a tenant running one 1 GiB workload volume: the other 30 GiB is a single worker's boot disk. Two workers on a 30Gi tier consume 60 GiB of the tenant's budget before a workload claims anything, and every cluster they create moves the number again — so the quota stops meaning what it says, and raising it stops being arithmetic anyone can do.

Two classes over one backend costs nothing but a second StorageClass object:

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: platform-storage        # underclusterClassName
provisioner: rook-ceph.rbd.csi.ceph.com
parameters:
  clusterID: rook-ceph
  pool: replicapool             # the same pool as the tenant class, deliberately
  # ... the rest identical to the tenant class
allowVolumeExpansion: true

Machine disks then count against a key nobody's quota names, and the tenant's budget is the tenant's workloads. Bound the machine side where it belongs instead — on the cluster and machine counts the tenant may ask for, which is what actually drives it. See Cap what a tenant can consume.

One VolumeSnapshotClass still has to match the platform class

Whatever you name here, clusterClass.values.dataVolume.volumeSnapshotClass must be served by the same driver as underclusterClassName — that is the snapshot the golden volume is taken with. Same driver is the requirement, not the same class name.

Pinning CDI's capabilities

CDI provisions the volumes tenant machines boot from, and it decides how by reading the storage's advertised capabilities. On a backend advertising more than one combination, its derived order can be wrong for kMetal:

spec:
  storage:
    tenantClaimPropertySets:
      - accessModes:
          - ReadWriteOnce
        volumeMode: Filesystem

The operator turns this into a CDI StorageProfile for tenantClassName.

Order is significant, and the failure is opaque

CDI takes the first entry for a DataVolume that asks for nothing specific.

A backend that advertises ReadWriteMany + Block first has CDI provision raw-block volumes, and the importer pod then fails on a permission error that names neither the storage class nor the volume mode.

Leave the list empty on a single-capability backend, where auto-derivation can only produce one answer. Set it on anything else, before the first tenant volume.

CDI's filesystem overhead is pinned to 0 on the tenant class by the platform, and is not a field. The default 6% reservation pushes a whole-GiB DataVolume just past the next GiB boundary, and a driver that rounds up from there — Ceph RBD does — provisions twice what the tenant asked for. Only that one class is touched, so an overhead set for any other class is left alone.

Both pins land on the tenant class only

tenantClaimPropertySets renders a StorageProfile for tenantClassName, and the overhead pin names tenantClassName. Neither reaches underclusterClassName — which is where the golden images and every machine root disk now live.

On a platform class distinct from the tenant class, those volumes therefore take CDI's auto-derived capabilities and its default 6% overhead. On a whole-GiB boot tier that is the same rounding problem, one size up: a 30Gi tier asks the backend for ~31.8 GiB, and a driver rounding to the next power-of-two-friendly size provisions 32 GiB per machine. Set the overhead for that class yourself if it bites, on the CDI object — kMetal leaves other classes alone precisely so that edit survives:

kubectl patch cdi cdi --type=merge \
  -p '{"spec":{"config":{"filesystemOverhead":{"storageClass":{"platform-storage":"0"}}}}}'

Read what CDI derived for the class before assuming its defaults are wrong — cloneStrategy is the line that decides whether machine creation is a snapshot or a copy:

$ kubectl get storageprofile platform-storage \
    -o jsonpath='{.status.cloneStrategy}{"\t"}{.status.snapshotClass}{"\n"}'
csi-clone   platform-storage-snapshot

Swapping the backend

Because the classes are named in the spec rather than created by kMetal, changing backend is a storage operation followed by a one-line spec change:

  1. Install the new CSI driver on the under cluster, with its StorageClass and VolumeSnapshotClass.
  2. Verify it provisions and snapshots — a driver without a working VolumeSnapshotClass cannot boot tenant machines.
  3. Point the class you are moving at the new class: storage.tenantClassName for tenant workloads, storage.underclusterClassName for the images and machine disks.
  4. Moving underclusterClassName also means clusterClass.values.dataVolume.volumeSnapshotClass has to move to a class on the new driver, in the same change, or golden snapshots stop becoming ready.
  5. Revisit tenantClaimPropertySets for the new backend's capabilities.

Existing volumes do not migrate. Data already provisioned stays on the old backend, bound to its PVs, until you move it deliberately — plan a swap as a migration, not a cutover.

Backend examples

None of the following is kMetal configuration — it is a starting point for standing up a backend that meets the requirements above.

NFS

One option among many

kMetal bundles no NFS provisioner. The example below uses nfs-subdir-external-provisioner as a common pick; substitute whatever NFS CSI driver your environment runs.

Note that most NFS provisioners ship no VolumeSnapshotClass, which makes them unsuitable for tenantClassName.

# nfs-provisioner.yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: nfs-storage
provisioner: nfs.csi.k8s.io
parameters:
  server: nfs-server.company.com
  share: /exports/kubernetes
  mountPermissions: "0755"
reclaimPolicy: Retain
volumeBindingMode: Immediate
mountOptions:

  - hard
  - nfsvers=4.1
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: nfs-provisioner
  namespace: kube-system
spec:
  replicas: 2
  selector:
    matchLabels:
      app: nfs-provisioner
  template:
    metadata:
      labels:
        app: nfs-provisioner
    spec:
      serviceAccountName: nfs-provisioner
      containers:

      - name: nfs-provisioner
        image: registry.k8s.io/sig-storage/nfs-subdir-external-provisioner:v4.0.2
        volumeMounts:

        - name: nfs-client-root
          mountPath: /persistentvolumes
        env:

        - name: PROVISIONER_NAME
          value: nfs.csi.k8s.io

        - name: NFS_SERVER
          value: nfs-server.company.com

        - name: NFS_PATH
          value: /exports/kubernetes
      volumes:

      - name: nfs-client-root
        nfs:
          server: nfs-server.company.com
          path: /exports/kubernetes

Tenant Cluster Storage

Tenants get persistent storage in their cluster through the KubeVirt CSI driver, which maps onto the class named by storage.tenantClassName. The credentials for the backend stay on the under cluster — nothing inside a tenant cluster holds a key to it. See Tenant Storage (CSI) for how that is wired, and how the per-tenant cap is enforced on both sides.

Per-tenant storage quota

Cap each tenant's total storage with a ResourceQuota in the tenant's under-cluster namespace, keyed on tenant-storage-class.storageclass.storage.k8s.io/requests.storage:

apiVersion: v1
kind: ResourceQuota
metadata:
  name: alpha-storage
  namespace: alpha
spec:
  hard:
    # <tenantClassName>.storageclass.storage.k8s.io/requests.storage
    tenant-storage-class.storageclass.storage.k8s.io/requests.storage: 10Gi

The key is derived from storage.tenantClassName, so it changes if you change that. It counts every volume in the namespace on that class — including machine root disks, if underclusterClassName names the same class. See Keep the platform class and the tenant class apart.

The quota lives here, on the under cluster — but tenants work inside their own cluster, where storage simply appears and nothing natively knows a quota exists.

kMetal closes that gap with the kmetal-webhook component: one process on the under cluster serving every tenant cluster, which installs itself into each as an admission webhook. A PVC that would take a tenant past this quota is rejected when it is created, with a message saying so — rather than accepted and left Pending while the tenant tries to work out whether the platform is broken.

# See what tenants are consuming
kubectl get resourcequota -A -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,USED:.status.used,HARD:.status.hard \
  | grep -i storage

# Raise a tenant's cap
kubectl patch resourcequota alpha-storage -n alpha --type=merge \
  -p '{"spec":{"hard":{"tenant-storage-class.storageclass.storage.k8s.io/requests.storage":"50Gi"}}}'

CDI on the under cluster

KubeVirt's Containerized Data Importer (CDI) handles golden-image imports and the DataVolume lifecycle on the under cluster. kMetal installs it as part of the platform, so there is nothing to deploy.

It matters here because every tenant volume ultimately becomes a CDI-managed DataVolume on the tenant class — which is why pinning CDI's capabilities is the one piece of CDI configuration worth getting right before the first tenant volume exists.

Distributed Storage

Your choice entirely

kMetal ships no storage at all, so what backs it is your pick — Rook-Ceph, Longhorn, a vendor CSI driver, or anything else compatible with the under cluster's Kubernetes version and shipping a VolumeSnapshotClass. The examples below are illustrative starting points, not kMetal-specific configuration.

Rook-Ceph

# rook-ceph-cluster.yaml
apiVersion: ceph.rook.io/v1
kind: CephCluster
metadata:
  name: rook-ceph
  namespace: rook-ceph
spec:
  cephVersion:
    image: quay.io/ceph/ceph:v18.2.0
  dataDirHostPath: /var/lib/rook
  mon:
    count: 3
    allowMultiplePerNode: false
  mgr:
    count: 2
    allowMultiplePerNode: false
  dashboard:
    enabled: true
    ssl: true
  storage:
    useAllNodes: true
    useAllDevices: false
    deviceFilter: "^sd[b-z]"
    config:
      osdsPerDevice: "1"
---
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: rook-ceph-block
provisioner: rook-ceph.rbd.csi.ceph.com
parameters:
  clusterID: rook-ceph
  pool: replicapool
  imageFormat: "2"
  imageFeatures: layering
  csi.storage.k8s.io/provisioner-secret-name: rook-csi-rbd-provisioner
  csi.storage.k8s.io/provisioner-secret-namespace: rook-ceph
  csi.storage.k8s.io/node-stage-secret-name: rook-csi-rbd-node
  csi.storage.k8s.io/node-stage-secret-namespace: rook-ceph
  csi.storage.k8s.io/fstype: ext4
allowVolumeExpansion: true
reclaimPolicy: Delete

Longhorn

# longhorn-storage.yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: longhorn
  annotations:
    storageclass.kubernetes.io/is-default-class: "true"
provisioner: driver.longhorn.io
allowVolumeExpansion: true
reclaimPolicy: Delete
volumeBindingMode: Immediate
parameters:
  numberOfReplicas: "3"
  staleReplicaTimeout: "2880"
  fromBackup: ""
  fsType: "ext4"
  dataLocality: "disabled"

Object Storage

MinIO

# minio-deployment.yaml
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: minio
  namespace: minio-system
spec:
  serviceName: minio
  replicas: 4
  selector:
    matchLabels:
      app: minio
  template:
    metadata:
      labels:
        app: minio
    spec:
      containers:

      - name: minio
        image: quay.io/minio/minio:latest
        args:

        - server
        - --console-address
        - ":9001"
        - http://minio-{0...3}.minio.minio-system.svc.cluster.local/data
        env:

        - name: MINIO_ROOT_USER
          valueFrom:
            secretKeyRef:
              name: minio-credentials
              key: rootUser

        - name: MINIO_ROOT_PASSWORD
          valueFrom:
            secretKeyRef:
              name: minio-credentials
              key: rootPassword
        ports:

        - containerPort: 9000
          name: api

        - containerPort: 9001
          name: console
        volumeMounts:

        - name: data
          mountPath: /data
        resources:
          requests:
            cpu: "500m"
            memory: "1Gi"
  volumeClaimTemplates:

  - metadata:
      name: data
    spec:
      accessModes: [ "ReadWriteOnce" ]
      storageClassName: <your-storage-class>
      resources:
        requests:
          storage: 100Gi

Backup Storage

Operator's choice

kMetal does not bundle a backup tool. The examples below show Velero — a common operator choice — configured against S3-compatible object storage. Adapt to whatever backup tooling your environment uses.

Velero with S3

# velero-backup-location.yaml
apiVersion: velero.io/v1
kind: BackupStorageLocation
metadata:
  name: default
  namespace: velero
spec:
  provider: aws
  objectStorage:
    bucket: kmetal-backups
    prefix: velero
  config:
    region: us-west-2
    s3ForcePathStyle: "false"
    s3Url: https://s3.us-west-2.amazonaws.com
---
apiVersion: velero.io/v1
kind: VolumeSnapshotLocation
metadata:
  name: default
  namespace: velero
spec:
  provider: aws
  config:
    region: us-west-2

Velero with MinIO

# velero-minio-backup.yaml
apiVersion: velero.io/v1
kind: BackupStorageLocation
metadata:
  name: minio
  namespace: velero
spec:
  provider: aws
  objectStorage:
    bucket: velero
  config:
    region: minio
    s3ForcePathStyle: "true"
    s3Url: http://minio.minio-system:9000
    publicUrl: https://minio.company.com

Storage Performance

Benchmark Storage

# Deploy fio for benchmarking
kubectl run fio-benchmark --rm -it --image=ljishen/fio -- /bin/bash

# Run sequential write test
fio --name=seqwrite --rw=write --bs=1M --size=1G --numjobs=1 --runtime=60 --time_based --filename=/data/test

# Run random write test
fio --name=randwrite --rw=randwrite --bs=4K --size=1G --numjobs=4 --runtime=60 --time_based --filename=/data/test

# Run read test
fio --name=read --rw=read --bs=1M --size=1G --numjobs=1 --runtime=60 --time_based --filename=/data/test

Monitor Storage Usage

# Check PV usage
kubectl get pv

# Check PVC usage
kubectl get pvc -A

# Storage capacity
kubectl top pv

# Per-node storage usage
kubectl get --raw /api/v1/nodes/<node-name>/proxy/stats/summary | \
  jq '.node.fs'

Storage Troubleshooting

Debug PV/PVC Issues

# Check PVC status
kubectl get pvc -A
kubectl describe pvc <pvc-name> -n <namespace>

# Check PV binding
kubectl get pv
kubectl describe pv <pv-name>

# Check storage class
kubectl get storageclass
kubectl describe storageclass <storage-class-name>

# Check provisioner logs
kubectl logs -n kube-system -l app=<provisioner-name> -f

CSI Driver Debugging

The controller side is kubevirt-csi-driver-operator — one Deployment on the under cluster serving every tenant cluster, with the CSI controller running in-process rather than as per-tenant pods and sidecars. Only the node DaemonSet runs inside the tenant cluster.

# --- Controller side (under cluster) ---
kubectl get deploy kubevirt-csi-driver-operator -n kmetal-system
kubectl logs -n kmetal-system deploy/kubevirt-csi-driver-operator

# What it provisioned for one cluster, on the under cluster
kubectl get datavolumes -n <environment> -l csi.kubevirt.io/cluster=<cluster>

# --- Node side (tenant cluster) ---
kubectl --kubeconfig=<tenant>.kubeconfig -n kube-system get ds kubevirt-csi-node
kubectl --kubeconfig=<tenant>.kubeconfig -n kube-system logs -l app=kubevirt-csi-node -c csi-driver

# Volume attachments (tenant cluster)
kubectl --kubeconfig=<tenant>.kubeconfig get volumeattachment

For non-kubevirt-csi drivers (NFS, Rook-Ceph, vendor CSI, etc.), the controller usually runs in kube-system or a vendor-specific namespace on the under cluster — adapt the selectors accordingly.

Storage Migration

Requires a snapshot-capable CSI driver

Snapshots need a driver that ships a VolumeSnapshotClass — the same requirement kMetal has for tenantClassName. Update volumeSnapshotClassName and storageClassName below to match what your driver provides.

# Create snapshot — replace volumeSnapshotClassName with a class your CSI driver provides
kubectl create -f - <<EOF
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshot
metadata:
  name: my-snapshot
  namespace: default
spec:
  volumeSnapshotClassName: <your-csi-snapshot-class>
  source:
    persistentVolumeClaimName: my-pvc
EOF

# Restore from snapshot — storageClassName must match a class that exists on the cluster
kubectl create -f - <<EOF
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: restored-pvc
spec:
  storageClassName: <your-storage-class>
  dataSource:
    name: my-snapshot
    kind: VolumeSnapshot
    apiGroup: snapshot.storage.k8s.io
  accessModes:

    - ReadWriteOnce
  resources:
    requests:
      storage: 10Gi
EOF