Storage Configuration¶
kMetal installs no storage: no CSI driver, no provisioner, no StorageClass, no VolumeSnapshotClass.
You bring storage the way you bring the primary CNI — it exists before kMetal is installed — and name it in the KMetal spec.
This page is about what the platform requires of that storage, and what you configure once it exists. For the model, see Storage Concepts. For the tenant-facing wiring, see Tenant Storage (CSI).
What the platform requires¶
| Requirement | Why |
|---|---|
| Dynamic provisioning | Every tenant volume and every tenant worker disk is provisioned on demand. |
A VolumeSnapshotClass on the platform class's driver |
The image builder snapshots the disk it installed the OS onto, and CDI clones every machine's root disk from that snapshot. A class that cannot snapshot leaves the artifact with no golden volume and machines with no disk to boot; a clone the backend cannot take becomes a full image copy per machine. |
allowVolumeExpansion |
Tenants resize volumes; the CSI layer passes the request through. |
The snapshot requirement moved
It used to sit on tenantClassName, because that is where the golden images were built.
They are on underclusterClassName now, together with the machine root disks cloned from them — so a node-local platform class, which used to be a defensible choice, no longer boots a single tenant machine.
That is what etcdClassName below exists to give back.
Anything meeting those can back kMetal — a vendor array, Rook-Ceph, Longhorn, an NFS CSI driver in a lab. The examples further down this page are illustrative starting points, not kMetal-specific configuration.
Naming the classes¶
spec:
storage:
underclusterClassName: platform-storage # golden images + machine root disks
etcdClassName: local-path # optional; defaults to the class above
tenantClassName: tenant-storage-class # handed to tenant workloads
Three fields, because the consumers pull in different directions. Pointing everything at one class silently gets one of them wrong.
| Field | Required | Backs | Wants |
|---|---|---|---|
underclusterClassName |
yes | The golden OS images in kmetal-image-builder, and the root disk of every tenant machine cloned from them. |
Snapshots. Capacity for one volume per machine plus one golden volume per image. |
etcdClassName |
no | The Kamaji datastore holding every hosted control plane's etcd data. Defaults to underclusterClassName. |
Low latency, small. Node-local is defensible, but ties the durability of every tenant control plane to one node. |
tenantClassName |
yes | The volumes tenant workloads claim, through the KubeVirt CSI driver. | Network-attached, surviving the loss of any single node. |
A node-local class under underclusterClassName or tenantClassName makes every tenant worker node a machine that dies permanently with its host — which tenants experience as their cluster losing nodes.
Under etcdClassName it is a deliberate trade instead: latency and simplicity, paid for with the etcd data of every hosted control plane if the node holding it is lost.
Set etcdClassName only where one class cannot serve both platform jobs.
On a cluster whose single class snapshots and is fast enough for etcd, leave it unset and say so once in a comment.
All three are hard to change once tenants exist, because their data is on them. Decide before the first tenant — see Configuration Best Practices.
Keep the platform class and the tenant class apart¶
Give them two StorageClass objects even where the same driver and the same pool back both.
The reason is quota, not performance.
A tenant's storage quota is a ResourceQuota keyed on <class>.storageclass.storage.k8s.io/requests.storage, so it counts by class name rather than by who the volume belongs to.
Machine root disks are provisioned in the tenant's own namespace on underclusterClassName, which means one shared name has the platform's own volumes billed to the tenant:
$ kubectl get resourcequota -n dynamo-dev -o yaml | grep -A2 'hard:'
hard:
tenant-storage-class.storageclass.storage.k8s.io/requests.storage: 59Gi
used:
tenant-storage-class.storageclass.storage.k8s.io/requests.storage: "33285996544"
31 GiB used on a tenant running one 1 GiB workload volume: the other 30 GiB is a single worker's boot disk. Two workers on a 30Gi tier consume 60 GiB of the tenant's budget before a workload claims anything, and every cluster they create moves the number again — so the quota stops meaning what it says, and raising it stops being arithmetic anyone can do.
Two classes over one backend costs nothing but a second StorageClass object:
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: platform-storage # underclusterClassName
provisioner: rook-ceph.rbd.csi.ceph.com
parameters:
clusterID: rook-ceph
pool: replicapool # the same pool as the tenant class, deliberately
# ... the rest identical to the tenant class
allowVolumeExpansion: true
Machine disks then count against a key nobody's quota names, and the tenant's budget is the tenant's workloads. Bound the machine side where it belongs instead — on the cluster and machine counts the tenant may ask for, which is what actually drives it. See Cap what a tenant can consume.
One VolumeSnapshotClass still has to match the platform class
Whatever you name here, clusterClass.values.dataVolume.volumeSnapshotClass must be served by the same driver as underclusterClassName — that is the snapshot the golden volume is taken with.
Same driver is the requirement, not the same class name.
Pinning CDI's capabilities¶
CDI provisions the volumes tenant machines boot from, and it decides how by reading the storage's advertised capabilities. On a backend advertising more than one combination, its derived order can be wrong for kMetal:
The operator turns this into a CDI StorageProfile for tenantClassName.
Order is significant, and the failure is opaque
CDI takes the first entry for a DataVolume that asks for nothing specific.
A backend that advertises ReadWriteMany + Block first has CDI provision raw-block volumes, and the importer pod then fails on a permission error that names neither the storage class nor the volume mode.
Leave the list empty on a single-capability backend, where auto-derivation can only produce one answer. Set it on anything else, before the first tenant volume.
CDI's filesystem overhead is pinned to 0 on the tenant class by the platform, and is not a field.
The default 6% reservation pushes a whole-GiB DataVolume just past the next GiB boundary, and a driver that rounds up from there — Ceph RBD does — provisions twice what the tenant asked for.
Only that one class is touched, so an overhead set for any other class is left alone.
Both pins land on the tenant class only
tenantClaimPropertySets renders a StorageProfile for tenantClassName, and the overhead pin names tenantClassName.
Neither reaches underclusterClassName — which is where the golden images and every machine root disk now live.
On a platform class distinct from the tenant class, those volumes therefore take CDI's auto-derived capabilities and its default 6% overhead.
On a whole-GiB boot tier that is the same rounding problem, one size up: a 30Gi tier asks the backend for ~31.8 GiB, and a driver rounding to the next power-of-two-friendly size provisions 32 GiB per machine.
Set the overhead for that class yourself if it bites, on the CDI object — kMetal leaves other classes alone precisely so that edit survives:
kubectl patch cdi cdi --type=merge \
-p '{"spec":{"config":{"filesystemOverhead":{"storageClass":{"platform-storage":"0"}}}}}'
Read what CDI derived for the class before assuming its defaults are wrong — cloneStrategy is the line that decides whether machine creation is a snapshot or a copy:
Swapping the backend¶
Because the classes are named in the spec rather than created by kMetal, changing backend is a storage operation followed by a one-line spec change:
- Install the new CSI driver on the under cluster, with its
StorageClassandVolumeSnapshotClass. - Verify it provisions and snapshots — a driver without a working
VolumeSnapshotClasscannot boot tenant machines. - Point the class you are moving at the new class:
storage.tenantClassNamefor tenant workloads,storage.underclusterClassNamefor the images and machine disks. - Moving
underclusterClassNamealso meansclusterClass.values.dataVolume.volumeSnapshotClasshas to move to a class on the new driver, in the same change, or golden snapshots stop becoming ready. - Revisit
tenantClaimPropertySetsfor the new backend's capabilities.
Existing volumes do not migrate. Data already provisioned stays on the old backend, bound to its PVs, until you move it deliberately — plan a swap as a migration, not a cutover.
Backend examples¶
None of the following is kMetal configuration — it is a starting point for standing up a backend that meets the requirements above.
NFS¶
One option among many
kMetal bundles no NFS provisioner. The example below uses nfs-subdir-external-provisioner as a common pick; substitute whatever NFS CSI driver your environment runs.
Note that most NFS provisioners ship no VolumeSnapshotClass, which makes them unsuitable for tenantClassName.
# nfs-provisioner.yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: nfs-storage
provisioner: nfs.csi.k8s.io
parameters:
server: nfs-server.company.com
share: /exports/kubernetes
mountPermissions: "0755"
reclaimPolicy: Retain
volumeBindingMode: Immediate
mountOptions:
- hard
- nfsvers=4.1
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: nfs-provisioner
namespace: kube-system
spec:
replicas: 2
selector:
matchLabels:
app: nfs-provisioner
template:
metadata:
labels:
app: nfs-provisioner
spec:
serviceAccountName: nfs-provisioner
containers:
- name: nfs-provisioner
image: registry.k8s.io/sig-storage/nfs-subdir-external-provisioner:v4.0.2
volumeMounts:
- name: nfs-client-root
mountPath: /persistentvolumes
env:
- name: PROVISIONER_NAME
value: nfs.csi.k8s.io
- name: NFS_SERVER
value: nfs-server.company.com
- name: NFS_PATH
value: /exports/kubernetes
volumes:
- name: nfs-client-root
nfs:
server: nfs-server.company.com
path: /exports/kubernetes
Tenant Cluster Storage¶
Tenants get persistent storage in their cluster through the KubeVirt CSI driver, which maps onto the class named by storage.tenantClassName.
The credentials for the backend stay on the under cluster — nothing inside a tenant cluster holds a key to it.
See Tenant Storage (CSI) for how that is wired, and how the per-tenant cap is enforced on both sides.
Per-tenant storage quota¶
Cap each tenant's total storage with a ResourceQuota in the tenant's under-cluster namespace, keyed on tenant-storage-class.storageclass.storage.k8s.io/requests.storage:
apiVersion: v1
kind: ResourceQuota
metadata:
name: alpha-storage
namespace: alpha
spec:
hard:
# <tenantClassName>.storageclass.storage.k8s.io/requests.storage
tenant-storage-class.storageclass.storage.k8s.io/requests.storage: 10Gi
The key is derived from storage.tenantClassName, so it changes if you change that.
It counts every volume in the namespace on that class — including machine root disks, if underclusterClassName names the same class.
See Keep the platform class and the tenant class apart.
The quota lives here, on the under cluster — but tenants work inside their own cluster, where storage simply appears and nothing natively knows a quota exists.
kMetal closes that gap with the kmetal-webhook component: one process on the under cluster serving every tenant cluster, which installs itself into each as an admission webhook.
A PVC that would take a tenant past this quota is rejected when it is created, with a message saying so — rather than accepted and left Pending while the tenant tries to work out whether the platform is broken.
# See what tenants are consuming
kubectl get resourcequota -A -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,USED:.status.used,HARD:.status.hard \
| grep -i storage
# Raise a tenant's cap
kubectl patch resourcequota alpha-storage -n alpha --type=merge \
-p '{"spec":{"hard":{"tenant-storage-class.storageclass.storage.k8s.io/requests.storage":"50Gi"}}}'
CDI on the under cluster¶
KubeVirt's Containerized Data Importer (CDI) handles golden-image imports and the DataVolume lifecycle on the under cluster.
kMetal installs it as part of the platform, so there is nothing to deploy.
It matters here because every tenant volume ultimately becomes a CDI-managed DataVolume on the tenant class — which is why pinning CDI's capabilities is the one piece of CDI configuration worth getting right before the first tenant volume exists.
Distributed Storage¶
Your choice entirely
kMetal ships no storage at all, so what backs it is your pick — Rook-Ceph, Longhorn, a vendor CSI driver, or anything else compatible with the under cluster's Kubernetes version and shipping a VolumeSnapshotClass. The examples below are illustrative starting points, not kMetal-specific configuration.
Rook-Ceph¶
# rook-ceph-cluster.yaml
apiVersion: ceph.rook.io/v1
kind: CephCluster
metadata:
name: rook-ceph
namespace: rook-ceph
spec:
cephVersion:
image: quay.io/ceph/ceph:v18.2.0
dataDirHostPath: /var/lib/rook
mon:
count: 3
allowMultiplePerNode: false
mgr:
count: 2
allowMultiplePerNode: false
dashboard:
enabled: true
ssl: true
storage:
useAllNodes: true
useAllDevices: false
deviceFilter: "^sd[b-z]"
config:
osdsPerDevice: "1"
---
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: rook-ceph-block
provisioner: rook-ceph.rbd.csi.ceph.com
parameters:
clusterID: rook-ceph
pool: replicapool
imageFormat: "2"
imageFeatures: layering
csi.storage.k8s.io/provisioner-secret-name: rook-csi-rbd-provisioner
csi.storage.k8s.io/provisioner-secret-namespace: rook-ceph
csi.storage.k8s.io/node-stage-secret-name: rook-csi-rbd-node
csi.storage.k8s.io/node-stage-secret-namespace: rook-ceph
csi.storage.k8s.io/fstype: ext4
allowVolumeExpansion: true
reclaimPolicy: Delete
Longhorn¶
# longhorn-storage.yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: longhorn
annotations:
storageclass.kubernetes.io/is-default-class: "true"
provisioner: driver.longhorn.io
allowVolumeExpansion: true
reclaimPolicy: Delete
volumeBindingMode: Immediate
parameters:
numberOfReplicas: "3"
staleReplicaTimeout: "2880"
fromBackup: ""
fsType: "ext4"
dataLocality: "disabled"
Object Storage¶
MinIO¶
# minio-deployment.yaml
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: minio
namespace: minio-system
spec:
serviceName: minio
replicas: 4
selector:
matchLabels:
app: minio
template:
metadata:
labels:
app: minio
spec:
containers:
- name: minio
image: quay.io/minio/minio:latest
args:
- server
- --console-address
- ":9001"
- http://minio-{0...3}.minio.minio-system.svc.cluster.local/data
env:
- name: MINIO_ROOT_USER
valueFrom:
secretKeyRef:
name: minio-credentials
key: rootUser
- name: MINIO_ROOT_PASSWORD
valueFrom:
secretKeyRef:
name: minio-credentials
key: rootPassword
ports:
- containerPort: 9000
name: api
- containerPort: 9001
name: console
volumeMounts:
- name: data
mountPath: /data
resources:
requests:
cpu: "500m"
memory: "1Gi"
volumeClaimTemplates:
- metadata:
name: data
spec:
accessModes: [ "ReadWriteOnce" ]
storageClassName: <your-storage-class>
resources:
requests:
storage: 100Gi
Backup Storage¶
Operator's choice
kMetal does not bundle a backup tool. The examples below show Velero — a common operator choice — configured against S3-compatible object storage. Adapt to whatever backup tooling your environment uses.
Velero with S3¶
# velero-backup-location.yaml
apiVersion: velero.io/v1
kind: BackupStorageLocation
metadata:
name: default
namespace: velero
spec:
provider: aws
objectStorage:
bucket: kmetal-backups
prefix: velero
config:
region: us-west-2
s3ForcePathStyle: "false"
s3Url: https://s3.us-west-2.amazonaws.com
---
apiVersion: velero.io/v1
kind: VolumeSnapshotLocation
metadata:
name: default
namespace: velero
spec:
provider: aws
config:
region: us-west-2
Velero with MinIO¶
# velero-minio-backup.yaml
apiVersion: velero.io/v1
kind: BackupStorageLocation
metadata:
name: minio
namespace: velero
spec:
provider: aws
objectStorage:
bucket: velero
config:
region: minio
s3ForcePathStyle: "true"
s3Url: http://minio.minio-system:9000
publicUrl: https://minio.company.com
Storage Performance¶
Benchmark Storage¶
# Deploy fio for benchmarking
kubectl run fio-benchmark --rm -it --image=ljishen/fio -- /bin/bash
# Run sequential write test
fio --name=seqwrite --rw=write --bs=1M --size=1G --numjobs=1 --runtime=60 --time_based --filename=/data/test
# Run random write test
fio --name=randwrite --rw=randwrite --bs=4K --size=1G --numjobs=4 --runtime=60 --time_based --filename=/data/test
# Run read test
fio --name=read --rw=read --bs=1M --size=1G --numjobs=1 --runtime=60 --time_based --filename=/data/test
Monitor Storage Usage¶
# Check PV usage
kubectl get pv
# Check PVC usage
kubectl get pvc -A
# Storage capacity
kubectl top pv
# Per-node storage usage
kubectl get --raw /api/v1/nodes/<node-name>/proxy/stats/summary | \
jq '.node.fs'
Storage Troubleshooting¶
Debug PV/PVC Issues¶
# Check PVC status
kubectl get pvc -A
kubectl describe pvc <pvc-name> -n <namespace>
# Check PV binding
kubectl get pv
kubectl describe pv <pv-name>
# Check storage class
kubectl get storageclass
kubectl describe storageclass <storage-class-name>
# Check provisioner logs
kubectl logs -n kube-system -l app=<provisioner-name> -f
CSI Driver Debugging¶
The controller side is kubevirt-csi-driver-operator — one Deployment on the under cluster serving every tenant cluster, with the CSI controller running in-process rather than as per-tenant pods and sidecars. Only the node DaemonSet runs inside the tenant cluster.
# --- Controller side (under cluster) ---
kubectl get deploy kubevirt-csi-driver-operator -n kmetal-system
kubectl logs -n kmetal-system deploy/kubevirt-csi-driver-operator
# What it provisioned for one cluster, on the under cluster
kubectl get datavolumes -n <environment> -l csi.kubevirt.io/cluster=<cluster>
# --- Node side (tenant cluster) ---
kubectl --kubeconfig=<tenant>.kubeconfig -n kube-system get ds kubevirt-csi-node
kubectl --kubeconfig=<tenant>.kubeconfig -n kube-system logs -l app=kubevirt-csi-node -c csi-driver
# Volume attachments (tenant cluster)
kubectl --kubeconfig=<tenant>.kubeconfig get volumeattachment
For non-kubevirt-csi drivers (NFS, Rook-Ceph, vendor CSI, etc.), the controller usually runs in kube-system or a vendor-specific namespace on the under cluster — adapt the selectors accordingly.
Storage Migration¶
Requires a snapshot-capable CSI driver
Snapshots need a driver that ships a VolumeSnapshotClass — the same requirement kMetal has for tenantClassName. Update volumeSnapshotClassName and storageClassName below to match what your driver provides.
# Create snapshot — replace volumeSnapshotClassName with a class your CSI driver provides
kubectl create -f - <<EOF
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshot
metadata:
name: my-snapshot
namespace: default
spec:
volumeSnapshotClassName: <your-csi-snapshot-class>
source:
persistentVolumeClaimName: my-pvc
EOF
# Restore from snapshot — storageClassName must match a class that exists on the cluster
kubectl create -f - <<EOF
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: restored-pvc
spec:
storageClassName: <your-storage-class>
dataSource:
name: my-snapshot
kind: VolumeSnapshot
apiGroup: snapshot.storage.k8s.io
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 10Gi
EOF