Skip to content

Data Protection Tasks

Setting a tenant up so they can back their own clusters up.

Backups are tenant-driven: a ClusterBackup is an object in the tenant's Environment, created by the tenant (or on their behalf), protecting a cluster they own. The admin side is making that possible — a reachable destination, the credentials to use it, and the permission to create the object at all: backup.kmetal.io is in no role by default, so a tenant who owns their own backups needs kmetal-tenant-data-protector on their Tenant.

For what a backup does and does not capture, see Cluster Backup & Restore.

A restore builds a new cluster, it never repairs one

Nothing in kMetal puts a cluster back the way it was. A tenant creates a new Cluster carrying backup.kmetal.io/restore-from: <ClusterBackup> on the manifest they first apply, and the archive is replayed into it once its control plane is up. A Cluster that was already Provisioned when the annotation appeared is rejected outright, so an archive can never land on a cluster serving traffic.

That is immutable infrastructure, and it is why the archive holds as little as it does:

Layer In the archive Because
API objects yes The only part of a cluster a fresh one cannot produce by itself.
Node objects never — the resource is excluded unconditionally Worker nodes are KubeVirt VMs booted from a golden image and joined by CAPI. A replacement is minutes away and identical, so nodes are cattle; a restored Node would only fight the real one.
Certificates and CA material no A restored cluster gets a fresh identity by design — which also means a new endpoint for the tenant to point clients at.
CNI, CSI and the other platform add-ons no The platform redelivers them to any cluster it provisions.
Volumes conditionally The bytes never lived in the tenant cluster — see below.

What that model asks of you is that the replacement cluster is buildable on the day a tenant needs it:

  • the tier the archived cluster used is still enabled, and its ClusterClass still exists,
  • a Kubernetes version inside the restore window is still offered, with a worker image built for it — the target must be the same major version, no older than the source, and at most three minor versions ahead,
  • the tenant still has quota, a Ready SubnetClaim and a free control-plane address.

None of the three is checked when the backup is taken, and an archive whose cluster cannot be rebuilt is not a backup. See Restoring for the tenant-side steps.

Provide a destination

Backups are written to S3, or to anything S3-compatible reachable from the under cluster. The bucket must already exist — nothing in kMetal creates one.

# Example against an S3-compatible endpoint
s3cmd mb --host=<endpoint> --host-bucket= s3://dynamo-backups

Decide per tenant or per platform:

Shape Trade-off
A bucket per tenant Clean blast radius and per-tenant lifecycle rules. More to provision at onboarding.
One bucket, prefix per tenant Less to provision. Archives are separated by path, not by credential, so a leaked credential reaches more.

Archives are laid out <tenantUID>-<tenantName>/<namespaceUID>-<namespaceName>/<clusterUID>-<clusterName>/<backupUID>-<backupName>/, so a shared bucket is navigable either way.

Give the tenant credentials

Two Secrets, in the same Environment as the ClusterBackup — so the tenant can read them and nobody else can:

# S3 credentials
kubectl create secret generic app-backup-s3-credentials \
  --namespace dynamo-prod \
  --from-literal=accessKeyId=<key> \
  --from-literal=secretAccessKey=<secret>

# Encryption key: exactly 32 raw bytes
kubectl create secret generic app-backup-encryption-key \
  --namespace dynamo-prod \
  --from-literal=key="$(openssl rand -base64 24)"

Insist on encryption

The archive contains the full contents of every Secret in the tenant's cluster, written to object storage. An unencrypted backup is a copy of the tenant's secrets in a bucket.

Encryption is optional in the API, so it is a platform policy question rather than a technical constraint. Provision the key alongside the credentials and the tenant has no reason to skip it.

Scope the S3 credential to the tenant's bucket or prefix. A credential that can read the whole bucket lets one tenant read another's archives, which is the one thing this arrangement is supposed to prevent.

Account for backup storage

Archives do not count against the tenant's tenantClassName quota — they live in object storage, outside Kubernetes. Bound them where the object store bounds them: bucket quotas, lifecycle rules, or per-prefix limits.

spec.delete: true on a ClusterBackup (the default) removes the archive when the object is deleted, so tenants pruning their own backups reclaims space. Nothing enforces retention — old ClusterBackup objects hold their archives indefinitely.

Volumes move with a restore, if the policy allows it

A backup captures objects, not bytes — and a restore can still hand the new cluster the same volumes the old one had.

The data was never inside the tenant cluster. Every tenant PersistentVolume is backed by a DataVolume in the under cluster, provisioned by the KubeVirt CSI driver on storage.tenantClassName, in the same namespace as the Cluster. That object holds no owner reference to the cluster that asked for it, so it outlives the tenant cluster and a restore only has to re-point it.

Whether that happens is decided by the PV's reclaim policy, at backup time:

Reclaim policy In the archive At restore
Retain The PV is captured, with claimRef cleared so it is Available again. The PV is applied before everything else, the restored PVC re-binds to it by name, and the backing DataVolume is relabelled csi.kubevirt.io/cluster: <new cluster> so the new cluster's CSI controller owns it.
Delete (the default) The PV is skipped; the PVC is captured with spec.volumeName cleared. The PVC provisions fresh, empty storage. The old volume went away with the cluster that owned it.

Two things to know before a tenant needs them:

  • The relabelling resolves the DataVolume in the target Cluster's own namespace, so restoring into the same Environment brings the volumes and restoring into a different one brings only the objects.
  • The default StorageClass kMetal projects into tenant clusters is Delete, and a StorageClass's reclaim policy is immutable, so offering Retain means an additional class in the tenant's cluster rather than an edit to that one. Handing tenants that manifest at onboarding costs less than explaining it at restore time.

A volume already in flight can still be saved, because the policy is mutable on the PV itself:

kubectl patch pv <pv-name> -p '{"spec":{"persistentVolumeReclaimPolicy":"Retain"}}'

That runs inside the tenant cluster, and it has to happen before teardown — nothing flips it for you, and it cannot be applied to a volume that is already gone.

Protecting the bytes is not kMetal's job

kMetal installs no storage backend; it consumes the class storage.tenantClassName points at. Protecting what is inside those volumes — replication, snapshots to another site, a backup product pointed at the storage system — belongs to that layer, and none of it is driven by a ClusterBackup.

Inside their own cluster a tenant does have VolumeSnapshot, and it works without anything from you: kubevirt-csi-driver-operator projects a kubevirt-csi-snapclass VolumeSnapshotClass into every tenant cluster alongside the node bundle. That is a different thing from the under cluster, where kMetal installs the snapshot controller and the CRDs and no class — the class there is your backend's, named in spec.storage.

Tenant snapshots land on storage.tenantClassName and count against the tenant's storage quota, so a tenant retaining snapshots indefinitely consumes the same budget as one holding volumes. Nothing prunes them. See Volume Snapshots and Storage.

Verify a tenant can back up

kubectl get clusterbackups -A
kubectl get clusterbackup <name> -n <environment> -o jsonpath='{.status.conditions}'

A backup is one-shot with no retry: any failure is terminal and the condition says why. The common causes are a bucket that does not exist, a credential that cannot write to it, and an encryption key that is not exactly 32 bytes.