Cluster Backup & Restore¶
A backup of a cluster is an object in your Environment, next to the cluster it protects. You create it, you own it, and it reaches nothing outside your own namespaces.
This is the mechanism for losing a cluster. For losing the data inside a volume, you want Volume Snapshots instead — the two are complementary, and the comparison is worth two minutes before you rely on either.
Everything on this page runs with your Environment credentials, not your cluster's kubeconfig.
What a backup is, and is not¶
A ClusterBackup captures your cluster's API objects: Deployments, ConfigMaps, Secrets, Services, Ingresses, CRDs and the custom resources they serve.
The controller fetches your cluster's kubeconfig, enumerates those objects, archives them, encrypts the archive, and uploads it to S3.
It captures objects, not bytes. That distinction decides what actually survives:
| Captured | Everything the API server holds, minus what a controller will recreate — ReplicaSets, ControllerRevisions and similar are skipped because their owners rebuild them. |
| Not captured | Certificates and CA material. A restore produces a cluster with a fresh identity, by design. |
| Not captured | Node objects. Your nodes are virtual machines the platform rebuilds from an image in minutes, so they are replaced rather than restored — a restored cluster brings up its own. |
| Not captured | The contents of kube-system, kube-public, kube-node-lease, and anything you label backup.kmetal.io/ignore: "true". |
| Conditional | PersistentVolumes — see below. |
Volumes need Retain¶
A PV only round-trips if its backing volume is guaranteed to outlive the cluster, which means a reclaim policy of Retain.
- A PV with
persistentVolumeReclaimPolicy: Retainis captured, with itsclaimRefcleared so it can be re-bound. - A PV with
Delete— the common default — is not. Its volume dies with the cluster. - The PVC is captured either way. If its PV was not captured,
spec.volumeNameis cleared, and the restored PVC provisions fresh storage instead.
Set Retain on anything whose data must survive, before you need the backup:
Nothing flips this for you, and it cannot be applied retroactively to a volume that has already been deleted.
That patch runs inside your cluster, against the PersistentVolume rather than the claim.
If you want Retain to be the default for a class of workload instead of something you remember, ask your platform team for a second StorageClass — see Persistent Storage.
Before your first backup¶
Two Secrets in the same Environment as the ClusterBackup:
S3 credentials — accessKeyId and secretAccessKey. The bucket must already exist; the controller does not create it.
An encryption key — exactly 32 raw bytes, under the key key:
kubectl create secret generic app-backup-encryption-key \
--namespace dynamo-prod \
--from-literal=key="$(openssl rand -base64 24)"
Encrypt it
Encryption is optional in the API and effectively mandatory in practice: the archive contains the full contents of every Secret in your cluster, and it is being written to object storage.
Unencrypted is for a local development loop and nothing else.
Verify the key length with printf '%s' "$KEY" | wc -c → 32. openssl rand -base64 24 produces exactly that.
Taking a backup¶
apiVersion: backup.kmetal.io/v1alpha1
kind: ClusterBackup
metadata:
name: app
namespace: dynamo-prod
spec:
clusterRef:
name: app # the Cluster in this namespace
backend:
type: S3
s3:
bucket: dynamo-backups
region: us-east-1
# endpoint: http://minio.storage.svc:80 # S3-compatible only; omit for AWS S3
credentialsSecretRef:
name: app-backup-s3-credentials
encryption:
secretKeyRef:
name: app-backup-encryption-key
key: key
delete: true # also remove the archive when this object is deleted
kubectl apply -f clusterbackup.yaml
kubectl get clusterbackup app -n dynamo-prod -o jsonpath='{.status.conditions}'
Wait for the Ready condition.
A backup is one-shot, and failure is terminal
There is no retry. If a backup fails, the Ready condition says why; delete the object and create it again.
Retrying is always safe — a backup that never completed never uploaded anything.
For scheduled backups, drive the creation of these objects from whatever scheduler you already run.
Restoring¶
Restore is not a separate object. You create a new cluster and tell it which backup to replay:
apiVersion: cluster.x-k8s.io/v1beta2
kind: Cluster
metadata:
name: app-restored
namespace: dynamo-prod
annotations:
backup.kmetal.io/restore-from: app # a ClusterBackup in this namespace
spec:
# ... your normal Cluster spec: ClusterClass, version, subnet, workers
kubectl apply -f restored-cluster.yaml
kubectl get cluster app-restored -n dynamo-prod -o jsonpath='{.metadata.annotations}'
What happens, in order:
backup.kmetal.io/restore-acceptedappears once the freshness check passes.- The cluster provisions normally — new CA, new identity, new certificates.
- Once its control plane is available, the archive is decrypted and applied: Namespaces, then PersistentVolumes, then CRDs, then everything else.
backup.kmetal.io/restore-completedis set, which is your signal that GitOps can adopt the cluster.
The annotation must be on the manifest you first apply
The restore is only accepted if the Cluster was not yet provisioned when the annotation appeared.
Annotating an existing or running cluster is rejected outright, with no retry.
This is what stops a stray annotation from replaying an archive over a cluster that is serving traffic.
Once applying begins, failure is also terminal — backup.kmetal.io/restore-rejected is set and nothing is retried.
A restore is refused if the new cluster's Kubernetes version is more than three minor versions ahead of the source.
Volume snapshots are the other half¶
A ClusterBackup captures objects, not bytes.
The data inside your volumes is protected by a different mechanism, in your own cluster, with its own page: Volume Snapshots.
The short version of how they fit together:
Retainon aPersistentVolumeis what carries a volume through a restore into the new cluster.- A
VolumeSnapshotis what recovers the contents of a volume that is still there.
Neither replaces the other, and a restore plan needs both.
Disaster recovery¶
Losing a cluster entirely is the case the two mechanisms are designed for together:
- Take
ClusterBackups on a schedule, and keepRetainon the volumes that matter. - To recover, create a new
Clusterwith therestore-fromannotation. - Re-attach or restore volume data —
Retainvolumes are re-bound by the restore; anything else comes back from aVolumeSnapshot, or not at all.
The result is a new cluster with your workloads, not a resurrection of the old one: new certificates, new identity, and a new API endpoint for your users and CI to be pointed at.
If the platform itself is what you lost rather than your cluster, that is your platform team's Disaster Recovery procedure, not this one.
See Also: Volume Snapshots · Persistent Storage · Delete Clusters