Platform Upgrades¶
Procedures for upgrading kMetal platform components and the under cluster.
Upgrade Strategy¶
kMetal upgrades happen at three layers, decoupled from each other:
- Under cluster Kubernetes version — the machines kMetal itself runs on. Upgraded declaratively by
yaki-operator, which kMetal installs: you create aKubernetesNodeUpgradeand it plans, sequences and verifies the rollout. See Node Upgrades. - kMetal platform components — managed by the
kmetalHelm umbrella chart. A chart-version bump pins the tested combination of every sub-chart (Kamaji, KubeVirt, Kube-OVN, MetalLB, cert-manager, CAPI providers, …) — you upgrade them together as one atomic chart release. - Tenant cluster versions — a field on each
Cluster, upgraded per tenant and independent of both layers above (see User Guide: Upgrade Clusters).
Nothing couples them, so the order is yours to choose — with one constraint: a kMetal release supports a range of under cluster Kubernetes versions, and the chart's kubeVersion constraint is what states it.
Check the target against that before moving either of the first two layers.
Under Cluster Upgrade¶
Pre-upgrade checks¶
kubectl get nodes
kubectl get pods -A | grep -v Running | grep -v Completed
# Take an etcd snapshot first
kubectl -n kube-system exec -it etcd-<node> -- etcdctl snapshot save /tmp/etcd-backup.db
Kubernetes version upgrade¶
Do not drain nodes and run kubeadm by hand.
kMetal installs yaki-operator precisely so the control plane moves first, one node at a time, and the workers follow within the version skew policy — with the plan, the outcome and the per-node detail readable from the API afterwards:
apiVersion: yaki.clastix.io/v1alpha1
kind: KubernetesNodeUpgrade
spec:
version: v1.35.8
target:
role: ControlPlane
The full procedure — the seed node, worker concurrency and its 60% ceiling, why nothing is drained, what an upgrade does not touch, and how to recover from a failed one — is in Node Upgrades.
Platform Component Upgrades¶
The umbrella chart pins each component version. To bump components, update the chart release (which pins the tested combination of all sub-chart versions):
helm upgrade kmetal-operator oci://ghcr.io/clastix/charts/kmetal-operator \
--version <new-chart-version> \
--namespace kmetal-operator-system \
--values values.yaml \
--wait --timeout=20m
helm history kmetal-operator -n kmetal-operator-system
Before any production upgrade, diff the rendered output:
helm template kmetal-operator oci://ghcr.io/clastix/charts/kmetal-operator --version <new> \
--values values.yaml > rendered-new.yaml
helm get manifest kmetal-operator -n kmetal-operator-system > rendered-current.yaml
diff rendered-current.yaml rendered-new.yaml
If the upgrade leaves bad state, roll back:
t.b.d. — Per-component independent upgrade paths (overriding a single sub-chart version while keeping the rest pinned) are t.b.d. in this section.
Infrastructure Provider Upgrades¶
t.b.d.
Tenant Cluster Upgrades¶
User Guide Content
Tenant cluster upgrade procedures (control plane version bump, worker node rollout, verification) are documented in the User Guide: Upgrade Clusters.
This section covers only platform-level upgrades.
Upgrade Verification¶
Post-upgrade checks¶
# Cluster version
kubectl version
kubectl get nodes -o wide
# Control-plane health (componentstatuses is deprecated)
kubectl get --raw='/readyz?verbose'
# Chart release
helm status kmetal-operator -n kmetal-operator-system
helm history kmetal-operator -n kmetal-operator-system
# Component pods healthy
kubectl get pods -A | grep -v Running | grep -v Completed
# Tenant control planes still healthy
kubectl get tenantcontrolplanes -A
Rollback¶
# Roll back the chart release
helm rollback kmetal-operator <previous-revision> -n kmetal-operator-system
# If the under-cluster Kubernetes upgrade failed, restore from etcd snapshot
kubectl -n kube-system exec -it etcd-<node> -- \
etcdctl snapshot restore /tmp/etcd-backup.db
t.b.d. — Full disaster-recovery flow after a botched upgrade is t.b.d. in this section; see Disaster Recovery for the broader procedure.