Under Cluster Node Upgrades¶
Moving the under cluster's own Kubernetes version, declaratively, from inside the cluster.
The under cluster is self-managed.
No node-pool API can move its kubelets, and by hand each node is a kubeadm invocation, a kubelet restart and a readiness check — in an order that matters, on machines you are also running every tenant's control plane on.
kMetal installs yaki-operator so that sequence is a resource you create rather than a runbook someone has to remember.
It ships with nothing configurable and sits idle until a KubernetesNodeUpgrade appears.
| Namespace | kmetal-yaki-operator — also where the upgrade Jobs are created |
| Category | NodeLifecycle. A component here being unready means node upgrades are unavailable, not that any tenant is affected |
| API | kubernetesnodeupgrades.yaki.clastix.io, short name knu, cluster-scoped like the nodes it acts on |
| Depends on | cert-manager, which issues the certificate for its defaulting webhook |
$ kubectl get km dev-env -o json | jq -r '.status.components[] | select(.category=="NodeLifecycle")'
{
"category": "NodeLifecycle",
"desiredVersion": "0.6.0",
"installedVersion": "0.6.0",
"name": "yaki-operator",
"phase": "Ready"
}
Bootstrap with YAKI, and the upgrades follow¶
kMetal's recommended way to build an under cluster is YAKI — a single script that runs init on the first control-plane node and join on every other — rather than driving kubeadm yourself.
See Under Cluster Setup.
The reason is not the install.
kubeadm by hand gets you the same cluster.
The reason is that the operator's upgrade path is YAKI's own upgrade flow, fetched at run time and executed in the host namespace of each selected node, so a cluster bootstrapped with YAKI is a cluster whose upgrade is already accounted for.
What that buys you is a shorter list of moving parts for day-2:
| Keeping the under cluster aligned | What it takes |
|---|---|
YAKI + yaki-operator |
One object in the API server of the cluster being upgraded. No inventory, no external control plane. |
| Terraform, Ansible or another configuration-management pass | A second source of truth for the fleet, and a run to trigger from outside the cluster. |
| Metal3 / Ironic, or another declarative bare-metal provisioner | A seed cluster and a hardware-management stack, for the benefit of re-provisioning rather than upgrading. |
Metal3 remains a reasonable choice for a large under cluster where machines are provisioned and re-provisioned often. It is not something you need in order to move a kubelet, which is the case this page covers.
Adopting YAKI later is not the same as bootstrapping with it
YAKI uses no package manager.
It fetches kubeadm, kubelet, kubectl, containerd, runc, crictl and the CNI plugin binaries from their upstream release URLs and places them under /usr/local, with systemd units in /usr/lib/systemd/system and /usr/local/lib/systemd/system.
On a node whose kubelet came from a distribution package, an upgrade adds a second copy of the binaries and leaves the systemd unit and PATH to decide which one actually runs.
Bootstrapping with YAKI avoids that question entirely; adopting it on an apt-managed node is a migration, and should be treated as one.
Upgrade the control plane first¶
apiVersion: yaki.clastix.io/v1alpha1
kind: KubernetesNodeUpgrade
spec:
version: v1.35.8
target:
role: ControlPlane
Control-plane nodes always move one at a time, and the first entry in the plan is the seed — the only node whose Job runs kubeadm upgrade apply, the command that moves the cluster's version.
Every other node, control plane or worker, gets the equivalent of kubeadm upgrade node.
The plan is name-sorted and persisted before any Job is created, so you can read which node will be the seed before anything happens:
$ kubectl get knu
NAME VERSION ROLE PHASE COMPLETED FAILED
kubernetesnodeupgrade-control-plane-bwn85 v1.35.8 ControlPlane Upgrading
$ kubectl get knu kubernetesnodeupgrade-control-plane-bwn85 -o jsonpath='{range .status.plan[*]}{.order}{"\t"}{.name}{"\tseed="}{.seed}{"\t"}{.sourceVersion}{" -> "}{.targetVersion}{"\t"}{.state}{"\n"}{end}'
0 kmetal-controller seed=true v1.35.2 -> v1.35.8 Running
Use kubectl create, not apply: a defaulting webhook renames every resource to kubernetesnodeupgrade-<role>-<random>, so you do not choose the name and kubectl get knu accumulates one row per attempt — the upgrade history of the under cluster, in its own API.
Check the target version against the kMetal release before you create anything. The platform pins a supported Kubernetes range, and the under cluster's version is not something a platform upgrade will fix for you — see Platform Upgrades.
A single-control-plane under cluster loses its API server for the duration
Lab and evaluation under clusters have one control-plane node, and upgrading it takes the under cluster's api-server down while kubeadm runs.
Flux, Kamaji and the Cluster API controllers stop reconciling for that window, and kubectl against the under cluster fails.
Tenants do not notice: hosted control-plane pods and virt-launcher pods are already running, and a kubelet restart does not stop them.
Production under clusters have three control-plane nodes, and the operator moves them one at a time for exactly this reason.
Then the workers¶
apiVersion: yaki.clastix.io/v1alpha1
kind: KubernetesNodeUpgrade
spec:
version: v1.35.8
target:
role: Worker
workerConcurrencyStrategy:
maxUnavailable: 1
A worker upgrade refuses to start unless every control-plane node in the cluster is already Ready at exactly spec.version.
It does not wait for one to catch up: it fails immediately, marks every entry Failed with the reason, and freezes.
That is the version skew policy enforced by construction, and a worker upgrade issued too early is almost always an ordering mistake rather than something to be patient about.
maxUnavailable is how many workers may have an active Job at once — an integer or a percentage, defaulting to 1, and never more than 60% of the selected workers.
On a small under cluster, that ceiling means one at a time
Three workers is the common production shape, and there maxUnavailable: 2 is rejected — two of three is 66%.
A percentage does not get you around it either: 40% of three rounds up to two, crosses the ceiling, and is clamped back to one.
So a three-worker under cluster upgrades serially whatever you ask for, and that is the intended answer rather than a limitation to work around.
The worker hosting the operator's own pod is planned last and held back until the other selected workers are done, so the upgrade cannot evict the controller driving it.
With maxUnavailable: 1 it waits for every other entry to reach a terminal state; with a higher value it waits only on entries actually in flight, so one stuck node cannot starve it.
Where the operator runs decides whether that rule applies
The controller lands wherever spec.placement on the KMetal object puts the platform's components.
Placed on control-plane nodes — a common choice — its pod is never part of a worker plan at all, and the ordering above is simply moot.
The operator does not drain, and on this platform that matters¶
YAKI restarts the kubelet on the host. It never talks to the api-server, the operator never cordons or drains, and neither consults a PodDisruptionBudget. Running pods stay running through the restart.
That is the gentler option here, and it is worth understanding why before adding a drain step out of habit.
An under cluster worker carries tenant hosted control planes and the virt-launcher pods backing tenant worker VMs.
Those VMs are created without an evictionStrategy, so evicting a virt-launcher pod does not live-migrate anything — the VM stops and is restarted elsewhere, which a tenant experiences as one of their nodes rebooting.
| Operation | What a tenant sees |
|---|---|
| Kubelet restart, as the operator does it | Nothing. Control-plane pods and worker VMs keep running. |
kubectl drain of an under cluster worker |
Their worker nodes restart, and their hosted control-plane pods move. |
Emptying a node is therefore a separate, tenant-visible maintenance operation — for a kernel update, a reboot or hardware work — and not part of a version bump.
Keep maxUnavailable low regardless: every entry restarts a kubelet on a machine several tenants are running on.
What an upgrade does not touch¶
The upgrade flow deliberately skips the host preparation that init and join do, and changes as little as it can.
- Only
kubeadm,kubeletandkubectlmove. containerd,runc,crictland the CNI plugin binaries are left where they are, and nothing warns you that the runtime is overdue. - No reboot, so a kernel update is yours to schedule.
- No platform components. kMetal's own versions come from the release you install; see Platform Upgrades.
- No tenant clusters. Their Kubernetes version is a
Clusterfield, upgraded independently — see Upgrade Clusters.
Follow it, and read the history¶
# Phase, and how far it got
kubectl get knu -w
# Per-node detail: order, seed, versions, state
kubectl get knu <name> -o jsonpath='{range .status.plan[*]}{.name}{"\t"}{.state}{"\t"}{.sourceVersion}{" -> "}{.targetVersion}{"\n"}{end}'
# What is running right now, and its log
kubectl get knu <name> -o jsonpath='{range .status.plan[?(@.state=="Running")]}{.name}{"\t"}{.jobName}{"\n"}{end}'
kubectl logs -n kmetal-yaki-operator job/<jobName> -f
A two-step upgrade of a four-node under cluster, mid-flight, reads like this:
$ kubectl get knu
NAME VERSION ROLE PHASE COMPLETED FAILED
kubernetesnodeupgrade-control-plane-bwn85 v1.35.8 ControlPlane Succeeded 1
kubernetesnodeupgrade-worker-v59fx v1.35.8 Worker Upgrading 1
$ kubectl get knu kubernetesnodeupgrade-worker-v59fx \
-o jsonpath='{range .status.plan[*]}{.order}{" "}{.name}{" "}{.state}{" "}{.sourceVersion}{" -> "}{.targetVersion}{"\n"}{end}'
0 kmetal-worker-01 Succeeded v1.35.2 -> v1.35.8
1 kmetal-worker-02 Running v1.35.2 -> v1.35.8
2 kmetal-worker-03 Pending v1.35.2 -> v1.35.8
The control-plane resource is finished and will never change again; the worker resource is walking its plan one entry at a time, in the order it committed to before the first Job existed.
A resource that reaches Succeeded or Failed is frozen: never reconciled again, and immutable.
Editing one does nothing, which is what makes kubectl get knu a log rather than a set of live objects.
Each attempt is a new resource.
When one fails¶
Failed means at least one selected node did not reach the target version — the phase never grades a partial rollout as a qualified success.
status.plan says which node, and status.completed / status.failed say how far it got.
An upgrade Job is never retried automatically, and a failed Job's retention is raised to five minutes so its logs outlive the upgrade that created it:
The reason is almost always in YAKI's output, because the operator's guarantees are about selection, ordering and verification — the upgrade itself is still kubeadm, running as root on your host.
Fix the host, then create a new resource.
Two escape hatches exist per resource, and neither relaxes the ordering rules:
| Field | Effect |
|---|---|
spec.force.allowDowngrade |
Permits a target below a node's current version, still within one minor. |
spec.force.skipVersionCheck |
Waives the major and multi-minor span checks, and re-runs nodes already at the target instead of skipping them. |
For a canary, narrow the target instead of forcing anything: spec.target.nodeNames or spec.target.nodeSelector move one node or one labelled pool, and a target that matches nothing fails rather than quietly completing as a no-op.
Who may create one¶
Creating a KubernetesNodeUpgrade runs privileged code as root on the selected nodes.
The Jobs use hostPID, hostNetwork, a hostPath mount of / and the SYS_CHROOT capability, because that is what upgrading a kubelet from inside the cluster costs.
Treat it as a platform-administrator capability:
- It is in no tenant-facing role, and must stay that way.
yaki.clastix.iois not part of any ClusterRole a CapsuleTenantbinds — see What a tenant cannot see at all. - Restrict
createonkubernetesnodeupgradesto the same subjects that hold cluster-admin on the under cluster.get/listis a reasonable thing to hand out more widely, since the resources are the upgrade history. - Pod Security Admission has to allow it. kMetal creates
kmetal-yaki-operatorwith no PSA labels, so a cluster-wide default that enforcesbaselineorrestrictedwill block the upgrade Jobs. Label the namespacepod-security.kubernetes.io/enforce=privilegedif you enforce a default.
Air-gapped under clusters¶
Each Job downloads the YAKI script from https://goyaki.clastix.io at run time, and the script then fetches the Kubernetes binaries from their upstream release URLs.
Three things therefore need to be reachable from the nodes, not from your workstation: that script, the upstream binaries, and the digest-pinned Alpine image the Job runs.
Mirror all three, or node upgrades are the one operation an otherwise air-gapped platform cannot perform.
Verify¶
# The operator, and what kMetal thinks of it
kubectl get deploy yaki-operator -n kmetal-yaki-operator
kubectl get km -o json | jq -r '.items[].status.components[] | select(.name=="yaki-operator")'
# Where every node actually is
kubectl get nodes -o custom-columns='NAME:.metadata.name,VERSION:.status.nodeInfo.kubeletVersion,RUNTIME:.status.nodeInfo.containerRuntimeVersion'
# The upgrade history
kubectl get knu
# Anything still in flight
kubectl get jobs -n kmetal-yaki-operator
A fleet where the control-plane version is ahead of the workers is the normal intermediate state of a two-step upgrade.
A fleet that stays that way is a worker upgrade that was never created, or one that failed — kubectl get knu distinguishes the two.