Skip to content

Cluster Maintenance

Routine care of a running cluster: reading its health, taking a node out of service, and parking a cluster you are not using.

Reading health

Two views, and they answer different questions.

Is the platform happy with my cluster? — from your Environment:

$ kubectl get cluster,machinedeployment,machine -n dynamo-prod
NAME                              PHASE         AGE   VERSION
cluster.cluster.x-k8s.io/app      Provisioned   4d    v1.34.1

NAME                                              REPLICAS   READY   UPDATED   AGE
machinedeployment.cluster.x-k8s.io/app-md-0-x7k    3          3       3        4d

Is my cluster happy with itself? — with the cluster's kubeconfig:

kubectl --kubeconfig app.kubeconfig get nodes
kubectl --kubeconfig app.kubeconfig get pods -A --field-selector=status.phase!=Running

kubectl get nodes returns workers only. Your control plane runs as pods on the platform, so it is not a node and never appears — that is correct, not a missing node.

Your control plane's pods live in your Environment alongside the Cluster:

kubectl get pods -n dynamo-prod -l kamaji.clastix.io/name=app

You can read them. You cannot edit them, and there is nothing there you should want to.

Your workers are replaced without you

The platform watches your worker machines and replaces one that stops being healthy. The policy is a MachineHealthCheck written by your platform team — you can see it, and you cannot change it:

kubectl get machinehealthcheck -n dynamo-prod

So a worker disappearing and a new one appearing in its place is normal operation, not an incident. What that asks of you is only that your workloads tolerate it: PodDisruptionBudgets on anything that needs a quorum, and no state kept on a node.

A machine that keeps being replaced is worth reporting — the loop is doing its job and the underlying cause is not something you can see.

Draining a node

Ordinary Kubernetes, run inside your cluster:

kubectl --kubeconfig app.kubeconfig cordon <node>
kubectl --kubeconfig app.kubeconfig drain <node> \
  --ignore-daemonsets --delete-emptydir-data
kubectl --kubeconfig app.kubeconfig uncordon <node>

Two things specific to this platform:

Volumes follow the pod. A drained pod's volume is detached from that node's VM and hot-plugged onto wherever the pod lands. You do not need to do anything for storage to follow — see Persistent Storage.

Do not fix a node, replace it. There is nothing to patch on a worker: it is built from an image, and anything you change by hand on it is lost the next time it is replaced.

Forcing a replacement is not yours to do, though — Machine objects are the platform's, and you hold no verb on them. What you have is the health check, which replaces an unhealthy worker on its own, and your platform team for the cases where it does not. Drain the node so nothing is running on it, and hand over the machine name.

Certificates

Your control plane's certificates are rotated by the platform. There is nothing to do and no expiry to watch.

Inside the cluster, kubelet certificates renew themselves as normal. CertificateSigningRequests should be approved automatically:

kubectl --kubeconfig app.kubeconfig get csr

A pile of Pending CSRs is worth reporting rather than approving by hand — it usually means something in the cluster's own bootstrap path is not working.

Your kubeconfig is not rotated for you

The admin kubeconfig in your Environment is a long-lived cluster-admin credential. It stays valid, which is convenient and is also the risk — there is no expiry that will quietly stop a leaked copy from working.

Treat it like a root password: fetch it when you need it rather than distributing it, and use your own RBAC inside the cluster for everything routine.

Parking a cluster

To stop paying for a cluster's workers without losing the cluster, scale the workers to zero:

kubectl patch cluster app -n dynamo-prod --type=json \
  -p '[{"op":"replace","path":"/spec/topology/workers/machineDeployments/0/replicas","value":0}]'

Bring it back the same way with a non-zero count.

What survives and what does not:

Survives The control plane and every API object in it. Your Deployments are still there, at zero available replicas.
Survives PersistentVolumeClaims and the volumes behind them, and the quota they consume.
Goes The worker VMs, and the capacity they were using.

So this is the right move for a cluster you will come back to, and it is not a way to reduce storage consumption — the volumes are still yours and still counted.

The control plane cannot be scaled to zero

The kubevirt-kubeadm ClusterClass does not expose control-plane replicas as a topology variable, so there is no supported way to park the control plane itself.

A cluster at zero workers still runs its control plane. If that is not acceptable, the option is to back it up and delete it — see Cluster Backup & Restore and Delete Clusters.

When something is wrong


See Also: Scale Clusters · Upgrade Clusters · Data Protection