Cluster Maintenance¶
Routine care of a running cluster: reading its health, taking a node out of service, and parking a cluster you are not using.
Reading health¶
Two views, and they answer different questions.
Is the platform happy with my cluster? — from your Environment:
$ kubectl get cluster,machinedeployment,machine -n dynamo-prod
NAME PHASE AGE VERSION
cluster.cluster.x-k8s.io/app Provisioned 4d v1.34.1
NAME REPLICAS READY UPDATED AGE
machinedeployment.cluster.x-k8s.io/app-md-0-x7k 3 3 3 4d
Is my cluster happy with itself? — with the cluster's kubeconfig:
kubectl --kubeconfig app.kubeconfig get nodes
kubectl --kubeconfig app.kubeconfig get pods -A --field-selector=status.phase!=Running
kubectl get nodes returns workers only.
Your control plane runs as pods on the platform, so it is not a node and never appears — that is correct, not a missing node.
Your control plane's pods live in your Environment alongside the Cluster:
You can read them. You cannot edit them, and there is nothing there you should want to.
Your workers are replaced without you¶
The platform watches your worker machines and replaces one that stops being healthy.
The policy is a MachineHealthCheck written by your platform team — you can see it, and you cannot change it:
So a worker disappearing and a new one appearing in its place is normal operation, not an incident.
What that asks of you is only that your workloads tolerate it: PodDisruptionBudgets on anything that needs a quorum, and no state kept on a node.
A machine that keeps being replaced is worth reporting — the loop is doing its job and the underlying cause is not something you can see.
Draining a node¶
Ordinary Kubernetes, run inside your cluster:
kubectl --kubeconfig app.kubeconfig cordon <node>
kubectl --kubeconfig app.kubeconfig drain <node> \
--ignore-daemonsets --delete-emptydir-data
kubectl --kubeconfig app.kubeconfig uncordon <node>
Two things specific to this platform:
Volumes follow the pod. A drained pod's volume is detached from that node's VM and hot-plugged onto wherever the pod lands. You do not need to do anything for storage to follow — see Persistent Storage.
Do not fix a node, replace it. There is nothing to patch on a worker: it is built from an image, and anything you change by hand on it is lost the next time it is replaced.
Forcing a replacement is not yours to do, though — Machine objects are the platform's, and you hold no verb on them.
What you have is the health check, which replaces an unhealthy worker on its own, and your platform team for the cases where it does not.
Drain the node so nothing is running on it, and hand over the machine name.
Certificates¶
Your control plane's certificates are rotated by the platform. There is nothing to do and no expiry to watch.
Inside the cluster, kubelet certificates renew themselves as normal. CertificateSigningRequests should be approved automatically:
A pile of Pending CSRs is worth reporting rather than approving by hand — it usually means something in the cluster's own bootstrap path is not working.
Your kubeconfig is not rotated for you
The admin kubeconfig in your Environment is a long-lived cluster-admin credential.
It stays valid, which is convenient and is also the risk — there is no expiry that will quietly stop a leaked copy from working.
Treat it like a root password: fetch it when you need it rather than distributing it, and use your own RBAC inside the cluster for everything routine.
Parking a cluster¶
To stop paying for a cluster's workers without losing the cluster, scale the workers to zero:
kubectl patch cluster app -n dynamo-prod --type=json \
-p '[{"op":"replace","path":"/spec/topology/workers/machineDeployments/0/replicas","value":0}]'
Bring it back the same way with a non-zero count.
What survives and what does not:
| Survives | The control plane and every API object in it. Your Deployments are still there, at zero available replicas. |
| Survives | PersistentVolumeClaims and the volumes behind them, and the quota they consume. |
| Goes | The worker VMs, and the capacity they were using. |
So this is the right move for a cluster you will come back to, and it is not a way to reduce storage consumption — the volumes are still yours and still counted.
The control plane cannot be scaled to zero
The kubevirt-kubeadm ClusterClass does not expose control-plane replicas as a topology variable, so there is no supported way to park the control plane itself.
A cluster at zero workers still runs its control plane. If that is not acceptable, the option is to back it up and delete it — see Cluster Backup & Restore and Delete Clusters.
When something is wrong¶
- A cluster that will not come up, or nodes that stay
NotReady— Cluster Troubleshooting. - Workloads that will not run, storage that will not bind, a Service with no address — Workload Troubleshooting.
See Also: Scale Clusters · Upgrade Clusters · Data Protection