Skip to content

Cluster Troubleshooting

A cluster that will not come up, workers that never join, a control plane you cannot reach.

Everything here runs with your Environment credentials unless it says otherwise. Your control plane runs as pods in your Environment, next to the Cluster — not in a platform namespace, and not on a node.

Start here

kubectl get cluster,machinedeployment,machine -n dynamo-prod
kubectl get events -n dynamo-prod --sort-by='.lastTimestamp' | tail -30

Those two answer "how far did it get", which decides everything below.

What you see Read on
Cluster never leaves Provisioning Control plane never becomes available
Cluster Provisioned, Machines Running, kubectl get nodes empty Workers provision but never join
Nodes present but NotReady Nodes stay NotReady
Nothing was created at all The create was refused

The create was refused

If kubectl apply returned an error rather than creating the object, nothing is broken — you were told no.

The two common reasons are quota (clusters, or storage — which on some platforms includes the disks a cluster's workers boot from) and permissions.

kubectl get resourcequota -n dynamo-prod
kubectl auth can-i create clusters.cluster.x-k8s.io -n dynamo-prod

Both are conversations with your platform team rather than things to retry.

Control plane never becomes available

kubectl describe cluster my-cluster -n dynamo-prod
kubectl get pods -n dynamo-prod -l kamaji.clastix.io/name=my-cluster

Those pods are your api-server, controller-manager, scheduler and konnectivity-server. You can read them and their logs; you cannot edit them, and restarting them is not your call.

Pods are Usually
Pending The platform is out of capacity, or your quota is exhausted.
Restarting A datastore problem behind them. Platform-side.
Absent entirely The control plane was never templated — look at the Cluster's conditions and events.

The other frequent cause is the control-plane address: an address that is not actually available leaves the Service in front of the control plane without one, and nothing downstream can connect.

kubectl get svc -n dynamo-prod | grep my-cluster

An EXTERNAL-IP of <pending> there is a platform-side problem — that address comes from a pool you cannot see or change.

Workers provision but never join

The machines exist, they are running, and the cluster has no nodes. This is almost always the worker being unable to reach something rather than the worker being broken.

kubectl get machine -n dynamo-prod
kubectl describe machine <machine-name> -n dynamo-prod
kubectl get virtualmachineinstance -n dynamo-prod

Three candidates, in order of likelihood:

The subnet. No Ready subnet, or an exhausted one, and a worker has no address.

kubectl get vpcclaims,subnetclaims -n dynamo-prod

The control-plane endpoint, from the worker's side. A worker joins by talking to the API server. An address that answers from your laptop but not from inside the tenant network gets exactly this symptom.

The bootstrap server. Workers fetch their join configuration from it. If it is unreachable, they boot and sit there.

The last two are platform-side. Confirm the subnet yourself, then hand over the cluster name and one machine name.

Nodes stay NotReady

Expected for the first minute or two of a new cluster: nodes are NotReady until the platform delivers the CNI into it.

kubectl --kubeconfig my-cluster.kubeconfig get nodes
kubectl --kubeconfig my-cluster.kubeconfig get pods -n kube-system

The CNI is not something you install, and it is re-applied if it is removed. Nodes still NotReady after several minutes, or CNI pods crash-looping, is a report rather than a fix — see Deploying Applications.

The kubeconfig does not work

Re-fetch it before anything else. It is a Secret in your Environment and it is always current:

kubectl get secret my-cluster-admin-kubeconfig -n dynamo-prod \
  -o jsonpath='{.data.admin\.conf}' | base64 -d > my-cluster.kubeconfig

export KUBECONFIG=my-cluster.kubeconfig
kubectl cluster-info
Symptom Means
x509: certificate signed by unknown authority You are using a kubeconfig from a different cluster. Each cluster has its own CA, and a restored cluster gets a new one.
Connection timeout The endpoint is not reachable from where you are. Network, not credentials.
Unauthorized Almost always the wrong file. Re-fetch.

A restore is the case that catches people: the restored cluster is a new cluster with a new identity, so every kubeconfig, CI credential and client config pointing at the old one has to be repointed.

Scaling does nothing

You changed the replica count and nothing happened, or it reverted.

kubectl get cluster my-cluster -n dynamo-prod \
  -o jsonpath='{.spec.topology.workers.machineDeployments[0].replicas}{"\n"}'
kubectl get machinedeployment -n dynamo-prod

If the topology says what you wanted and the MachineDeployment does not, wait — reconciliation is not instant. If the topology does not say it, you patched the generated MachineDeployment instead of the Cluster, and it was reverted. See Scale Clusters.

New machines stuck Provisioning rather than reverting is capacity or quota.

Deletion hangs

Deletion is not instant. Worker VMs and their volumes take as long as they take to tear down, so give it minutes before concluding it is stuck.

kubectl get cluster,machine -n dynamo-prod
kubectl describe cluster my-cluster -n dynamo-prod

Do not strip finalizers

Removing a finalizer from a Cluster does not finish the deletion — it abandons it. Worker VMs, volumes and addresses are left allocated on the platform with nothing owning them, still consuming your quota, and now with no controller that will ever clean them up.

A genuinely stuck deletion is a platform-side diagnosis. Report it with the cluster name and leave the object as it is.

Collecting something useful to hand over

Where your visibility ends is real, and the fastest escalation is one that includes what you could see:

kubectl get cluster,machinedeployment,machine -n dynamo-prod -o yaml > cluster.yaml
kubectl get events -n dynamo-prod --sort-by='.lastTimestamp' > events.txt
kubectl get pods -n dynamo-prod -l kamaji.clastix.io/name=my-cluster -o yaml > controlplane.yaml
kubectl get vpcclaims,subnetclaims,eipclaims -n dynamo-prod -o yaml > claims.yaml

Include the Environment name and the cluster name. Everything else your platform team can look up from those.


See Also: Workload Troubleshooting · Create Clusters · Maintenance