Skip to content

Workload Troubleshooting

Workloads that will not run, storage that will not bind, Services with no address.

Generic Kubernetes failures — ImagePullBackOff, CrashLoopBackOff, a Service whose selector matches nothing — behave here exactly as they do anywhere, and the upstream debugging guide is better than a copy of it would be.

This page is the failures that are specific to running on kMetal, where the usual reflex points at the wrong thing.

First: which cluster are you asking?

Half of "that command shows nothing" is the wrong context.

kubectl config current-context

Your workloads, PVCs and Services are in your cluster (its own kubeconfig). Your Cluster object, addresses, quota and backups are in your Environment (your platform credentials).

A PVC that is rejected at kubectl apply

$ kubectl apply -f pvc.yaml
Error from server: admission webhook denied the request: ...

This is not a normal Kubernetes failure and it is not a broken cluster. An admission webhook checks your tenant's storage quota when you create the claim, deliberately, so you find out immediately rather than watching a claim sit Pending.

The message is about Do
Quota, or an exhausted budget Check the real number in your Environment (below) and free something or ask for more.
Reaching the webhook, or a call failing Report it. The webhook is registered to fail closed, so this denies your claim rather than skipping the check. Nothing to fix from your side.
kubectl get resourcequota -n dynamo-prod      # your Environment

The number that matters is aggregated across every Environment and cluster your tenant owns, so it is not the sum of what this cluster asked for. See Persistent Storage.

A PVC accepted, then stuck Pending

The claim is in, so the quota was fine. Provisioning on the platform has not finished.

kubectl describe pvc <name>       # events carry the provisioner's own message

There is a real limit to what you can see here: the volume is an object on the platform's side, on storage you have no visibility into. If the events say nothing useful, this is a platform-side diagnosis rather than something to keep digging at.

Give your platform team the claim's namespace, name and UID — their object is named after the same UID:

kubectl get pvc <name> -o jsonpath='{.metadata.uid}{"\n"}'

A pod Pending with a volume node-affinity conflict

A volume is hot-plugged onto the worker VM its pod was scheduled to, and is attached to one VM at a time.

So two pods that both mount the same ReadWriteOnce claim cannot run on different nodes — the second stays Pending. That is standard Kubernetes behaviour, but it surfaces more often here because there is one StorageClass and ReadWriteOnce is what to plan for.

kubectl describe pod <name>
kubectl get pvc <name> -o jsonpath='{.spec.accessModes}{"\n"}'

Fixes, in order of preference: one writer, or a StatefulSet with a claim per replica, or ask your platform team what the backend behind the class can actually offer for shared access.

A LoadBalancer Service with EXTERNAL-IP: <pending>

Three distinct causes, and the events tell them apart.

No events at all — spec.loadBalancerClass: kmetal is missing. The platform is not looking at your Service, so nothing will ever happen. This is the most common one, because the field is easy to leave out and produces total silence.

kubectl get svc <name> -o jsonpath='{.spec.loadBalancerClass}{"\n"}'

no free EipClaim available in namespace "..." — every address your Environment holds is already bound. Bindings are exclusive, one Service per claim.

Waiting, with a warning on the claim itself — the claim you named by annotation is held by another Service.

kubectl describe svc <name>                   # in your cluster
kubectl get eipclaims -n dynamo-prod          # in your Environment

Read BOUND and READY on the claims. A claim that is not Ready cannot serve anything. See Exposing Workloads.

A LoadBalancer Service denied at kubectl apply

Same webhook as the PVC case — it also resolves the address binding. A denial here is platform-side; report it rather than working around it.

The address answers from inside but not from outside

The Service is bound and traffic reaches your pods from within your own VPC, but nothing arrives from elsewhere.

That is the edge of what you can see. The address lives on the platform's provider network, and whether the wider network routes to it is the platform team's and the network team's business.

What you can confirm before escalating:

# The binding actually happened, and to which address
kubectl get eipclaim <claim> -n dynamo-prod -o jsonpath='{.status.binding}{"\n"}'

# Your backends are healthy
kubectl get endpointslices -l kubernetes.io/service-name=<svc>

With a bound claim and healthy endpoints, the remaining problem is routing.

Something you cannot reach at all

Two destinations are blocked below your cluster, so no amount of NetworkPolicy or DNS work will open them:

  • Another tenant's anything. No route exists.
  • The platform's own API server and services. Blocked from tenant workloads by policy.

Your egress to the internet and to whatever your platform team has routed works normally. A connection to one of those two, though, is not a misconfiguration to debug — see Networking.

An add-on that keeps coming back after you delete it

The platform delivers your cluster's CNI, the CSI node components and the storage-quota webhook configuration, and it re-applies them on drift.

Deleting one is reverted, and narrowing the webhook configuration is reverted specifically. That is intentional: a cluster whose CNI a tenant removed is a support call.

kubectl get pods -n kube-system

If one of them is genuinely misbehaving, that is a report, not a thing to remove.


See Also: Cluster Troubleshooting · Persistent Storage · Exposing Workloads