Platform Validation¶
After installing kMetal, validate that the under cluster's platform components are healthy before provisioning tenant clusters.
The platform's own report¶
Start here, not with pods. The operator publishes one status entry per component, so the platform tells you whether it converged:
READY is True, and VERSION is the release you installed, only when every component reconciled at the version that release pins.
# Per component, in install order
kubectl get km kmetal -o jsonpath='{range .status.components[*]}{.name}{"\t"}{.phase}{"\t"}{.message}{"\n"}{end}'
Every entry should be Ready.
The reconciler halts at the first component that is not, so a single Failed or Installing entry followed by Pending ones tells you exactly where to look — the message names what it is waiting on.
See Component Reference for the full component list and what each category costs you.
The operator itself¶
The release should be deployed, and the operator Pod Running.
A Pod stuck Pending here is usually the node selector: the chart pins the operator to control-plane nodes by default.
Component pods¶
Only worth looking at once kubectl get km has named a component as the problem.
Component namespaces are pinned by the release rather than configured, so find the failing component's namespace from its Flux objects:
kubectl get helmreleases,kustomizations,ocirepositories -n kmetal-flux
# Anything not Running or Completed, anywhere
kubectl get pods -A | grep -v Running | grep -v Completed
MetalLB¶
# Pool(s) should match the overlay you applied
kubectl get ipaddresspool -n kmetal-metallb
kubectl get l2advertisement -n kmetal-metallb
# Smoke-test allocation
kubectl create service loadbalancer test-lb --tcp=80:80
kubectl get service test-lb -w # should get EXTERNAL-IP from the pool
kubectl delete service test-lb
StorageClasses¶
kMetal installs no storage: the classes named in the spec must already exist and provision.
kubectl get storageclass
kubectl get volumesnapshotclass
kubectl get km kmetal -o jsonpath='{.spec.storage}{"\n"}'
storage.underclusterClassName and storage.tenantClassName must both be present, and etcdClassName too if you named one.
The platform class's driver has to ship a VolumeSnapshotClass: golden OS images are snapshotted on it and every machine's root disk is cloned from that snapshot, so a driver without one cannot provision workers.
Confirm CDI agrees, rather than reading it off the driver's documentation:
$ kubectl get storageprofile platform-storage \
-o jsonpath='{.status.cloneStrategy}{"\t"}{.status.snapshotClass}{"\n"}'
csi-clone platform-storage-snapshot
An empty snapshotClass is a platform that installs cleanly and then never boots a tenant machine.
The two names should also differ, even where one backend serves both — storage quota counts by class name, and machine root disks are provisioned in the tenant's own namespace. See Keep the platform class and the tenant class apart.
See Storage for what the platform expects of the backend.
Compute nodes for tenant VMs¶
Tenant machines run only where KubeVirt's virt-handler runs, and kMetal pins it to nodes labelled node-role.kubernetes.io/compute.
$ kubectl get nodes -l node-role.kubernetes.io/compute
NAME STATUS ROLES AGE VERSION
kmetal-worker-01 Ready compute,network 4d5h v1.35.8
kmetal-worker-02 Ready compute,network 4d5h v1.35.8
kmetal-worker-03 Ready compute,network 4d5h v1.35.8
$ kubectl get pods -n system-kubevirt -l kubevirt.io=virt-handler -o wide
An empty first result is a platform that looks healthy and cannot start a single tenant machine — or build a golden image, since the installer is itself a VM. Label the nodes that carry tenant workloads, and keep them off the control plane deliberately: the platform is supposed to survive the VMs it hosts.
Tenancy enforcement¶
Capsule being Running says nothing about whether it enforces anything: it acts only on requests from subjects named in spec.multiTenancy, and kMetal names none by default.
kubectl get km kmetal -o jsonpath='{.spec.multiTenancy}{"\n"}'
kubectl get capsuleconfigurations -o yaml
An empty result on the first command is a platform where tenancy is installed and inert — every Tenant will be accepted and then ignored.
Whether that is right depends on whether you have onboarded anyone yet; it is the wrong state to discover from a tenant's missing quota.
The functional check is a namespace creation impersonating a subject the spec names:
A rejection is the passing result. A subject Capsule matches but that owns no Tenant is refused, with a message saying so — which is proof admission saw the request.
If the namespace is instead created and carries no capsule.clastix.io/tenant label, Capsule was never told about that subject:
kubectl get namespace tenancy-probe -o jsonpath='{.metadata.labels}{"\n"}'
kubectl delete namespace tenancy-probe
See Make Capsule aware of the owner.
CAPI providers¶
kubectl get coreproviders,bootstrapproviders,infrastructureproviders,controlplaneproviders -n kmetal-capi-providers
# Expect (one of each): core (cluster-api), bootstrap-kubeadm, infrastructure-kubevirt (CAPK), control-plane-kamaji (CACPK)
KubeVirt readiness¶
CRDs¶
Spot-check the APIs tenants will actually use:
# Cluster lifecycle
kubectl get crd clusters.cluster.x-k8s.io
kubectl get crd kamajicontrolplanes.controlplane.cluster.x-k8s.io
kubectl get crd virtualmachines.kubevirt.io
# The kMetal claims a tenant writes for their networking
kubectl get crd vpcclaims.network.kmetal.io
kubectl get crd subnetclaims.network.kmetal.io
kubectl get crd eipclaims.network.kmetal.io
# The underlay segments the operator creates from spec.networking.providerNetworks
kubectl get crd providernetworks.network.kmetal.io
Check the claim APIs rather than Kube-OVN's own vpcs.kubeovn.io and subnets.kubeovn.io.
Those exist too, but they are the platform's to write — a tenant never touches them, so a missing claim CRD is what actually blocks onboarding.
Troubleshooting¶
If any component fails:
# Pod status in the affected namespace
kubectl get pods -n <namespace> --sort-by=.metadata.creationTimestamp
# Events
kubectl get events -A --sort-by='.lastTimestamp' | tail -50
# Resource pressure
kubectl top nodes
kubectl top pods -A --sort-by=cpu | head -20
See Troubleshooting for component-specific guidance.