Skip to content

Platform Validation

After installing kMetal, validate that the under cluster's platform components are healthy before provisioning tenant clusters.

The platform's own report

Start here, not with pods. The operator publishes one status entry per component, so the platform tells you whether it converged:

kubectl get km

READY is True, and VERSION is the release you installed, only when every component reconciled at the version that release pins.

# Per component, in install order
kubectl get km kmetal -o jsonpath='{range .status.components[*]}{.name}{"\t"}{.phase}{"\t"}{.message}{"\n"}{end}'

Every entry should be Ready. The reconciler halts at the first component that is not, so a single Failed or Installing entry followed by Pending ones tells you exactly where to look — the message names what it is waiting on.

See Component Reference for the full component list and what each category costs you.

The operator itself

helm status kmetal-operator -n kmetal-operator-system
kubectl get pods -n kmetal-operator-system

The release should be deployed, and the operator Pod Running. A Pod stuck Pending here is usually the node selector: the chart pins the operator to control-plane nodes by default.

Component pods

Only worth looking at once kubectl get km has named a component as the problem. Component namespaces are pinned by the release rather than configured, so find the failing component's namespace from its Flux objects:

kubectl get helmreleases,kustomizations,ocirepositories -n kmetal-flux

# Anything not Running or Completed, anywhere
kubectl get pods -A | grep -v Running | grep -v Completed

MetalLB

# Pool(s) should match the overlay you applied
kubectl get ipaddresspool -n kmetal-metallb
kubectl get l2advertisement -n kmetal-metallb

# Smoke-test allocation
kubectl create service loadbalancer test-lb --tcp=80:80
kubectl get service test-lb -w     # should get EXTERNAL-IP from the pool
kubectl delete service test-lb

StorageClasses

kMetal installs no storage: the classes named in the spec must already exist and provision.

kubectl get storageclass
kubectl get volumesnapshotclass
kubectl get km kmetal -o jsonpath='{.spec.storage}{"\n"}'

storage.underclusterClassName and storage.tenantClassName must both be present, and etcdClassName too if you named one. The platform class's driver has to ship a VolumeSnapshotClass: golden OS images are snapshotted on it and every machine's root disk is cloned from that snapshot, so a driver without one cannot provision workers.

Confirm CDI agrees, rather than reading it off the driver's documentation:

$ kubectl get storageprofile platform-storage \
    -o jsonpath='{.status.cloneStrategy}{"\t"}{.status.snapshotClass}{"\n"}'
csi-clone   platform-storage-snapshot

An empty snapshotClass is a platform that installs cleanly and then never boots a tenant machine.

The two names should also differ, even where one backend serves both — storage quota counts by class name, and machine root disks are provisioned in the tenant's own namespace. See Keep the platform class and the tenant class apart.

See Storage for what the platform expects of the backend.

Compute nodes for tenant VMs

Tenant machines run only where KubeVirt's virt-handler runs, and kMetal pins it to nodes labelled node-role.kubernetes.io/compute.

$ kubectl get nodes -l node-role.kubernetes.io/compute
NAME               STATUS   ROLES             AGE    VERSION
kmetal-worker-01   Ready    compute,network   4d5h   v1.35.8
kmetal-worker-02   Ready    compute,network   4d5h   v1.35.8
kmetal-worker-03   Ready    compute,network   4d5h   v1.35.8

$ kubectl get pods -n system-kubevirt -l kubevirt.io=virt-handler -o wide

An empty first result is a platform that looks healthy and cannot start a single tenant machine — or build a golden image, since the installer is itself a VM. Label the nodes that carry tenant workloads, and keep them off the control plane deliberately: the platform is supposed to survive the VMs it hosts.

kubectl label node <node> node-role.kubernetes.io/compute=""

Tenancy enforcement

Capsule being Running says nothing about whether it enforces anything: it acts only on requests from subjects named in spec.multiTenancy, and kMetal names none by default.

kubectl get km kmetal -o jsonpath='{.spec.multiTenancy}{"\n"}'
kubectl get capsuleconfigurations -o yaml

An empty result on the first command is a platform where tenancy is installed and inert — every Tenant will be accepted and then ignored. Whether that is right depends on whether you have onboarded anyone yet; it is the wrong state to discover from a tenant's missing quota.

The functional check is a namespace creation impersonating a subject the spec names:

kubectl create namespace tenancy-probe --as=<subject-in-multiTenancy-users>

A rejection is the passing result. A subject Capsule matches but that owns no Tenant is refused, with a message saying so — which is proof admission saw the request. If the namespace is instead created and carries no capsule.clastix.io/tenant label, Capsule was never told about that subject:

kubectl get namespace tenancy-probe -o jsonpath='{.metadata.labels}{"\n"}'
kubectl delete namespace tenancy-probe

See Make Capsule aware of the owner.

CAPI providers

kubectl get coreproviders,bootstrapproviders,infrastructureproviders,controlplaneproviders -n kmetal-capi-providers
# Expect (one of each): core (cluster-api), bootstrap-kubeadm, infrastructure-kubevirt (CAPK), control-plane-kamaji (CACPK)

KubeVirt readiness

kubectl get kubevirt -n system-kubevirt
# .status.phase should be Deployed

CRDs

Spot-check the APIs tenants will actually use:

# Cluster lifecycle
kubectl get crd clusters.cluster.x-k8s.io
kubectl get crd kamajicontrolplanes.controlplane.cluster.x-k8s.io
kubectl get crd virtualmachines.kubevirt.io

# The kMetal claims a tenant writes for their networking
kubectl get crd vpcclaims.network.kmetal.io
kubectl get crd subnetclaims.network.kmetal.io
kubectl get crd eipclaims.network.kmetal.io

# The underlay segments the operator creates from spec.networking.providerNetworks
kubectl get crd providernetworks.network.kmetal.io

Check the claim APIs rather than Kube-OVN's own vpcs.kubeovn.io and subnets.kubeovn.io. Those exist too, but they are the platform's to write — a tenant never touches them, so a missing claim CRD is what actually blocks onboarding.

Troubleshooting

If any component fails:

# Pod status in the affected namespace
kubectl get pods -n <namespace> --sort-by=.metadata.creationTimestamp

# Events
kubectl get events -A --sort-by='.lastTimestamp' | tail -50

# Resource pressure
kubectl top nodes
kubectl top pods -A --sort-by=cpu | head -20

See Troubleshooting for component-specific guidance.

Next Steps