Creating Clusters¶
Create tenant Kubernetes clusters in kMetal. Each tenant cluster is a CAPI Cluster resource that references the platform-provided kubevirt-kubeadm ClusterClass — the ClusterClass drives the rest (hosted control plane, worker VMs, networking).
Prerequisites¶
Four things have to exist before a Cluster can be created, and all four live in your Environment rather than in the cluster:
| What it is | Whose | |
|---|---|---|
| An Environment | The namespace you create Cluster objects in |
Your platform team gives you this |
| A subnet | The IP space your worker VMs attach to, from a SubnetClaim |
Theirs, or yours to claim |
| A control-plane address | The endpoint your kubeconfig and workers connect to | Theirs — it comes from a pool you cannot see |
| A bootstrap-server endpoint | What your workers fetch their join configuration from | Theirs |
If any of them is missing or not Ready, that is the conversation to have before writing a Cluster — a cluster whose subnet is not ready comes up with a control plane and no workers.
You do not need an EipClaim to create a cluster.
Those are for exposing workloads once it is running — see Exposing Workloads.
Create via YAML¶
# my-cluster.yaml
apiVersion: cluster.x-k8s.io/v1beta2
kind: Cluster
metadata:
name: my-cluster
namespace: <tenant-namespace>
spec:
clusterNetwork:
pods:
cidrBlocks:
- 10.200.0.0/16 # tenant cluster pod CIDR
services:
cidrBlocks:
- 10.201.0.0/16 # tenant cluster service CIDR
topology:
classRef:
name: kubevirt-kubeadm
namespace: kmetal-capi-providers
version: v1.34.1 # tenant Kubernetes version
variables:
- name: controlPlane
value:
dataStoreName: default
network:
certSANs:
- <public-fqdn-or-ip>
serviceAddress: <metallb-vip> # e.g. 172.21.204.201
serviceType: LoadBalancer
- name: machineSize
value: small # small | medium | large | custom (defined by ClusterClass)
- name: bootstrapConfig
value:
serverAddress: <bootstrap-server> # e.g. 10.100.0.254
- name: network
value:
subnet: <ovn-subnet-name> # tenant OVN Subnet
workers:
machineDeployments:
- class: default-worker
name: md-0
replicas: 2
Apply it:
machineSize picks a worker shape: small (2 CPU, 4 GiB), medium (4 CPU, 8 GiB), large (8 CPU, 16 GiB), or custom.
With custom, add a machineSpecs variable naming cpuCores (1-64) and memoryGiB (2-128); the platform may not offer every size, and your quota still bounds the total.
The ClusterClass topology drives the rest:
- A
KamajiControlPlane(the CAPI-managed Kamaji control plane) is auto-created. - A
KubevirtCluster+ KubevirtMachineTemplate are auto-created. - A
MachineDeploymentis auto-created for the workers. - Kamaji provisions the tenant control plane pods (api-server, controller-manager, scheduler, konnectivity-server).
- CAPK provisions tenant worker VMs from the ClusterClass tier.
Provisioning typically takes a few minutes. Watch progress:
Access Cluster¶
Get Kubeconfig¶
# Kamaji creates the kubeconfig secret as <cluster-name>-admin-kubeconfig in the cluster's namespace
kubectl get secret my-cluster-admin-kubeconfig \
-n <tenant-namespace> \
-o jsonpath='{.data.admin\.conf}' | base64 -d > my-cluster.kubeconfig
export KUBECONFIG=my-cluster.kubeconfig
kubectl get nodes
Verify Cluster¶
# Under-cluster view
kubectl get cluster -n <tenant-namespace> # CAPI Cluster
kubectl get kamajicontrolplane -n <tenant-namespace> # control plane status
kubectl get tenantcontrolplane -n <tenant-namespace> # underlying Kamaji TCP
kubectl get machinedeployment,machines -n <tenant-namespace> # workers
kubectl get kubevirtcluster -n <tenant-namespace> # CAPK infrastructure
# Tenant view (with kubeconfig)
kubectl get nodes
kubectl get pods -n kube-system
kubectl get storageclass
When it does not come up¶
Two failures dominate, and they are diagnosed differently:
# The cluster and its control plane
kubectl describe cluster my-cluster -n dynamo-prod
kubectl get pods -n dynamo-prod -l kamaji.clastix.io/name=my-cluster
# The workers
kubectl get machinedeployment,machine -n dynamo-prod
kubectl describe machine <machine-name> -n dynamo-prod
A control plane that never becomes available is usually quota, or the control-plane address.
Workers that provision but never join — a Machine that reaches Running while kubectl get nodes stays empty — is usually the subnet, the address being unreachable from the worker side, or the bootstrap server.
Nodes NotReady is expected for the first minute or two
A fresh cluster's nodes are NotReady until the platform delivers the CNI into it.
That happens automatically and takes a moment; it is not something you install.
Nodes still NotReady after several minutes is worth reporting.
See Cluster Troubleshooting for the rest, and give your platform team the cluster's namespace and name — most of what is left is on their side of the boundary.
Next steps¶
- Persistent Storage — claim a volume.
- Exposing Workloads — get traffic to it.
- Data Protection — before it holds anything you care about.