Configuration Best Practices¶
The KMetal spec is small, and most of it can be edited on a running platform.
A handful of fields cannot — not because the API rejects a change, but because tenants and their data come to depend on them.
This page is about telling those two groups apart.
Decide these before the first tenant¶
Every field here is cheap to set and expensive to revisit. Get them from the network and storage teams before the install, not after.
networking.serviceCIDR¶
Must match the --service-cluster-ip-range the under cluster's control plane was bootstrapped with.
That value is fixed at bootstrap, so this is not really a decision — it is a value to look up and copy correctly.
Kube-OVN needs it to keep service traffic off the overlay; a mismatch produces service traffic that leaves through the wrong path and is hard to attribute.
networking.kubeOVN.podCIDR¶
The ovn-default subnet — not the cluster pod network, which the primary CNI owns.
The two must not overlap.
Size it for the tenant workloads that will attach to the overlay, and treat it as permanent: renumbering it renumbers everything attached.
networking.kubeOVN.centralNodes¶
The OVSDB Raft cluster. Use an odd count so it keeps quorum on node loss — three tolerates one failure, one tolerates none. Every node listed must carry the overlay interface, and the control plane usually does not, so it usually does not belong here.
Nodes can be changed later, but each change is a Raft membership change; do it deliberately, one at a time.
networking.providerNetworks[].subnets[].cidrBlock¶
Two things depend on these ranges permanently: every tenant VPC blocks egress to them automatically, and tenant-facing addresses are allocated from them.
Be generous with excludeIPs from the start.
The gateway belongs there, as does anything already assigned on the segment — addresses reserved for DNAT to tenant control planes, or for reaching the under cluster itself.
An address in use elsewhere that is not excluded will be handed to a tenant, which takes out whatever was using it.
storage.underclusterClassName, storage.etcdClassName and storage.tenantClassName¶
Three fields, because the consumers pull in different directions. Pointing everything at one class silently gets one of them wrong.
underclusterClassNameholds the golden OS images and the root disk of every tenant machine cloned from them. It must supportVolumeSnapshots, which rules out most node-local provisioners: without snapshots there is no golden volume and no machine boots.etcdClassNameholds the Kamaji datastore — tenant control-plane etcd data. Small, private, never touched by tenants, and it defaults to the class above. Set it where the snapshotting class is the wrong class for etcd; a node-local class here is a defensible choice, but be deliberate, because node-local means the etcd data of every hosted control plane is lost with the node.tenantClassNameis handed to tenants. It should be network-attached and survive the loss of any single node.
Give the platform class and the tenant class two names even when one backend serves both. Storage quota is keyed on the class name and machine root disks live in the tenant's namespace, so a shared name bills the platform's own volumes to the tenant — see Keep the platform class and the tenant class apart.
All of them are hard to change once tenants exist, because their data is on them.
storage.tenantClaimPropertySets¶
Set this before the first tenant volume, not after diagnosing a failure.
On a backend advertising more than one capability combination, CDI's own derived order can put ReadWriteMany + Block first, and the importer pod then fails on a permission error that names neither the storage class nor the volume mode.
Safe to change on a running platform¶
- Address pool ranges (
networking.loadBalancer.*) — growing a pool is routine. Shrinking one below what is already allocated is not. console— add or remove it freely; nothing else depends on it.placement— move the platform's controllers to a different class of node. Label the nodes first, or every platform Pod sitsPending.clusterClass.values— affects clusters created or reconciled afterwards.registry— rotate credentials by updating the Secret; the operator redistributes the copies.extraBlockedEgressCIDRs— additive, and takes effect on the next reconcile.
Let the release own the versions¶
The component set and every chart version travel with the operator release. There is no supported way to bump one component, and that is the point: a release is a combination that has been tested together.
The corollary is that upgrading kMetal is upgrading the operator. Plan those upgrades as platform events rather than as values edits — see Upgrades.
Keep the spec in version control¶
The KMetal object is the platform's definition.
Treat it the way you would treat any other production definition: in a repository, reviewed, applied from there.
That also gives you the diff when something changes behaviour, which the cluster itself will not tell you retrospectively.
Two things must never go in it:
- Credentials.
registrynames Secrets; it does not carry them. The object is cluster-scoped and readable by anyone with read access to it. - Anything generated per-run. The spec should render identically every time.
Validate before applying, and watch after¶
kubectl apply --dry-run=server -f kmetal.yaml # schema, defaults, CEL validation
kmetal render -f kmetal.yaml --from-cluster # the objects that would be applied
Then watch the components converge, and let the halt point tell you what went wrong:
kubectl get km kmetal -o jsonpath='{range .status.components[*]}{.name}{"\t"}{.phase}{"\t"}{.message}{"\n"}{end}'
The reconciler stops at the first component that is not Ready; everything after it reports Pending with the name of what it is waiting on.
Nothing is rolled back, so a failed change leaves the previous working state serving — which is the behaviour you want, but it does mean a stalled component is your signal to act rather than something that will clear itself.
Test where a mistake is cheap¶
The failure modes that matter here — a double-tagged VLAN, an exhausted address pool, a raw-block DataVolume — do not show up in a render. They show up on a cluster, usually at the moment a tenant tries to use something.
Keep a non-production platform with the same network shape, and make changes there first.
See Also: Platform Configuration · Component Configuration · Platform Configuration Reference