Skip to content

Platform Configuration

kMetal's platform configuration is a single cluster-scoped object, KMetal. This page is about how to approach it: what belongs in it, what deliberately does not, and how to change it safely.

For the field-by-field surface — every key, its default and its constraints — see Platform Configuration Reference.

One object, two ways to deliver it

The KMetal object can be created by the operator's Helm chart, as kmetal.spec in the values:

kmetal:
  name: kmetal
  spec:
    networking: { ... }        # exactly what goes under spec: on the object
    storage: { ... }

Or applied separately, with the chart installing only the operator — which is the right shape when the platform is described in a GitOps tree:

kmetal:
  spec: {}                     # nothing is created while this is empty

The chart passes the spec through verbatim. It is not modelled in values.yaml, because the CustomResourceDefinition already carries the schema, the defaults and the validation — a second copy would only be a place for the two to disagree.

What the spec covers

Seven blocks, and only two of them are required:

Block Required What it decides
networking yes The overlay interface and its central nodes, the provider networks tenants egress through, the service CIDR, and the LoadBalancer address pools.
storage yes Which existing StorageClasses the platform uses for its own state and hands to tenants.
clusterClass no A passthrough to the chart tenant clusters are templated from.
multiTenancy no Which identities Capsule enforces tenancy for. Absent means it enforces nothing.
console no The web console. Absent means it is not installed.
registry no Secrets holding credentials for kMetal's own images and charts.
placement no Where the platform's controllers run. Defaults to the control plane.

What is deliberately absent

The spec carries values with no sane default, plus a few explicit escape hatches. Everything a kMetal release pins is missing from the API on purpose:

  • Component versions. Every chart and artifact version travels with the operator release. Upgrading kMetal means bumping the operator — see Upgrades.
  • The component set and its order. Which components are installed, and in what sequence, is decided and tested per release.
  • Safety-critical settings. Kube-OVN's non-primary CNI flags, its CNI config priority, the CDI filesystem overhead on the tenant class. Getting any of these wrong takes down pod networking or silently doubles what a tenant's volumes consume.
  • Public names. The kmetal-flux namespace, the tenant-cp-pool and mgmt-pool address pools, the kmetal-ca issuer. Specs elsewhere reference these by name, so making them configurable would strand whatever already points at them.
  • Capsule's namespace naming policy. forceTenantPrefix and protectedNamespaceRegex change the rule namespaces are admitted under rather than who is subject to it, so flipping either on a running platform starts rejecting namespace names that were accepted the day before. multiTenancy exposes the entitlement side only.

If you find yourself needing one of these, that is a signal to talk to us rather than a gap to work around — see Component Configuration for the escape hatches that do exist.

Nothing that entitles a subject gets a default

multiTenancy is the one block where an empty field is a deliberate refusal rather than a convenience.

Who owns a tenant is the identity provider's answer, not the platform's. A subject a kMetal release wrote into multiTenancy.users would be an authorization in force on every cluster that nobody asked for, expressed in a place — a pinned release value — where a site has no way to say no. So kMetal names nobody: not a user, not an administrator, not an exempt group, not even the group covering the ServiceAccounts its own GitOps tooling makes Tenant owners out of.

The cost of that position is that the block is easy to forget, and forgetting it fails quietly: a Tenant whose owner is not listed in multiTenancy.users is invisible to Capsule, so that owner's Environments are created outside every tenant with no quota and no constraint applied, and nothing reports it. Set it in the same change that onboards the first tenant — see Multi-Tenancy Tasks.

Decide these before the first install

Some fields are cheap to change later; a few are not, because tenants and their data end up depending on them.

Field Why it is hard to change
networking.kubeOVN.podCIDR The overlay's own subnet. Changing it renumbers everything attached to it.
networking.serviceCIDR Must match the control plane's --service-cluster-ip-range, which is fixed at bootstrap.
networking.providerNetworks[].subnets[].cidrBlock Every tenant VPC blocks egress to these ranges; tenant addresses are allocated from them.
storage.underclusterClassName Holds the golden OS images and every tenant machine's root disk. Moving it means rebuilding the images and rolling every machine.
storage.etcdClassName Holds tenant control-plane etcd data. Moving it means migrating that data.
storage.tenantClassName Tenant volumes live on it, and the tenant's storage quota is keyed on its name.

The rest — the address pool ranges, the console, placement, multiTenancy, the ClusterClass values — can be edited on a running platform and reconciled in place.

See Configuration Best Practices for the reasoning behind each of these.

Changing the configuration

The spec is reconciled continuously, so an edit is applied by editing the object:

kubectl edit km kmetal

Or, if the chart owns it, by upgrading the release with the new values:

helm upgrade kmetal-operator oci://ghcr.io/clastix/charts/kmetal-operator \
  --namespace kmetal-operator-system \
  --reuse-values -f values.yaml

Then watch the components converge:

kubectl get km kmetal -o jsonpath='{range .status.components[*]}{.name}{"\t"}{.phase}{"\t"}{.message}{"\n"}{end}'

The reconciler halts at the first component that is not ready, so a spec the platform cannot satisfy shows up as one named component stalling — not as a silent partial apply.

Validating before you apply

The API server carries the validation, so it is the fastest check:

kubectl apply --dry-run=server -f kmetal.yaml

To see the objects the operator would actually create, render the catalogue instead:

kmetal render -f kmetal.yaml --from-cluster -o install.yaml

A render is what the operator would apply, built from the same component catalogue. What it does not reproduce is the operator's behaviour — no readiness gating, no halt on failure, no drift correction.


See Also: Platform Configuration Reference · Component Configuration · Configuration Best Practices