Platform Configuration Reference¶
kMetal's undercluster is installed by the kMetal operator from a single cluster-scoped KMetal object.
There are two configuration surfaces, and they do different jobs:
| Surface | What it configures | Schema owner |
|---|---|---|
kmetal-operator chart values |
How the operator itself runs — replicas, image, RBAC, CRD lifecycle | the chart's values.yaml |
KMetal spec (clastix.io/v1) |
The undercluster: networking, storage, ClusterClass, tenancy, console, credentials | the CustomResourceDefinition |
The KMetal spec carries only values that have no sane default, plus explicit escape hatches.
The component set, every chart version, and the install order are pinned by the operator release, so they are deliberately absent from the API — upgrading kMetal means bumping the operator.
See Upgrades.
Chart values¶
helm upgrade --install kmetal-operator oci://ghcr.io/clastix/charts/kmetal-operator \
--namespace kmetal-operator-system --create-namespace \
--values values.yaml
| Key | Default | Notes |
|---|---|---|
replicaCount |
1 |
More than one needs leaderElection.enabled; buys failover, not throughput. |
image.repository |
ghcr.io/clastix/kmetal-operator |
|
image.tag |
chart appVersion |
Overriding it changes which component set and chart versions get installed. |
image.pullPolicy |
IfNotPresent |
|
imagePullSecrets |
[] |
Also used by the CRD hook, which runs the same image. |
rbac.create |
true |
The role is broad by necessity — the operator installs arbitrary platform components. |
serviceAccount.create / .name / .annotations |
true / "" / {} |
|
crds.install |
true |
pre-install,pre-upgrade hook running kmetal install-crds. |
crds.timeout |
2m |
How long the hook waits for the definitions to be Established. |
kmetal.name |
kmetal |
Name of the singleton KMetal object. |
kmetal.spec |
{} |
The KMetal spec, passed through verbatim. Empty installs the operator alone. |
kmetal.singleInstance |
false |
Passes --instance-name. Turn on where a cluster carries more than one object. |
kmetal.keepOnUninstall |
true |
Leave on: every component is owned by the object, so deleting it uninstalls the platform. |
leaderElection.enabled |
true |
Required above one replica, harmless with one. |
logLevel |
info |
1 adds ownership, credential distribution and install progress; 2 adds every object applied. |
metrics.port / metrics.service.* |
8080 / ClusterIP |
|
healthProbe.port |
8081 |
|
resources |
100m/128Mi requests, 512Mi limit |
No CPU limit on purpose — throttling a controller mid-reconcile looks like a stalled install. |
nodeSelector |
node-role.kubernetes.io/control-plane: "" |
Clear it on a cluster whose control plane is not labelled this way, or the Pod sits Pending with no other symptom. |
tolerations |
control-plane NoSchedule |
A no-op on an untainted control plane; what makes the default nodeSelector work on a tainted one. |
affinity, topologySpreadConstraints, priorityClassName |
{}, [], "" |
|
podSecurityContext, securityContext |
non-root, read-only rootfs, all caps dropped | |
podAnnotations, podLabels, commonLabels |
{} |
|
nameOverride, fullnameOverride |
"" |
|
extraArgs |
[] |
Appended to the operator command. |
No delete path anywhere
Removing the KMetal CustomResourceDefinition removes the KMetal object, and every component is owned by that object — so tidying up the CRD tears down the whole undercluster. The chart grants the CRD hook no delete, and annotates the object with helm.sh/resource-policy: keep, so helm uninstall removes the operator without removing the platform.
kmetal.spec is deliberately not modelled in values.yaml: the CRD already carries the schema, the defaults and the validation, and a second copy would only be a place for the two to disagree. What goes there is exactly what goes under spec: on the object.
KMetal spec¶
spec.networking¶
Required. The undercluster network substrate.
networking.kubeOVN¶
Kube-OVN runs as a non-primary CNI: the undercluster's primary CNI is a prerequisite you install beforehand, and Kube-OVN never takes over pod eth0. What it adds is the tenant-facing substrate — VPCs, EIPs, SNAT through the OVN distributed gateway.
| Field | Required | Default | Notes |
|---|---|---|---|
podCIDR |
yes | — | The ovn-default subnet, not the cluster pod network. Must not overlap what the primary CNI owns. |
podGateway |
no | first usable address of podCIDR |
|
tunnelInterface |
yes | — | Node interface carrying the overlay, typically a kernel VLAN sub-interface (bond0.205). |
tunnelType |
no | geneve |
geneve | vxlan | stt |
centralNodes |
yes | — | Nodes running ovn-central and kube-ovn-controller, forming the OVSDB Raft cluster. |
centralNodes is resolved to each Node's InternalIP and passed explicitly rather than discovered by label — discovery races node labelling during a cold bootstrap and yields an empty cluster. Use an odd count so the cluster keeps quorum on node loss; 3 tolerates one failure. Every node listed must carry tunnelInterface.
The invariants that keep the non-primary split safe — NON_PRIMARY_CNI, and the CNI config priority that stops Kube-OVN becoming Multus' default delegate — are pinned by the release and are not fields. Getting them wrong takes the cluster's pod network down.
networking.providerNetworks¶
Required, 1–32 entries. The underlay segments tenant VPCs egress through. A VPC that names no provider network egresses through defaultProviderSubnet; declaring more gives a tenant a segment of its own, which is how per-tenant egress isolation is done.
| Field | Required | Default | Notes |
|---|---|---|---|
name |
yes | — | How tenants reference this network. |
interface |
yes | — | Taken wholesale into OVN's provider bridge, so it must carry no node IP. Never the interface the node is managed through. |
autoCreateVLANSubinterface |
no | false |
Has the kMetal networking daemon create interface on each participating node instead of expecting the host to have it. |
excludeNodes |
no | [] |
Every node without the physical NIC behind interface — typically the control plane. A node left in without the NIC gets an OVN bridge with nothing behind it. |
subnets |
yes | — | 1–32 underlay ranges carried on this network. |
Each entry in subnets:
| Field | Required | Default | Notes |
|---|---|---|---|
name |
yes | — | The Kube-OVN Subnet name. Kube-OVN additionally requires a subnet literally called external to exist before it reconciles external access on tenant VPCs. |
vlanID |
no | 0 |
0–4094. |
cidrBlock |
yes | — | Canonical CIDR. Every tenant VPC has egress to it blocked automatically. |
gateway |
yes | — | Next hop off this segment, normally the edge firewall. |
excludeIPs |
no | [] |
Up to 32 entries, each a single address or an inclusive first..last range. |
Do not double-tag, and do not under-declare excludeIPs
Leave vlanID at 0 when interface is a kernel VLAN sub-interface, which already tags — tagging again at the localnet port double-tags the frame and the traffic is silently dropped. Set it only when handing OVN an untagged interface and asking it to do the tagging.
excludeIPs holds addresses back from OVN's IPAM. The gateway belongs there, as does anything already assigned on the segment — addresses reserved for DNAT to tenant control planes, or for reaching the undercluster itself. Omitting an address that is in use lets OVN hand it to a tenant, which takes out whatever was using it.
Remaining networking fields¶
| Field | Required | Default | Notes |
|---|---|---|---|
defaultProviderSubnet |
yes | — | A Kube-OVN Subnet name declared under providerNetworks, not a provider network name. The subnet a tenant VPC egresses through when it names none itself. |
serviceCIDR |
yes | — | Must match the --service-cluster-ip-range the control plane was bootstrapped with. Canonical (masked) CIDR. Kube-OVN needs it to keep service traffic off the overlay. |
extraBlockedEgressCIDRs |
no | [] |
Up to 64 CIDRs dropped from every tenant VPC's egress in addition to the provider subnets, which are blocked automatically. Use it for ranges that are not provider networks but that tenants still must not reach. |
loadBalancer.tenantControlPlanes.addresses |
yes | — | Pool backing the Services that expose tenant control planes. |
loadBalancer.management.addresses |
no | — | Pool backing operator-facing Services, such as the console. Without it no management pool is created and any Service requesting one stays Pending. |
Each address entry is either a CIDR (192.0.2.0/24) or an inclusive start-end range (192.0.2.10-192.0.2.20).
The two pools are separate because they sit on different network segments with different exposure; collapsing them would let tenant control-plane VIPs and operator VIPs consume each other's space. Keep the tenant range on a segment dedicated to tenant traffic: putting it on the provider segment — the one carrying OVN's external gateway — would give tenants a path to SNAT egress and to each other's traffic.
Services opt into a pool by annotation:
spec.storage¶
Required. kMetal installs no CSI driver, provisioner or StorageClass, and manages no storage cluster — bring your own, the same way you bring your own primary CNI, and name it here.
| Field | Required | Default | Notes |
|---|---|---|---|
underclusterClassName |
yes | — | The platform's own volumes: the golden OS images, and the root disk of every tenant machine cloned from them. Also the Kamaji datastore, unless etcdClassName names another class. |
etcdClassName |
no | underclusterClassName |
The Kamaji datastore holding every hosted control plane's etcd data. |
tenantClassName |
yes | — | The class tenant workloads get by default. |
tenantClaimPropertySets |
no | [] |
Pins how CDI provisions volumes on tenantClassName, as a CDI StorageProfile. |
underclusterClassName must support VolumeSnapshots: the image builder snapshots the golden image volume and CDI clones each machine's root disk from it. A class that cannot snapshot leaves the OSArtifact with no golden PVC and every tenant machine with no disk to boot — which rules out most node-local provisioners here.
etcdClassName exists for exactly that reason. The two platform consumers want opposite things and only one of them can be satisfied by an arbitrary class: the class above has to snapshot, which the low-latency node-local classes generally cannot, while etcd wants that latency and can make no use of a snapshot even where the class offers one. Leave it unset on a cluster with a single class that does both. Setting it to a node-local class means the etcd data of every hosted control plane is lost with the node — a deliberate trade of tenant control-plane durability for latency, and one to make here rather than inherit from what the golden images needed.
tenantClassName should be network-attached and survive the loss of any single node. Name a class of its own even where one backend serves both it and underclusterClassName: a tenant's storage quota is keyed on the class name, and machine root disks are provisioned in the tenant's own namespace — so one shared name has the platform's volumes counted against the tenant's workload budget. See Keep the platform class and the tenant class apart.
Each entry in tenantClaimPropertySets:
| Field | Required | Default | Notes |
|---|---|---|---|
accessModes |
yes | — | Access modes the class can satisfy for this combination. |
volumeMode |
no | Filesystem |
Filesystem | Block |
Order is significant
CDI takes the first entry for a DataVolume that does not ask for anything specific. Leave the list empty on a single-capability backend, where auto-derivation can only produce one answer. Set it on anything advertising more than one combination: a backend that advertises ReadWriteMany + Block first will have CDI provision raw-block volumes, and the importer pod then fails on a permission error naming neither the storage class nor the volume mode. Set it before the first tenant volume, not after diagnosing that error.
spec.clusterClass¶
Optional passthrough to the ClusterClass chart tenant clusters are created from.
| Field | Required | Notes |
|---|---|---|
values |
no | Merged over the values kMetal computes, and wins on conflict. |
kMetal computes only what it alone can know: where it installed the image builder, that the golden images and the machine root disks cloned from them both belong on underclusterClassName (osArtifact.storageClassName and dataVolume.storageClassName), and that tenant machines attach to the Kube-OVN overlay. Everything else — which tiers to offer, the worker image and its tag, DNS servers, disk sizes, the VolumeSnapshotClass used to clone boot volumes — is yours to set here. The chart is not mirrored into this API because its knobs move between versions, and every bump would break a mirror.
Two chart values have to be set here or tenant machines will not boot
Neither is among the values kMetal can compute:
dataVolume.volumeSnapshotClass— must serve the same driver asstorage.underclusterClassName, which is the class the golden image is snapshotted on. The chart ships no default for it from v1.10.0 and refuses to render without one.- the kubeadm tier's
containerDisk.tag— must match the Kubernetes version tenantClusters ask for.
Because these win on conflict, this is a genuine escape hatch: it can override computed values too, including ones holding the platform together.
spec.multiTenancy¶
Optional. Which identities Capsule enforces tenancy for.
Opt-in throughout. Every field is empty unless you set it, and kMetal fills in nothing on your behalf — omit the block and Capsule is installed enforcing nothing.
| Field | Required | Default | Notes |
|---|---|---|---|
users |
no | [] |
Subjects Capsule matches a Tenant owner against. Up to 64. |
administrators |
no | [] |
Subjects treated as an owner of every Tenant, present and future. Up to 64. |
ignoreUserWithGroups |
no | [] |
Requests carrying one of these groups skip Capsule entirely, whatever users says. Up to 64. |
allowServiceAccountPromotion |
no | false |
Lets a Tenant owner turn a ServiceAccount in one of the tenant's own namespaces into an owner of that tenant. |
Each entry in users and administrators:
| Field | Required | Notes |
|---|---|---|
kind |
yes | User | Group | ServiceAccount |
name |
yes | The name as the API server authenticates it, which is not always the name of an object. |
Nothing here has a default, and that is deliberate. Every entry entitles a subject, and who owns a tenant is the identity provider's answer rather than the platform's — a subject a kMetal release put here would be an authorization in force on every cluster that nobody asked for, expressed where the spec has no way to say no.
A Tenant owner missing from users fails silently
Listing a subject grants it nothing — a Tenant still has to name it as an owner — but it is the precondition for enforcement, and the failure runs the other way.
A Tenant owner absent from users is invisible to Capsule: their namespaces are created outside every tenant, with none of its quotas, NetworkPolicies or constraints applied.
Nothing rejects the Tenant, and nothing reports the namespace as unmanaged.
Nothing is added to what you set, so users is the whole set Capsule matches against.
Extend it when a new owner kind arrives, rather than after diagnosing an Environment that took no quota.
Writing a subject's name¶
| Kind | name is |
|---|---|
User |
The username the authenticator presents — for OIDC, the claim the API server is configured to read. |
Group |
The group name the authenticator presents. Every ServiceAccount in one namespace is covered by system:serviceaccounts:<namespace>. |
ServiceAccount |
system:serviceaccount:<namespace>:<name>. The bare name matches nothing. |
ServiceAccount-owned Tenants are the case worth stating, because they are how kMetal's own tooling owns a tenant and they are matched through a group rather than as themselves — the group covering whichever namespace those ServiceAccounts live in, listed like any other subject:
kmetal-tenants is a convention, not a platform name: kMetal neither creates that namespace nor adds its group for you.
administrators¶
Cluster-wide by construction, so it is for platform operators and GitOps ServiceAccounts — never for anyone who also owns a tenant of their own.
Reaching a tenant is still explicit: Capsule only reads a namespace operation as tenant-related when the request carries the Capsule tenant label, so an administrator who omits it is operating outside tenancy rather than on every tenant at once.
ignoreUserWithGroups¶
For the identity provider that puts every account in one group.
That group goes in users so tenant owners are matched, which also catches the platform team, and naming their own group here is what separates the two.
Without it a cluster administrator inherits tenant admission on every namespace they touch.
allowServiceAccountPromotion¶
Off by default because it delegates entitlement downwards: the promoted ServiceAccount becomes a tenant owner that nothing in this spec names, granted by whoever already owns the tenant. Capsule stops the chain at one step — a promoted ServiceAccount cannot promote another — but that step is the tenant's decision rather than the platform's.
Only entitlement is exposed
Namespace naming policy — Capsule's forceTenantPrefix and protectedNamespaceRegex — stays pinned by the release.
It changes the rule namespaces are admitted under rather than who is subject to it, so flipping it on a running platform starts rejecting namespace names that were accepted the day before.
spec.console¶
Optional. Present means install it; omit the block entirely to install no console. Nothing else depends on it.
| Field | Required | Default | Notes |
|---|---|---|---|
address |
no | assigned from the management pool | Must fall inside networking.loadBalancer.management. kMetal cannot check that, and an address outside the pool leaves the Service Pending. |
tls |
no | — | Omitting it serves plain HTTP. |
tls.issuerRef.name |
no | kmetal-ca |
|
tls.issuerRef.kind |
no | ClusterIssuer |
Issuer | ClusterIssuer. Defaults away from cert-manager's own Issuer default: a namespaced Issuer would have to live in the console's namespace, which kMetal creates. |
tls.issuerRef.group |
no | cert-manager.io |
|
tls.dnsNames |
no | [] |
address is added automatically as an IP SAN. Add any DNS record the console is reached by — a certificate is only valid for the names it carries. |
kmetal-ca is created by kMetal alongside cert-manager, so TLS works without bringing an issuer. It is trusted by nothing until you distribute its root; name your own issuer to be signed by something the wider network already trusts.
Serving the console over plain HTTP puts a cluster-admin-capable session token on the wire in clear text — reasonable on a lab bench, not much else.
Console cluster requirements
The kMetal console plugin ships as an image volume, pinned by the release. That needs Kubernetes 1.33+ (where image volumes are beta and on by default), or 1.31+ with the ImageVolume feature gate enabled on the kube-apiserver and every kubelet — plus containerd 2.0+ or CRI-O 1.31+. Without those the console pod does not start. Omit this block on such a cluster; nothing else in the platform has that requirement.
spec.registry¶
Optional. Credentials for the registries hosting kMetal's own images and charts.
| Field | Required | Notes |
|---|---|---|
imagePullSecretName |
no | A kubernetes.io/dockerconfigjson Secret. Used by kubelet to pull component images, surfaced to charts as an imagePullSecrets entry, and referenced by the Flux HelmRepository and OCIRepository sources for the private kMetal charts and artifacts. |
chartPullSecretName |
no | An Opaque Secret carrying OCI_USERNAME and OCI_PASSWORD, or OCI_ACCESS_TOKEN. Used as the Cluster API operator's configSecret when it fetches the Kamaji control-plane provider from the private OCI registry. |
Both are referenced, never inlined: they hold credentials, and a cluster-scoped spec is readable by anyone with read access to the object.
Both Secrets are read from the namespace the operator itself runs in, and copied from there into every namespace the platform installs into — a Secret is only usable from its own namespace. The operator's own namespace is the one guaranteed to hold them on a cluster installed from scratch: the operator's image comes from the same private registry, so the credential is already there before the operator starts.
Namespaces that do not exist yet are skipped rather than created, since each component's namespace is created by its own install. A missing source Secret is not an error either — a cluster pulling only public charts and images needs neither, and failing the reconcile over an unused credential would block an otherwise healthy install.
The two are distinct because the consumers want different shapes: kubelet and Flux want a dockerconfigjson, the Cluster API operator wants OCI_* keys.
spec.kubevirt¶
Optional. Whitelists PCI and mediated (mdev/vGPU) devices tenant VMs may request through a ClusterClass gpus variable entry. KubeVirt denies every host device by default; a device absent from here has nothing for a ClusterClass deviceName to match, and the VM's Pod sits Pending.
| Field | Required | Notes |
|---|---|---|
permittedHostDevices |
no | Real PCI devices. Each entry: pciVendorSelector (the device's vendor:device hex pair, e.g. 10de:2236 from lspci -nn on the node) and resourceName (what a ClusterClass gpus entry's deviceName matches against). |
mediatedDevices |
no | VFIO mediated devices (mdev/vGPU types). Each entry: mdevNameSelector (the mdev type's name, exactly as reported under mdev_supported_types/*/name on the node) and resourceName. |
kubevirt:
permittedHostDevices:
- pciVendorSelector: "10de:2236"
resourceName: "nvidia.com/A10"
mediatedDevices:
- mdevNameSelector: "nvidia-222"
resourceName: "nvidia.com/GRID_T4-2Q"
KubeVirt tracks two separate settings for mediated devices: which types are whitelisted for scheduling, and which types virt-handler itself keeps alive on nodes. Only the first is a field here; the second (mediatedDevicesConfiguration.mediatedDeviceTypes) is derived automatically from mediatedDevices, since there is no real case for whitelisting a type without also wanting it kept running. See GPU & Device Passthrough for the end-to-end setup, including the node-side prerequisites (IOMMU, vfio-pci) neither this CR nor KubeVirt can satisfy on their own.
spec.placement¶
Optional. Constrains where kMetal's platform components run. Omitted entirely, the defaults are exactly what is shown below.
placement:
nodeSelector:
node-role.kubernetes.io/control-plane: ""
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
It applies to the singleton controllers — cert-manager and its cainjector, webhook and startupapicheck, Kamaji and its etcd, Capsule, the CAPI operator, Sveltos, the kMetal operators, MetalLB's controller — every component that is control-plane workload rather than per-node workload. The default keeps them off tenant nodes, and the default tolerations are a no-op on an untainted control plane while being what makes the default selector work on a tainted one.
It deliberately does not reach the DaemonSets and per-node plugins: MetalLB's speaker has to answer for a VIP from any node that can carry it, and Kube-OVN's per-node workloads are placed by node label instead, since they need the overlay interface present.
Names that are not configurable¶
Some names are constants because they are a public interface — specs elsewhere reference them by name, so letting them drift would strand everything already pointing at the old one.
| Name | What it is |
|---|---|
kmetal-flux |
Namespace for every Flux object the operator emits, and the storageNamespace for its Helm releases. |
tenant-cp-pool |
MetalLB pool tenant control-plane Services draw from. |
mgmt-pool |
MetalLB pool operator-facing Services draw from. |
kmetal-ca |
The ClusterIssuer kMetal creates and issues from. Bring a different issuer by naming it explicitly instead. |
CDI's filesystem overhead is pinned to 0 on tenantClassName for the same class of reason, and is not a field: the default 6% reservation pushes a whole-GiB DataVolume just past the next GiB boundary, and a driver that rounds up from there — Ceph RBD does — provisions twice what the tenant asked for. Only that one class is touched.
Worked example¶
apiVersion: clastix.io/v1
kind: KMetal
metadata:
name: kmetal
spec:
networking:
kubeOVN:
# The ovn-default subnet — NOT the cluster pod network, which the
# primary CNI owns. The two must not overlap.
podCIDR: 10.16.0.0/16
podGateway: 10.16.0.1
# Kernel VLAN sub-interface carrying the Geneve overlay.
tunnelInterface: bond0.205
tunnelType: geneve
# Resolved to InternalIPs. Three members so OVSDB Raft keeps quorum
# through the loss of one. The control-plane node is excluded: it has
# no provider NIC and no tunnel interface.
centralNodes:
- kmetal-worker-01
- kmetal-worker-02
- kmetal-worker-03
# Must match the apiserver's --service-cluster-ip-range.
serviceCIDR: 10.96.0.0/16
providerNetworks:
# Shared default — direct routing, public addresses.
- name: provider
# OVN takes this interface wholesale into its provider bridge, so it
# must carry no node IP.
interface: bond0.206
excludeNodes:
- kmetal-controller # no provider NIC
subnets:
- name: external
# Zero: bond0.206 already tags at the kernel, and tagging again
# double-tags at the localnet port and drops the traffic.
vlanID: 0
cidrBlock: 198.51.100.0/27
gateway: 198.51.100.1
excludeIPs:
- 198.51.100.0 # network address
- 198.51.100.1 # gateway
- 198.51.100.2 # undercluster management
- 198.51.100.3 # platform reserve
- 198.51.100.10..198.51.100.20 # DNAT band for tenant CPs
- 198.51.100.31 # broadcast
# Dedicated Edge-NAT segment, one tenant.
- name: provider2
interface: bond0.210
autoCreateVLANSubinterface: true # daemon creates it per node
excludeNodes:
- kmetal-controller
subnets:
- name: external-provider2
vlanID: 0
cidrBlock: 172.21.210.0/24
gateway: 172.21.210.1
excludeIPs:
- 172.21.210.1
# A Kube-OVN Subnet name declared above, not a provider network name.
defaultProviderSubnet: external
# The provider subnets above are blocked for tenant egress automatically.
# This is only for ranges that are not provider networks but that tenants
# still must not reach.
# extraBlockedEgressCIDRs: []
loadBalancer:
# Segment dedicated to tenant traffic.
tenantControlPlanes:
addresses:
- 172.21.208.100-172.21.208.250
# Operator-facing segment.
management:
addresses:
- 172.21.204.200-172.21.204.250
# kMetal installs no storage: bring your own CSI and name its classes here.
storage:
# Golden OS images and the machine root disks cloned from them. Has to
# snapshot, so a node-local class cannot serve this one.
underclusterClassName: platform-storage
# Kamaji etcd for every hosted control plane. Node-local here, which ties
# tenant control-plane durability to the node — and is why it is named
# separately rather than following the class above.
etcdClassName: local-path
# Handed to tenant workloads. A name of its own, so machine root disks are
# not counted against the tenant's own storage quota.
tenantClassName: tenant-storage-class
# Order matters: CDI takes the first entry for a DataVolume that asks for
# nothing specific.
tenantClaimPropertySets:
- accessModes:
- ReadWriteOnce
volumeMode: Filesystem
# Merged over the values kMetal computes, and winning on conflict.
clusterClass:
values:
tiers:
kubeadm:
enabled: true
standard:
enabled: false
expert:
enabled: false
dataVolume:
# Must serve the same driver as storage.tenantClassName.
volumeSnapshotClass: tenant-storage-class-snapshot
storageTiers:
- 30Gi
kubeadm:
containerDisk:
image: quay.io/capk/ubuntu-2404-container-disk
# Must match the Kubernetes version tenant Clusters request.
tag: v1.34.1
# Who Capsule enforces tenancy for. Opt-in throughout: kMetal names nobody,
# so Capsule installed without this block enforces nothing.
multiTenancy:
# A Tenant owner missing from here is invisible to admission — its
# Environments are created outside every tenant, with no quota applied,
# and nothing complains. Names are as the API server authenticates them.
users:
# The ServiceAccounts that own Tenants, matched through the group
# covering their namespace. Nothing adds this for you.
- kind: Group
name: system:serviceaccounts:kmetal-tenants
# The identity provider group tenant owners come from.
- kind: Group
name: kmetal:tenant-owners
# Owners of every Tenant, present and future. Platform operators and
# GitOps ServiceAccounts, never someone who also owns a tenant.
administrators:
- kind: User
name: kubernetes-admin
# Requests carrying one of these groups skip Capsule entirely, whatever
# users says. This is what keeps the platform team out of tenant
# admission when one provider group covers everybody.
# ignoreUserWithGroups:
# - kmetal:platform-team
# Off: it lets a Tenant owner promote a ServiceAccount in one of their own
# Environments into an owner of that tenant.
# allowServiceAccountPromotion: false
# Omit this block entirely to install no console.
console:
# Must fall inside networking.loadBalancer.management above.
address: 172.21.204.200
tls:
issuerRef:
name: kmetal-ca
kind: ClusterIssuer
group: cert-manager.io
# The address above is added automatically as an IP SAN.
# dnsNames:
# - console.kmetal.example.com
# Referenced, never inlined. Both Secrets live in the operator's namespace.
registry:
imagePullSecretName: clastix-ghcr
chartPullSecretName: clastix-ghcr-oci
# Omitted entirely, the defaults are exactly these.
placement:
nodeSelector:
node-role.kubernetes.io/control-plane: ""
tolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
The same spec goes verbatim under kmetal.spec in the chart values when the object is created by the chart rather than applied separately.
Validating a spec¶
The CRD carries the schema, the defaults and the validation, so the API server is the authority:
kubectl explain kmetal.spec --recursive # the field reference
kubectl apply --dry-run=server -f kmetal.yaml # validation without installing
To see exactly what the operator would apply, render the catalogue instead of installing it:
kmetal render -f kmetal.yaml --from-cluster -o install.yaml
kmetal render -f kmetal.yaml --node worker-01=10.0.0.1 --node worker-02=10.0.0.2
The renderer reads the same component catalogue as the operator, so a render is what the operator would apply. Kube-OVN needs the address behind each central node, which a render has no cluster to look up — pass --from-cluster or a --node name=address for each. What a render does not reproduce is the operator's behaviour: applying the file is a single pass, with no readiness gating, no halt on failure and no drift correction.
Watching an install¶
$ kubectl get km
NAME VERSION FLUX READY AGE
kmetal v1.0.0 v2.8.8 True 4h
$ kubectl get km kmetal -o jsonpath='{range .status.components[*]}{.name}{"\t"}{.phase}{"\t"}{.message}{"\n"}{end}'
The status carries one entry per component in install order, each with an installed and a desired version. The reconciler walks the order front to back and halts at the first component that is not Ready; everything after it reports Pending with the name of what it is waiting on. Component phases are Pending, Installing, Upgrading, Ready and Failed — Upgrading is distinct from Installing because a failure there leaves a working older version behind, whereas a failed install leaves nothing.
Two conditions are worth gating scripts on:
FluxReady— component zero. Flux is applied by the operator itself, and every other component is a Flux object, so nothing progresses while this is false.Ready— every component reconciled at the version this release pins.status.versionadvances at the same moment, so a partially applied upgrade never reports as the new release.
See Also: Install kMetal, Upgrades, API Reference for the tenant-facing Cluster API surface.