Fleet Management Tasks¶
Delivering add-ons into tenant clusters, and keeping the platform itself in Git.
Two mechanisms, pointing in opposite directions:
| Reconciles | Who decides the content | |
|---|---|---|
| Flux | The platform, on the under cluster | kMetal. The operator installs Flux as component zero and owns every object it creates. |
| Sveltos | Add-ons, into tenant clusters | You. kMetal installs Sveltos and ships no profiles of its own. |
The second row is the one that catches people out. Everything a tenant cluster contains beyond a control plane and nodes — starting with its CNI — is a platform decision expressed as Sveltos objects, and there is no default.
Keep the platform definition in Git¶
The KMetal object is the platform's definition.
Treat it like any other production definition — in a repository, reviewed, applied from there.
Two shapes work:
The chart owns it. The spec travels as kmetal.spec in the operator chart's values, and the release is applied from your pipeline:
Applied separately. The chart installs only the operator, with kmetal.spec left empty, and the object is applied by your own GitOps tree:
The second is the better fit when the platform definition should live beside your other cluster manifests rather than inside a Helm release.
Never put credentials in it.
spec.registry names Secrets; the object is cluster-scoped and readable by anyone with read access to it.
Understand what Flux owns¶
Every platform component is a Flux object in the kmetal-flux namespace, and Flux converges them continuously — including correcting drift, so a hand-edit to a component's resources is reverted.
This is not a Flux installation you configure. The operator applies Flux itself as component zero, and owns the sequence the components go in — see Platform Components.
If you already run Flux
The operator detects another controller managing Flux, defers to it, and records the owner in status.flux.managedBy rather than fighting it.
In that state installedVersion may differ from desiredVersion indefinitely, and that is the correct outcome — two controllers reconciling one Deployment towards different versions never converge. Take kMetal out of the other controller's scope if you want it to own the version.
$ kubectl get km <name> -o jsonpath='{.status.flux}'
{"controllers":["source-controller","kustomize-controller","helm-controller","notification-controller"],"desiredVersion":"v2.8.8","installedVersion":"v2.8.8","lastTransitionTime":"2026-08-18T09:40:58Z","phase":"Ready"}
Your own Flux tree is also the natural home for the Sveltos objects below: they are ordinary manifests, and nothing in kMetal creates them.
Understand what Sveltos owns¶
Sveltos matches clusters by label and applies content into them — plain manifests, a Helm chart, a Kustomize overlay, or a URL pointing at somebody else's release YAML — and then keeps applying it.
Two things about how kMetal sets it up shape every profile you write:
Tenant clusters register themselves. Sveltos discovers CAPI Cluster objects natively, so a cluster a tenant created a minute ago is already matched by every profile whose selector fits it. There is no per-cluster wiring, which is the whole reason this and not a Kustomization per cluster.
The under cluster is registered too, as the undercloud SveltosCluster in kmetal-sveltos, labelled kubernetes.io/metadata.name=undercloud.
So a profile can target the platform itself, not only the tenants — which is what makes the event-driven pattern further down work.
kubectl get sveltosclusters -A # the undercloud registration
kubectl get clusterprofiles # what you have declared
kubectl get clustersummaries -A # one per profile × cluster, named <profile>-capi-<cluster>
Not everything goes through Sveltos.
Add-ons that need a per-tenant control loop rather than a set of manifests are dedicated multi-cluster controllers instead: kmetal-webhook delivers itself, the KubeVirt CSI node bundle comes from kubevirt-csi-driver-operator, and cloud-provider-ovn's controllers run inside kmetal-networking's manager.
You write no profiles for those, and when one is missing the answer is in the component's status rather than in a ClusterSummary.
Give tenant clusters a CNI¶
kMetal does not choose a CNI for tenant clusters and does not install one. Kube-OVN is the under cluster's tenant-facing substrate; it is not the pod network inside a tenant cluster.
So a freshly provisioned tenant cluster comes up with its nodes NotReady until something lands a CNI, and that something is a ClusterProfile you wrote.
Which CNI is yours to pick.
Flannel, into every CAPI cluster, is the shape:
apiVersion: config.projectsveltos.io/v1beta1
kind: ClusterProfile
metadata:
name: flannel-cni
spec:
clusterSelector:
matchExpressions:
- key: cluster.x-k8s.io/cluster-name
operator: Exists
values: []
syncMode: ContinuousWithDriftDetection
policyRefs:
- remoteURL:
url: "https://github.com/flannel-io/flannel/releases/download/v0.28.2/kube-flannel.yml"
interval: 10m
patches:
- patch: |-
- op: replace
path: /data/net-conf.json
value: |
{
"Backend": {"Type": "host-gw"},
"EnableNFTables": false,
"Network": "{{ index .Cluster.spec.clusterNetwork.pods.cidrBlocks 0 }}"
}
target:
kind: ConfigMap
name: kube-flannel-cfg
namespace: kube-flannel
version: v1
Worth reading field by field, because it is the pattern for everything else you deliver:
| Field | What it buys you |
|---|---|
clusterSelector |
cluster.x-k8s.io/cluster-name: Exists matches every CAPI cluster, including the ones created next month. This is the mandatory-add-on shape: no opt-in, nothing per-cluster. |
policyRefs[].remoteURL |
The upstream release manifest, fetched as published and re-read every interval. Nothing vendored, nothing repackaged. |
patches |
JSON patches over what was fetched, so upstream stays upstream and your deviation from it stays a readable diff. |
{{ index .Cluster.spec.clusterNetwork.pods.cidrBlocks 0 }} |
The template is instantiated per cluster against that cluster's own CAPI Cluster object. Flannel's Network becomes the pod CIDR CAPI assigned, which is why one profile serves the whole fleet with no per-cluster values to maintain. |
syncMode: ContinuousWithDriftDetection |
Sveltos watches what it delivered and puts it back. A tenant who edits or deletes the CNI gets it re-applied rather than a broken cluster. Continuous re-applies on change only; OneTime delivers and forgets. |
tier (default 100) |
Conflict order when two profiles deploy the same object: lowest tier wins. Untouched, first-to-arrive wins, which is not something to rely on. |
continueOnConflict |
Left false, the first conflict stops the rest of the profile — a loud failure instead of a partial delivery. |
stopMatchingBehavior |
WithdrawPolicies, the default, removes what was delivered when a cluster stops matching. LeavePolicies leaves it behind. |
Two things follow from the CNI being your choice:
- Pre-pull its images into the golden image. A worker whose first boot has to pull a CNI from the internet is the slowest and least reliable part of provisioning — see the image content list.
- The backend has to suit the machine network.
host-gwworks here because a cluster's workers all sit on one Kube-OVN subnet and are directly L2-adjacent; a topology that spreads workers across subnets needs an encapsulating backend instead.
Deliver to a subset¶
What you can select on, on a tenant Cluster:
| Label | Set by | Good for |
|---|---|---|
cluster.x-k8s.io/cluster-name |
CAPI | "every cluster", as above. |
topology.cluster.x-k8s.io/owned |
CAPI, on any cluster using a ClusterClass — so every kMetal cluster | The same thing, narrowed to topology-managed clusters. |
projectcapsule.dev/tenant, capsule.clastix.io/managed-by |
Capsule's mutating webhook, on every namespaced object created in a tenant's namespace | One tenant's clusters. The webhook re-stamps the correct value on every write, so it is not something a tenant can forge or drop. |
| Anything you add yourself | You | Opt-in tiers, canaries, one-off exclusions, or Bring-Your-Own-CNI (exclusion). |
Do not gate a mandatory add-on on a tenant-writable label
The Cluster is the tenant's object.
A selector matching a label the tenant writes lets them opt out of the add-on by editing their own manifest — including out of the CNI.
Select mandatory delivery on what CAPI or Capsule stamps, and keep tenant-set labels for what is genuinely optional.
By default, tenants cannot declare add-ons themselves: no tenant-facing role grants config.projectsveltos.io,
and ClusterProfile is a Cluster-Scoped object, not usable in the Multi-Tenancy design of kMetal.
If Tenants want to take advantage of Sveltos, they can rely on the namespaced Profile in their Environment:
appropriate RBAC must be distributed to them.
Profile is the type that fits tenancy because its reach is its namespace.
It matches only clusters in its own namespace, and its policyRefs may name only ConfigMaps and Secrets there — the API requires their namespace field to be empty and substitutes the Profile's own.
So the grant splits fleet management cleanly: the mandatory tier stays yours as ClusterProfiles, and optional add-ons become something a tenant declares and you never merge.
Two conditions come with the grant, and both are settled on your side:
- Give your mandatory profiles a low
tier. Conflicts resolve to the lowest tier and the default is 100 for everyone, so a platform profile left at the default can lose an object to a tenantProfilethat got there first. deploymentType: Localdelivers into the under cluster. It is a field rather than a resource, so RBAC cannot exclude it — pair the grant with Sveltos' own multi-tenancy (aRoleRequestplus theprojectsveltos.io/serviceaccount-nameand-namespacelabels, which deploy as that ServiceAccount) or with an admission policy that rejectsLocal.
See What a tenant can do on the under cluster for the role itself.
Without the grant, an optional add-on is a Profile you create in their Environment when they ask for it.
React to clusters instead of matching them¶
The event framework is the other half of Sveltos, and it delivers things that are not add-ons at all.
Machine health checks are the worked example: CAPI does not create a MachineHealthCheck for a topology, so the platform has to, once per tenant cluster.
Four objects, and no controller of your own:
- an
EventSourceselecting CAPIClusters that carrytopology.cluster.x-k8s.io/owned, - an
EventTriggerwhosesourceClusterSelectoriskubernetes.io/metadata.name: undercloud— it watches on the under cluster, because that is where theClusterobjects live, - a
ConfigMapannotatedprojectsveltos.io/instantiate: ok, holding a template that ranges over.MatchingResourcesand renders oneMachineHealthCheckper cluster it found, - a
ClusterProfileSveltos generates itself to deliver the rendered result back onto the under cluster.
apiVersion: lib.projectsveltos.io/v1beta1
kind: EventTrigger
metadata:
name: cluster-mhc
spec:
sourceClusterSelector:
matchLabels:
kubernetes.io/metadata.name: undercloud
eventSourceName: cluster-mhc
oneForEvent: false # one rendering over all matches, not one per event
policyRefs:
- name: cluster-mhc
namespace: kmetal-system
kind: ConfigMap
The ConfigMap is the part that does the work.
Its content is a Go template, and the annotation is what tells Sveltos to render it rather than deliver it verbatim:
apiVersion: v1
kind: ConfigMap
metadata:
name: cluster-mhc
namespace: kmetal-system
annotations:
projectsveltos.io/instantiate: ok # without this the template is delivered as text
data:
mhc.yaml: |-
{{- range $cluster := .MatchingResources }}
apiVersion: cluster.x-k8s.io/v1beta2
kind: MachineHealthCheck
metadata:
name: {{ $cluster.Name }}-workers
namespace: {{ $cluster.Namespace }}
spec:
clusterName: {{ $cluster.Name }}
selector:
matchLabels:
cluster.x-k8s.io/cluster-name: {{ $cluster.Name }}
topology.cluster.x-k8s.io/deployment-name: md-0
checks:
nodeStartupTimeoutSeconds: 900
unhealthyMachineConditions:
- type: NodeHealthy
status: "False"
timeoutSeconds: 300
- type: NodeHealthy
status: "Unknown"
timeoutSeconds: 300
---
{{- end }}
.MatchingResources is what the EventSource found, and with collectResources: false each entry carries identity only — apiVersion, kind, name, namespace.
That is enough here, because a MachineHealthCheck is addressed entirely by cluster name and namespace.
Set collectResources: true and the whole objects are collected as well (.Resources in the template), which is what you need if a rendered object has to read a field off the Cluster rather than just its name.
oneForEvent: false on the trigger is what makes the range correct: one rendering over every match, so the document is a ----separated list of every cluster's health check, rather than one delivery per event.
The health check itself is a set of platform decisions, and each is worth a second look before copying it:
| Setting | Reference value | What it means here |
|---|---|---|
nodeStartupTimeoutSeconds |
900 |
How long a Machine may take to produce a Ready node before it is remediated. CAPI defaults to 600. Worker boot includes cloning a golden volume, so the raise is about the storage backend, not the OS — measure a cold clone on your class before lowering it. |
unhealthyMachineConditions |
NodeHealthy False or Unknown for 300s |
Both statuses matter: False is a node reporting itself unhealthy, Unknown is a node that stopped reporting — which is what a lost VM, a lost host, or a lost tenant control plane looks like. |
selector |
deployment-name: md-0 |
Machines are matched by their MachineDeployment name, so this covers a cluster whose worker pool is called md-0 and nothing else. A tenant with a second pool, or a differently-named one, has no health check on it. |
Remediation is uncapped unless you cap it
spec.remediation is absent here, and CAPI reads that as always remediate.
A fleet-wide event — the under-cluster network, a tenant control plane going away — turns every worker NodeHealthy: Unknown at once, and every worker is then deleted and rebuilt at once.
remediation.triggerIf.unhealthyLessThanOrEqualTo: 40% (or unhealthyInRange) is what makes the check stop at a threshold instead, on the reasoning that when most of a cluster looks unhealthy the cluster is usually not the thing that broke.
This is also the point of the immutable-node model paying off: remediation deletes the Machine and CAPI builds a replacement from the same golden image, so a health check that fires is a rebuild rather than an investigation — see Nodes Management.
The profiles this produces are named sveltos-<hash> and carry eventtrigger.lib.projectsveltos.io/* labels naming the trigger that made them — which is the answer when kubectl get clusterprofiles lists something nobody wrote.
Reach for this whenever the platform needs an object per tenant cluster that CAPI does not create for you.
Verify the fleet¶
# The delivery components themselves
kubectl get km <name> -o jsonpath='{range .status.components[*]}{.name}{"\t"}{.phase}{"\n"}{end}' | grep -E 'sveltos|repositories'
# Every tenant cluster, and whether add-ons reached it
kubectl get clusters -A
kubectl get clustersummaries -A
# What actually landed, and whether it is current
kubectl get clustersummary <profile>-capi-<cluster> -n <environment> \
-o jsonpath='{.status.featureSummaries[*].status}{"\n"}{.status.deployedGVKs[*].deployedGroupVersionKind}'
A featureSummaries entry reads Provisioned when the profile is fully applied, and deployedGVKs lists what was created — the quickest way to tell a profile that delivered nothing from one that delivered the wrong thing.
A cluster missing its add-ons entirely is usually a selector that does not match: compare the ClusterProfile's clusterSelector against the Cluster's labels, and check status.matchingClusters on the profile.