Multi-Tenancy¶
Isolating tenants on the under cluster.
Tenant Isolation Model¶
kMetal provides tenant isolation at multiple layers:
- Control Plane Isolation: Each tenant has a dedicated Kubernetes control plane (TenantControlPlane) with its own etcd.
- Compute Isolation: Tenant workers run as KubeVirt VMs on the under cluster — kernel-level isolation via KVM.
- Network Isolation: Each tenant operates in an isolated OVN VPC; cross-tenant traffic is dropped by default.
- Namespace Isolation: Tenant control plane and platform resources are separated by Kubernetes namespaces on the under cluster.
- Resource Isolation: ResourceQuotas and LimitRanges on the under cluster prevent resource exhaustion across tenants.
- RBAC Isolation: Role-based access control per tenant on the under cluster.
Layers 4 to 6 are not objects you write.
They are fields on a Capsule Tenant, and Capsule materialises them into every Environment the tenant owns — which is why this page is about the Tenant and not about namespaces.
Write the policy on the Tenant, not in the namespace¶
An Environment is a namespace Capsule owns.
Writing a ResourceQuota, a RoleBinding or a NetworkPolicy into one by hand puts you in competition with the controller that reconciles it, and it puts the object in a namespace whose owner holds admin anyway.
Set the field instead:
| Reaching for | Set this |
|---|---|
| A namespace with labels | spec.namespaceOptions.additionalMetadataList — labels and annotations on every Environment |
ResourceQuota |
spec.resourceQuotas with scope: Tenant, aggregated across Environments — see Cap what a tenant can consume |
LimitRange |
spec.limitRanges, or a replication (below) — the field is deprecated in favour of replications |
RoleBinding |
spec.owners[].clusterRoles and spec.additionalRoleBindings — see RBAC |
NetworkPolicy |
A replication, so the same baseline lands in every Environment (below) |
| Pod Security Standards | Labels in additionalMetadataList; the standard is enforced by the api-server, not by Capsule |
| Allowed classes | spec.storageClasses, spec.priorityClasses, spec.runtimeClasses — allowlists Capsule enforces at admission |
| Node placement | spec.nodeSelector (below) |
| Namespace count | spec.namespaceOptions.quota |
Nothing stops a platform administrator from creating those objects directly. What stops it being useful is that the tenant's next Environment will not have them.
A baseline NetworkPolicy in every Environment¶
spec.networkPolicies on the Tenant still works and is deprecated.
The current mechanism is a replication — a GlobalTenantResource selecting tenants, whose rawItems are applied into each of their Environments and re-applied on drift:
apiVersion: capsule.clastix.io/v1beta2
kind: GlobalTenantResource
metadata:
name: tenant-baseline-netpol
spec:
resyncPeriod: 60s
settings:
adopt: false
tenantSelector: {} # every Tenant; narrow it per customer tier
resources:
- rawItems:
- apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: tenant-baseline
spec:
podSelector: {}
policyTypes: [Ingress, Egress]
ingress:
# Same tenant only.
- from:
- namespaceSelector:
matchLabels:
capsule.clastix.io/tenant: "{{ tenant.name }}"
egress:
- to:
- namespaceSelector:
matchLabels:
capsule.clastix.io/tenant: "{{ tenant.name }}"
# DNS.
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: kube-system
ports:
- { protocol: UDP, port: 53 }
- { protocol: TCP, port: 53 }
# kMetal's admission webhook — see the warning below.
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: kmetal-system
ports:
- { protocol: TCP, port: 443 }
{{ tenant.name }} is Capsule's own templating, not Helm's — the available keys are tenant.name and namespace, and they are substituted per Environment as the item is applied.
Egress to kmetal-webhook.kmetal-system is not optional
A tenant cluster's api-server runs as pods in the tenant's Environment, and kMetal pushes a ValidatingWebhookConfiguration into every tenant cluster pointing back at the kmetal-webhook Service in kmetal-system on the under cluster.
It is registered with failurePolicy: Fail, so if those pods cannot resolve and reach kmetal-webhook.kmetal-system.svc:443:
- tenants cannot create
PersistentVolumeClaims at all — the storage quota check is what that webhook does, so blocking it does not relax the quota, it denies the request, LoadBalancerServices withloadBalancerClass: kmetalare denied too, since the same webhook resolves theEipClaimbinding.
A deny-all-egress baseline without that rule looks like a working platform until the first tenant asks for a volume. Test it from a control-plane pod in an Environment rather than assuming.
Two more egress destinations to keep in mind before tightening further: the DNS rule above, and whatever the tenant's own workloads need out of the Environment. Everything else about a tenant cluster's own traffic is Kube-OVN's job, not this policy's — see Networking.
Confine a tenant to a pool of nodes¶
spec.nodeSelector puts the scheduler.alpha.kubernetes.io/node-selector annotation on every Environment, and Capsule refuses a tenant's attempt to change or remove it:
apiVersion: capsule.clastix.io/v1beta2
kind: Tenant
metadata:
name: dynamo
spec:
nodeSelector:
kmetal.io/tenant-pool: dynamo
This is the field to reach for when a customer's workloads must land on hardware you chose — a dedicated pool, a licensing boundary, a noisy-neighbour split.
It applies to everything in the Environment, which on this platform means both the hosted control-plane pods and the virt-launcher pods backing the tenant's worker VMs.
Two prerequisites, and both fail quietly:
- The nodes must carry the label. Nothing schedules in the Environment otherwise, including the control plane, and the pods sit
Pendingwith no obvious cause.
kubectl label node worker-1 worker-2 kmetal.io/tenant-pool=dynamo
kubectl get nodes -l kmetal.io/tenant-pool=dynamo
- The under cluster's api-server must run the
PodNodeSelectoradmission plugin. The annotation is inert without it, so theTenantlooks configured and pods land anywhere. It is an api-server flag, so this is a decision taken when the under cluster is built.
For the platform's own components the equivalent is spec.placement on the KMetal object, which is unrelated to tenancy — see Component Configuration.
A dedicated etcd datastore per tenant¶
Every hosted control plane keeps its state on a Kamaji DataStore, and by default they share the platform's default.
Sharing is fine until it is not: one datastore is one blast radius, one restore, and one set of credentials.
A per-tenant datastore is an ordinary kamaji-etcd release.
Deploy it as a HelmRelease so Flux owns it like everything else on the under cluster, and put it in a platform namespace rather than in the tenant's Environment — the tenant owner holds admin there and would be able to read the etcd client certificate:
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: etcd-dynamo
namespace: kmetal-flux
spec:
interval: 1h
targetNamespace: kmetal-etcd-dynamo # platform-owned, not the tenant's Environment
install:
createNamespace: true
chart:
spec:
chart: kamaji-etcd
version: 0.17.0 # 0.17.0 or newer
sourceRef:
kind: HelmRepository
name: clastix # already present in kmetal-flux
namespace: kmetal-flux
values:
replicas: 3
datastore:
enabled: true # creates the DataStore object for you
name: dynamo-etcd
# cert-manager instead of the cfssl/kubectl certificate jobs.
selfSignedCertificates:
enabled: false
certManager:
enabled: true
issuerRef:
name: kmetal-ca
kind: ClusterIssuer
group: cert-manager.io
persistentVolumeClaim:
storageClassName: platform-storage # same durability question as the platform's own
size: 8Gi
With certManager.enabled: true the chart provisions the etcd CA, server, peer and client certificates as cert-manager Certificates and points the generated DataStore at those Secrets, instead of minting them once with the cfssl pre-install Jobs.
Certificates then renew on cert-manager's schedule rather than at the age of the release, which is the reason to prefer it — and selfSignedCertificates.enabled: false is what stops the old jobs running alongside.
Use chart 0.17.0 or newer; it is the current release, and pinning it keeps the etcd version and this wiring at a combination that has been tested together.
DataStore is cluster-scoped, so what makes it a tenant's is who is allowed to name it.
Tenants cannot list datastores, and a tenant naming somebody else's would be pointing their own control plane at a store whose schema they do not own — so hand each tenant the name of theirs and set it on the Cluster:
kubectl get datastores
kubectl get tenantcontrolplanes -A -o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,DATASTORE:.status.storage.dataStoreName'
Changing a running control plane's datastore means migrating its etcd data, so this is a decision for onboarding rather than later — see Decide where control-plane state lives.
Verify the isolation¶
# What Capsule enforces for a tenant
kubectl get tenant dynamo -o jsonpath='{.spec}' | python3 -m json.tool
# What landed in an Environment
kubectl get resourcequota,limitrange,networkpolicy -n dynamo-prod
kubectl get namespace dynamo-prod -o jsonpath='{.metadata.annotations}'
# Replications, and whether they applied
kubectl get globaltenantresources,tenantresources -A
# The webhook every tenant cluster depends on
kubectl get svc kmetal-webhook -n kmetal-system
kubectl get endpoints kmetal-webhook -n kmetal-system
An Environment with no ResourceQuota and no NetworkPolicy is usually a namespace Capsule never adopted, which is a tenancy problem rather than a policy one — see the silent failure.