Skip to content

Multi-Tenancy

Isolating tenants on the under cluster.

Tenant Isolation Model

kMetal provides tenant isolation at multiple layers:

  1. Control Plane Isolation: Each tenant has a dedicated Kubernetes control plane (TenantControlPlane) with its own etcd.
  2. Compute Isolation: Tenant workers run as KubeVirt VMs on the under cluster — kernel-level isolation via KVM.
  3. Network Isolation: Each tenant operates in an isolated OVN VPC; cross-tenant traffic is dropped by default.
  4. Namespace Isolation: Tenant control plane and platform resources are separated by Kubernetes namespaces on the under cluster.
  5. Resource Isolation: ResourceQuotas and LimitRanges on the under cluster prevent resource exhaustion across tenants.
  6. RBAC Isolation: Role-based access control per tenant on the under cluster.

Layers 4 to 6 are not objects you write. They are fields on a Capsule Tenant, and Capsule materialises them into every Environment the tenant owns — which is why this page is about the Tenant and not about namespaces.

Write the policy on the Tenant, not in the namespace

An Environment is a namespace Capsule owns. Writing a ResourceQuota, a RoleBinding or a NetworkPolicy into one by hand puts you in competition with the controller that reconciles it, and it puts the object in a namespace whose owner holds admin anyway.

Set the field instead:

Reaching for Set this
A namespace with labels spec.namespaceOptions.additionalMetadataList — labels and annotations on every Environment
ResourceQuota spec.resourceQuotas with scope: Tenant, aggregated across Environments — see Cap what a tenant can consume
LimitRange spec.limitRanges, or a replication (below) — the field is deprecated in favour of replications
RoleBinding spec.owners[].clusterRoles and spec.additionalRoleBindings — see RBAC
NetworkPolicy A replication, so the same baseline lands in every Environment (below)
Pod Security Standards Labels in additionalMetadataList; the standard is enforced by the api-server, not by Capsule
Allowed classes spec.storageClasses, spec.priorityClasses, spec.runtimeClasses — allowlists Capsule enforces at admission
Node placement spec.nodeSelector (below)
Namespace count spec.namespaceOptions.quota

Nothing stops a platform administrator from creating those objects directly. What stops it being useful is that the tenant's next Environment will not have them.

A baseline NetworkPolicy in every Environment

spec.networkPolicies on the Tenant still works and is deprecated. The current mechanism is a replication — a GlobalTenantResource selecting tenants, whose rawItems are applied into each of their Environments and re-applied on drift:

apiVersion: capsule.clastix.io/v1beta2
kind: GlobalTenantResource
metadata:
  name: tenant-baseline-netpol
spec:
  resyncPeriod: 60s
  settings:
    adopt: false
  tenantSelector: {}                 # every Tenant; narrow it per customer tier
  resources:
    - rawItems:
        - apiVersion: networking.k8s.io/v1
          kind: NetworkPolicy
          metadata:
            name: tenant-baseline
          spec:
            podSelector: {}
            policyTypes: [Ingress, Egress]
            ingress:
              # Same tenant only.
              - from:
                  - namespaceSelector:
                      matchLabels:
                        capsule.clastix.io/tenant: "{{ tenant.name }}"
            egress:
              - to:
                  - namespaceSelector:
                      matchLabels:
                        capsule.clastix.io/tenant: "{{ tenant.name }}"
              # DNS.
              - to:
                  - namespaceSelector:
                      matchLabels:
                        kubernetes.io/metadata.name: kube-system
                ports:
                  - { protocol: UDP, port: 53 }
                  - { protocol: TCP, port: 53 }
              # kMetal's admission webhook — see the warning below.
              - to:
                  - namespaceSelector:
                      matchLabels:
                        kubernetes.io/metadata.name: kmetal-system
                ports:
                  - { protocol: TCP, port: 443 }

{{ tenant.name }} is Capsule's own templating, not Helm's — the available keys are tenant.name and namespace, and they are substituted per Environment as the item is applied.

Egress to kmetal-webhook.kmetal-system is not optional

A tenant cluster's api-server runs as pods in the tenant's Environment, and kMetal pushes a ValidatingWebhookConfiguration into every tenant cluster pointing back at the kmetal-webhook Service in kmetal-system on the under cluster.

It is registered with failurePolicy: Fail, so if those pods cannot resolve and reach kmetal-webhook.kmetal-system.svc:443:

  • tenants cannot create PersistentVolumeClaims at all — the storage quota check is what that webhook does, so blocking it does not relax the quota, it denies the request,
  • LoadBalancer Services with loadBalancerClass: kmetal are denied too, since the same webhook resolves the EipClaim binding.

A deny-all-egress baseline without that rule looks like a working platform until the first tenant asks for a volume. Test it from a control-plane pod in an Environment rather than assuming.

Two more egress destinations to keep in mind before tightening further: the DNS rule above, and whatever the tenant's own workloads need out of the Environment. Everything else about a tenant cluster's own traffic is Kube-OVN's job, not this policy's — see Networking.

Confine a tenant to a pool of nodes

spec.nodeSelector puts the scheduler.alpha.kubernetes.io/node-selector annotation on every Environment, and Capsule refuses a tenant's attempt to change or remove it:

apiVersion: capsule.clastix.io/v1beta2
kind: Tenant
metadata:
  name: dynamo
spec:
  nodeSelector:
    kmetal.io/tenant-pool: dynamo

This is the field to reach for when a customer's workloads must land on hardware you chose — a dedicated pool, a licensing boundary, a noisy-neighbour split. It applies to everything in the Environment, which on this platform means both the hosted control-plane pods and the virt-launcher pods backing the tenant's worker VMs.

Two prerequisites, and both fail quietly:

  • The nodes must carry the label. Nothing schedules in the Environment otherwise, including the control plane, and the pods sit Pending with no obvious cause.
kubectl label node worker-1 worker-2 kmetal.io/tenant-pool=dynamo
kubectl get nodes -l kmetal.io/tenant-pool=dynamo
  • The under cluster's api-server must run the PodNodeSelector admission plugin. The annotation is inert without it, so the Tenant looks configured and pods land anywhere. It is an api-server flag, so this is a decision taken when the under cluster is built.

For the platform's own components the equivalent is spec.placement on the KMetal object, which is unrelated to tenancy — see Component Configuration.

A dedicated etcd datastore per tenant

Every hosted control plane keeps its state on a Kamaji DataStore, and by default they share the platform's default. Sharing is fine until it is not: one datastore is one blast radius, one restore, and one set of credentials.

A per-tenant datastore is an ordinary kamaji-etcd release. Deploy it as a HelmRelease so Flux owns it like everything else on the under cluster, and put it in a platform namespace rather than in the tenant's Environment — the tenant owner holds admin there and would be able to read the etcd client certificate:

apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
  name: etcd-dynamo
  namespace: kmetal-flux
spec:
  interval: 1h
  targetNamespace: kmetal-etcd-dynamo      # platform-owned, not the tenant's Environment
  install:
    createNamespace: true
  chart:
    spec:
      chart: kamaji-etcd
      version: 0.17.0                      # 0.17.0 or newer
      sourceRef:
        kind: HelmRepository
        name: clastix                       # already present in kmetal-flux
        namespace: kmetal-flux
  values:
    replicas: 3
    datastore:
      enabled: true                        # creates the DataStore object for you
      name: dynamo-etcd
    # cert-manager instead of the cfssl/kubectl certificate jobs.
    selfSignedCertificates:
      enabled: false
    certManager:
      enabled: true
      issuerRef:
        name: kmetal-ca
        kind: ClusterIssuer
        group: cert-manager.io
    persistentVolumeClaim:
      storageClassName: platform-storage    # same durability question as the platform's own
      size: 8Gi

With certManager.enabled: true the chart provisions the etcd CA, server, peer and client certificates as cert-manager Certificates and points the generated DataStore at those Secrets, instead of minting them once with the cfssl pre-install Jobs. Certificates then renew on cert-manager's schedule rather than at the age of the release, which is the reason to prefer it — and selfSignedCertificates.enabled: false is what stops the old jobs running alongside. Use chart 0.17.0 or newer; it is the current release, and pinning it keeps the etcd version and this wiring at a combination that has been tested together.

DataStore is cluster-scoped, so what makes it a tenant's is who is allowed to name it. Tenants cannot list datastores, and a tenant naming somebody else's would be pointing their own control plane at a store whose schema they do not own — so hand each tenant the name of theirs and set it on the Cluster:

    variables:
      - name: controlPlane
        value:
          dataStoreName: dynamo-etcd
kubectl get datastores
kubectl get tenantcontrolplanes -A -o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,DATASTORE:.status.storage.dataStoreName'

Changing a running control plane's datastore means migrating its etcd data, so this is a decision for onboarding rather than later — see Decide where control-plane state lives.

Verify the isolation

# What Capsule enforces for a tenant
kubectl get tenant dynamo -o jsonpath='{.spec}' | python3 -m json.tool

# What landed in an Environment
kubectl get resourcequota,limitrange,networkpolicy -n dynamo-prod
kubectl get namespace dynamo-prod -o jsonpath='{.metadata.annotations}'

# Replications, and whether they applied
kubectl get globaltenantresources,tenantresources -A

# The webhook every tenant cluster depends on
kubectl get svc kmetal-webhook -n kmetal-system
kubectl get endpoints kmetal-webhook -n kmetal-system

An Environment with no ResourceQuota and no NetworkPolicy is usually a namespace Capsule never adopted, which is a tenancy problem rather than a policy one — see the silent failure.