Skip to content

Security Posture

What the architecture rules out, what it leaves to you, and where to point posture tooling.

Tenants share hardware here: the same hosts, the same storage backend, the same provider segment, and one under cluster that manages all of them. The isolation that makes that acceptable is structural rather than configured — so most of it cannot be misconfigured away, and the parts that can are worth knowing by name.

A tenant cluster has no control plane of its own

A tenant's api-server, controller-manager and scheduler are pods on the under cluster, managed by Kamaji. Their nodes are workers and nothing else.

That removes a whole class of exposure from the tenant's machines:

  • No etcd. No data directory, no snapshot on disk, no peer certificates to steal.
  • No cluster CA. The signing key never exists on a tenant node, so a compromised worker cannot mint identities for the cluster it belongs to.
  • No api-server flags to get wrong, and no static pod manifests or /etc/kubernetes/pki for anyone to edit.
  • No control-plane CIS surface. The control-plane half of a benchmark is the platform's responsibility and is not something a tenant can fail.

What a worker does hold is worth being clear about: its own kubelet credentials, the join material, the CNI and CSI node components, and whatever the tenant runs on it. A compromised worker is a node-level identity inside that one tenant cluster.

A tenant cannot reach the platform

Two independent boundaries, and both are default-on.

On the API. A tenant owner is bound into their own Environments and nowhere else — no cluster-scoped reads at all, so they cannot enumerate nodes, datastores, ClusterClasses or other tenants. RBAC has the verified list.

On the network. Every externally-reachable tenant VPC carries an OVN policy route that drops traffic to the platform's provider CIDRs, injected by the controller alongside the default route:

priority <isolation>  match "ip4.dst == <provider CIDRs>"  action drop

The tenant's VpcClaim controls one field — externalAccess — and nothing else about that routing. A VPC without external access has no egress path at all. So a compromised workload cannot pivot from a tenant cluster onto the provider segment, at another tenant's external addresses, or at platform services living there.

The blocked set is the provider subnets, plus what you add

kMetal blocks every provider subnet CIDR automatically, because that list is derivable and forgetting one would silently open a path between tenants.

It does not guess at your management network. If the under cluster's node and api-server addresses are not inside a provider subnet, they are reachable from a tenant VPC with external access until you list them:

spec:
  networking:
    extraBlockedEgressCIDRs:
      - 10.0.0.0/24        # under-cluster nodes and api-server
      - 10.0.10.0/24       # out-of-band management, BMCs, storage admin

This is the single most valuable posture check on the page — see Block a range tenants must not reach.

Between tenants there is nothing to configure: two VPCs have no route to each other, CIDRs may overlap freely, and traffic is Geneve-encapsulated between hosts rather than carried on per-tenant VLANs — see Networking.

The compute boundary is the hypervisor

Tenant workloads run in pods, inside VMs, on shared hosts. A container escape lands the attacker in the tenant's own kernel — a VM they already control — not on the host and not in another tenant's workload.

That is a stronger boundary than node pools give you, and it is why nothing on this platform relies on shared-kernel container isolation between tenants. spec.nodeSelector on a Tenant is still available where a customer needs dedicated hardware for other reasons — see Confine a tenant to a pool of nodes.

The Environment is inside the tenant's trust boundary

An Environment is a namespace on the under cluster, and the tenant owner typically holds admin in it. Their hosted control plane runs there, which means its certificates and kubeconfig Secrets are theirs to read — correctly, since it is their cluster.

The consequence is a placement rule rather than a permission one: nothing the tenant should not have goes in their Environment.

  • A dedicated etcd datastore belongs in a platform namespace, not in the Environment — see A dedicated etcd datastore per tenant.
  • Backup credentials are the deliberate exception: they sit in the Environment because the tenant creates the ClusterBackup, so scope them to that tenant's bucket or prefix.
  • Anything Sveltos delivers into a tenant cluster is likewise theirs to read — see Secrets Management.

Enforcement that survives the tenant

A tenant is cluster-admin of their own cluster, so anything the platform puts inside it is something the tenant can delete. Two mechanisms make that a delay rather than a bypass:

What How it comes back
The CNI, and any other mandatory add-on Sveltos ClusterProfiles run in ContinuousWithDriftDetection, which re-applies on drift
The kmetal-webhook ValidatingWebhookConfiguration in each tenant cluster Its controller watches for tenant-side edits and re-applies immediately, and deliberately reverts narrowing of namespaceSelector, objectSelector or matchPolicy that would let a tenant admit what the webhook denies

That webhook is what enforces the storage quota and the EipClaim binding rules from inside the tenant cluster, which is why it is worth knowing that narrowing it does not work — and why its egress path must stay open, or admission fails closed instead.

What is still yours to configure

Nothing above is a substitute for the under cluster's own posture. It is the trust root: cluster-admin there reaches every tenant's kubeconfig, every datastore and every Secret.

Concern Under cluster — yours Tenant cluster — theirs
Benchmarks (CIS) The whole thing, control plane included Node and workload sections only; the control-plane sections are yours, on their behalf
Encryption at rest Configure it before the first tenant Not exposed
Audit logging Yours, and it covers every tenant action against the platform Not exposed
Pod Security Standards Labels on Environments, via the Tenant Their own namespaces, their own policy
Admission policy beyond Capsule Yours to install Theirs, plus what the platform delivers
RBAC The Tenant's role set Ordinary RBAC on the credential Kamaji issued

Test restricted Pod Security on one Environment first

An Environment holds more than applications: the hosted control-plane pods live there, and so do the virt-launcher pods backing the tenant's worker VMs.

A policy those cannot satisfy does not degrade the tenant — it stops their cluster from having workers. Enforce it on a single Environment with a running cluster before applying it fleet-wide.

Where to run posture tooling

kMetal installs none, and the placement question has a clear answer in each direction:

  • On the under cluster, run what looks at the platform: kube-bench for the CIS benchmark, an image and manifest scanner for what you deploy, a runtime sensor for the hosts. This is where a finding matters most, because it is the trust root.
  • In tenant clusters, deliver the agent with a Sveltos ClusterProfile if you offer it as a platform service — the same mechanism as the CNI, so every cluster gets it including the ones created next month. Expect the control-plane checks to report as not applicable: there is no control plane there to check.
  • Do not expect a tenant-cluster scanner to see the platform. It cannot, by design, which is also why a finding about a missing control-plane hardening flag in a tenant cluster is noise rather than signal.

The limits worth stating plainly

Isolation is not independence. Tenants share failure domains even where they share no access:

  • A host failure takes the workloads that were on it, and the VMs are rebuilt by the machine health check — see React to clusters instead of matching them.
  • The storage backend is shared, so its availability is every tenant's availability.
  • An under-cluster outage stops reconciliation everywhere: running control planes and workloads keep serving, but nothing is created, scaled or repaired until it is back.
  • A shared etcd datastore is a shared blast radius, which is the argument for per-tenant datastores.

Check the posture

# Is the management network actually blocked for tenant egress?
kubectl get km <name> -o jsonpath='{.spec.networking.extraBlockedEgressCIDRs}'
kubectl get vpc -o custom-columns='NAME:.metadata.name,EXTERNAL:.spec.enableExternal,POLICIES:.spec.policyRoutes[*].match'

# What a tenant owner can actually do
kubectl auth can-i --list --as=<tenant owner> -n <environment>
kubectl auth can-i list nodes --as=<tenant owner>

# Enforcement inside tenant clusters
kubectl get clustersummaries -A
kubectl get svc,endpoints kmetal-webhook -n kmetal-system