Skip to content

Networking Tasks

Two halves, and they meet at the provider network.

The underlay is yours: spec.networking on the KMetal object, where the provider segments and the address pools are declared. Inside a tenant Environment there are exactly three objects, and every tenant networking question is one of them:

Claim Reconciles into Per Environment
VpcClaim the Environment's Kube-OVN Vpc, its static route and its isolation policy routes Exactly one, enforced at admission.
SubnetClaim a Subnet in that VPC, plus the NetworkAttachmentDefinition worker VMs attach through One or more — typically one per cluster.
EipClaim one address on the provider segment: the VPC's egress IP, or an address a workload is published on As many as you grant.

All three have immutable specs, and all three derive their VPC binding rather than declaring it.

For the surfaces these act on, see Networking Configuration. For the model, see Networking Concepts.

The underlay

Every surface in this half is a field of spec.networking on the KMetal object, and every field, default and constraint is in Networking Configuration. What follows is what changes safely on a running platform, and what to look at afterwards.

Add a provider network

Append an entry to spec.networking.providerNetworks and let the operator reconcile:

$ kubectl get providernetworks.network.kmetal.io
NAME        INTERFACE   READY   AGE
provider    bond0.206   True    26d
provider2   bond0.210   True    26d
provider3   bond0.211   True    26d

$ kubectl get km <name> -o jsonpath='{range .status.components[*]}{.name}{"\t"}{.phase}{"\n"}{end}' | grep provider-networks
provider-networks       Ready

Adding one is additive and safe on a running platform. Its CIDR is blocked for every tenant's egress automatically, so no tenant gains a path to it by its existence.

Two fields fail quietly when they are wrong, so re-read them before applying: interface must carry no node IP, and vlanID stays 0 on a kernel VLAN sub-interface that already tags. Both are spelled out under Provider networks.

Block a range tenants must not reach

The provider subnets are blocked automatically. extraBlockedEgressCIDRs covers everything else reachable from the provider segment that tenants must not touch — a corporate range, a metadata endpoint.

Additive, and applied on the next reconcile. It becomes a drop rule on every tenant VPC, so it takes effect for tenants whose VPC was created long before you added the range.

Expose tenant control planes

Tenant control planes are ordinary LoadBalancer Services in the tenant's Environment, owned by Kamaji's TenantControlPlane, and MetalLB answers for them — not Kube-OVN. That is the split to hold onto: MetalLB serves the under cluster, which includes every tenant's api-server endpoint; Kube-OVN serves tenant workloads, through the EipClaims further down this page.

Two pools, declared together under spec.networking.loadBalancer:

Pool Object name Draws from it
tenantControlPlanes tenant-cp-pool The Service in front of each tenant's api-server and konnectivity endpoint. Required.
management mgmt-pool Operator-facing Services — the console, a storage dashboard. Optional: leave it out and no pool is created, so a Service asking for one stays Pending.

The names are fixed, because Services reference them by name. Each pool is rendered as an IPAddressPool plus an L2Advertisement scoped to that pool, both in kmetal-metallb, and always as a pair — a pool without an advertisement assigns addresses that nothing on the network answers ARP for.

$ kubectl get ipaddresspools,l2advertisements -n kmetal-metallb
NAME                                      AUTO ASSIGN   ADDRESSES
ipaddresspool.metallb.io/mgmt-pool        true          ["172.21.204.200-172.21.204.250"]
ipaddresspool.metallb.io/tenant-cp-pool   true          ["172.21.208.100-172.21.208.250"]

NAME                                        IPADDRESSPOOLS
l2advertisement.metallb.io/mgmt-pool        ["mgmt-pool"]
l2advertisement.metallb.io/tenant-cp-pool   ["tenant-cp-pool"]

Layer 2 is the only mode: MetalLB answers ARP from whichever node holds the address, so the ranges have to be on segments the under-cluster nodes are on, and there is no BGP session to configure. The advertisement is scoped by pool name on purpose — an unscoped one advertises every pool, which would put tenant control-plane VIPs on the management segment and the reverse. The speaker runs on every node, and spec.placement deliberately does not reach it: any node may have to answer for a VIP.

Both ranges have to stay off the provider segment, and off each other — the reasons are under Address pools.

Pin or auto-assign a control-plane address

A tenant cluster that names its control-plane address gets it: the Cluster's control-plane service address becomes a metallb.io/loadBalancerIPs annotation on the Service, and MetalLB honours it as long as the address falls inside a pool. It then reports which pool it came from:

$ kubectl get svc -A --field-selector spec.type=LoadBalancer \
    -o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,IP:.status.loadBalancer.ingress[0].ip,POOL:.metadata.annotations.metallb\.io/ip-allocated-from-pool'
NS               NAME                      IP               POOL
dynamo-prod      solar                     172.21.208.100   tenant-cp-pool
photon-prod      quantum                   172.21.208.111   tenant-cp-pool
kmetal-console   kmetal-console-headlamp   172.21.204.200   mgmt-pool

Both pools are auto-assignable, so a Service that pins neither an address nor a pool can be served out of either one. Pin the address on anything whose endpoint is written into a kubeconfig, a certificate SAN, or a DNS record — which is every tenant control plane — or name the pool with the metallb.universe.tf/address-pool annotation, which is how the platform's own Services do it.

Two consequences worth knowing before handing addresses out:

  • The address is in the kubeconfig and in the api-server certificate. Moving a control plane to another address means re-issuing SANs and re-distributing kubeconfigs, so treat an assignment as permanent.
  • A public endpoint is a separate concern. Where tenants reach their control plane on a routable address, that is DNAT at the edge onto this VIP, and the public address has to be in the certificate's SANs. See Control Plane External Access.

Grow an address pool

Add the range to that pool's addresses. Growing is routine, and a second range in the same pool needs nothing else — the advertisement follows the pool, not the range.

Shrinking one below what is already allocated is not — MetalLB will not reclaim an address in use, and the pool ends up inconsistent with what Services hold.

kubectl get ipaddresspools -n kmetal-metallb
kubectl get svc -A --field-selector spec.type=LoadBalancer

Check the second command before shrinking anything.

A Service that will not get an address

kubectl describe svc <name> -n <environment> | tail -20
kubectl logs -n kmetal-metallb -l app.kubernetes.io/component=controller --tail=50
What you see Usually
Pending, no event The pool does not exist — management left out of the spec, or the pools component has not reconciled.
AllocationFailed The pinned address is outside every pool, or the pool is exhausted.
An address assigned, but nothing answers it The advertisement is missing for that pool, or no speaker runs on a node on that segment. Check the speaker DaemonSet.

VpcClaim: the Environment's routing domain

One per Environment, and the Vpc it creates takes the namespace's name. Its spec carries intent and nothing else: whether the tenant may egress, and optionally which segment they egress through.

apiVersion: network.kmetal.io/v1alpha1
kind: VpcClaim
metadata:
  name: main
  namespace: photon-prod
spec:
  externalAccess: true

The static route that makes egress work, and the policy routes that keep this VPC off the provider segment and away from its neighbours, are injected by the controller from your configuration. Neither is writable from the claim, which is the point: a tenant cannot ask for a route.

A second VpcClaim in the same namespace is rejected at admission — only one VpcClaim is allowed per namespace.

Give a tenant its own provider network

A VpcClaim that names no provider network egresses through defaultProviderSubnet from spec.networking. That is a fallback, not a segment tenants are placed on together — nothing is shared out by default. To put a tenant on a segment of its own, its VpcClaim names the provider network:

spec:
  externalAccess: true
  providerNetworkRef:
    name: provider2

That tenant's VPC then egresses through provider2 alone, and their EipClaims draw from that segment's subnet instead of the default one.

Three constraints:

  • providerNetworkRef requires externalAccess: true. A claim with one and not the other is rejected.
  • VpcClaim.spec is immutable. An existing tenant cannot be moved to a different provider network by editing their claim — the claim has to be deleted and re-created, which takes the VPC and everything attached to it. Decide at onboarding, not after.
  • If the provider network declares more than one subnet, the claim must also carry the network.kmetal.io/provider-network-subnet annotation naming which one to use. Admission rejects the claim otherwise, naming how many subnets it found.
$ kubectl get vpcclaim -n photon-prod
NAME   VPC           EXTERNAL   PROVIDERNETWORK   READY   AGE
main   photon-prod   true       provider2         True    44h

Delete one

Deletion is held by a finalizer while SubnetClaims still reference the VPC, and the claim reports SubnetClaimsPresent. Delete the subnets first — which means deleting the clusters attached to them first.

SubnetClaim: IP space inside the VPC

Each claim reconciles into one cluster-scoped Subnet named <namespace>-<claim>, and, unless you turn it off, a NetworkAttachmentDefinition named <claim>-net in the Environment. Worker VMs attach through that NAD, so the subnet is what a tenant cluster actually lands in.

An Environment may hold several. That is how one tenant gets separate dev and prod IP planes inside a single isolated VPC.

apiVersion: network.kmetal.io/v1alpha1
kind: SubnetClaim
metadata:
  name: production
  namespace: photon-prod
spec:
  cidrBlock: 10.100.0.0/24
  gateway: 10.100.0.1          # defaults to the first usable address
  protocol: IPv4               # IPv4 | IPv6 | Dual
  excludeIps:
    - 10.100.0.1
  enableSnatRule: true
Field Notes
cidrBlock Required. Must not overlap another subnet in the same VPC. Tenants may overlap each other freely — isolation is the routing domain, not a global allocation.
gateway Defaults to the first usable host in cidrBlock, and must be inside it.
protocol IPv4 by default.
excludeIps Addresses held back from the subnet's IPAM, up to 256. The gateway belongs here.
attachment.managed On by default: the controller creates the NAD and sets the Subnet's provider to <nad>.<namespace>.ovn. Turn it off only when something else owns the attachment.
attachment.name Overrides the generated <claim>-net NAD name.
enableSnatRule SNATs this subnet's egress through the VPC's router-port EIP. Requires externalAccess: true on the Environment's VpcClaim, since that is what creates the EIP — set on a VPC without it, the rule points at an address that does not exist.

The VPC and the namespaces the subnet serves are derived from the Environment's VpcClaim and are deliberately absent from the spec, so a subnet cannot be attached to a VPC that is not the tenant's.

Size it before applying it

SubnetClaim.spec is immutable, so the CIDR is a one-time decision: re-addressing means deleting the claim, which takes the Subnet and every workload interface on it. Size for the worker VMs the tenant will run, plus what is held back: a /24 leaves around 250 usable once the network and broadcast addresses, the gateway and any excludeIps are out.

Apply the VpcClaim first and let it go Ready. The overlap check resolves the VPC from that claim's status, and is skipped while the Environment has no VPC — so two overlapping SubnetClaims applied into an empty Environment are both admitted, and the collision only surfaces later, in Kube-OVN.

Once a VPC is there, an overlap is refused at admission with the range it collided with:

$ kubectl apply -f second-subnet.yaml
Error from server: admission webhook "vsubnetclaim.kb.io" denied the request:
cidrBlock "10.100.0.0/25" overlaps SubnetClaim "production" (10.100.0.0/24) in namespace "photon-prod"

Add a subnet to an existing tenant

Additive, and safe: a new claim in the same Environment gets its own Subnet and its own NAD, and nothing about the existing subnets changes.

$ kubectl get subnetclaims -n photon-prod
NAME         SUBNET                   VPC           CIDR            READY   AGE
production   photon-prod-production   photon-prod   10.100.0.0/24   True    45h
staging      photon-prod-staging      photon-prod   10.101.0.0/24   True    2m

A tenant Cluster joins one by the claim's own name, through the ClusterClass network.subnet variable — staging. The Kube-OVN Subnet name photon-prod-staging is printed for troubleshooting purposes, and must not be surfaced anywhere else.

Watch capacity

The claim mirrors the Subnet's counters, so you can see a tenant running out without reading Kube-OVN objects:

$ kubectl get subnetclaims -A \
    -o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,CIDR:.spec.cidrBlock,FREE:.status.v4AvailableIPs,USED:.status.v4UsingIPs'
NS            NAME         CIDR            FREE   USED
dynamo-prod   production   10.100.0.0/24   250    3
photon-prod   production   10.100.0.0/24   243    10
photon-prod   staging      10.101.0.0/24   251    2

A subnet close to exhaustion cannot be grown — the spec is immutable. The answer is a second SubnetClaim and new clusters placed on it, which is why it is worth watching before a tenant hits the wall.

The Ready condition is Kube-OVN's own, copied onto the claim, so its reasons come from there and read a little oddly — ResetLogicalSwitchAclSuccess on a healthy subnet is normal. WaitingForVpc is kMetal's, and means what it says.

EipClaim: addresses on the provider segment

An EipClaim is one address on the provider segment, and it is the unit of grant. The claims sitting in an Environment are the addresses that tenant has — a cluster with no free claim to bind cannot publish a LoadBalancer Service at all.

Two kinds, told apart by spec.type:

spec.type What consumes it
lrp The tenant's VPC, as the egress/SNAT address on its router port. The one the VPC uses is the well-known claim named external; any other lrp claim is inert.
nat A LoadBalancer Service in one of the tenant's clusters, DNATed straight to the workload. One Service at a time.

Addresses come from the VPC's external subnet — the provider subnet the VpcClaim resolved to, or defaultProviderSubnet when it named none. Putting a tenant on a provider network of their own therefore also moves where their addresses come from.

Grant one

apiVersion: network.kmetal.io/v1alpha1
kind: EipClaim
metadata:
  name: web
  namespace: photon-prod
spec:
  type: nat
  v4Ip: 198.51.100.24     # optional
Field Notes
type nat (the default) or lrp.
v4Ip Pins a specific address. Left out, Kube-OVN allocates one from the external subnet, minus its excludeIPs. Pin the addresses the edge router or DNS already knows about; let the rest float.
clusterSelector Restricts which of the Environment's Clusters may bind a Service to this claim. Empty or absent matches every cluster in the Environment.

EipClaim.spec is immutable, like the other claims. Re-addressing means deleting the claim and re-creating it, which takes the OVN NAT rules with it.

The external claim

A VpcClaim with externalAccess: true is gated on the claim named external, of type lrp, being Ready. The controller auto-creates it, owned by the VpcClaim, so neither you nor the tenant has to write it.

Pre-create it for one reason only: to pin the tenant's egress address with v4Ip before the VpcClaim gets there. Give it the wrong type and the VPC never becomes Ready — the condition reads EipClaim "external" must have spec.type: lrp.

How a tenant Service picks one up

Their side of it, worth knowing because your grant is what makes it work:

  • The Service carries spec.loadBalancerClass: kmetal. kMetal reconciles nothing else, which leaves any other LoadBalancer implementation in the tenant cluster alone.
  • With the annotation network.kmetal.io/eip-claim: web, it binds to that claim by name.
  • Without it, it binds to the first free Ready nat claim in the Environment, in name order.
  • Bindings are exclusive. A Service naming a claim another Service holds waits, and a warning event lands on the claim; a Service that finds nothing free waits too, with no free EipClaim available in namespace "photon-prod" on the Service.

An lrp claim is never picked up implicitly, so the external address cannot be taken by a workload.

See who holds what

status.binding names the consumer, which makes the Environment's claims an allocation table:

$ kubectl get eipclaims -n photon-prod
NAME       EIP                    V4              BOUND   READY   AGE
external   photon-prod-external   198.51.100.22   true    True    47h
web        photon-prod-web        198.51.100.24   true    True    47h

$ kubectl get eipclaim external -n photon-prod -o jsonpath='{.status.binding.type}{"\n"}'
Vpc

$ kubectl get eipclaim web -n photon-prod \
    -o jsonpath='{.status.binding.service.cluster.name}{" "}{.status.binding.service.service.namespace}{"/"}{.status.binding.service.service.name}{"\n"}'
solar demo/whoami

That is the cluster and the Service inside it, reported from the tenant cluster by the binding controller — you do not need the tenant's kubeconfig to see who is on which address.

The EIP column is the cluster-scoped OvnEip behind the claim, named <namespace>-<claim>. The external claim is the exception: its EIP is <vpc>-<external subnet>, matching Kube-OVN's own convention so it adopts the object as the VPC's gateway EIP.

Revoke one

Deleting a bound claim is rejected — the Service has to be deleted or unannotated first, and a claim attached to a VpcClaim cannot be deleted while that VPC exists. The exception is a binding to a cluster that is already gone, which is not enforced.

If you also let tenants create their own EipClaims, bound them with the count/eipclaims.network.kmetal.io quota rather than by hand — see Multi-tenancy. Either way, your remaining part is the edge router: the provider segment has to be routed to the under cluster for any of these addresses to answer. See Load Balancer for the two integration patterns.

Diagnose a claim that will not become Ready

kubectl get vpcclaims,subnetclaims,eipclaims -A
kubectl describe subnetclaim <name> -n <environment>
Condition reason On Means
WaitingForVpc SubnetClaim The Environment has no Ready VpcClaim yet. Fix the VPC first.
WaitingForEip VpcClaim External access is waiting on the external EipClaim — check it is type lrp and Ready.
ExternalGatewayNotReady VpcClaim The VPC exists, but its external gateway is not up. Look at the provider network and its subnet.
SubnetClaimsPresent VpcClaim Deletion is blocked because subnets still reference the VPC. Delete those, and the clusters on them, first.
EipNotReady EipClaim Kube-OVN has not assigned the address. Usually a v4Ip outside the external subnet or in its excludeIPs, or an exhausted subnet.
EipInUseByVpc EipClaim Deletion is blocked: the claim is the VPC's egress address. The VpcClaim goes first.

Three failures happen at admission instead, so they never reach a condition — the kubectl apply is what carries the message:

  • a second VpcClaim in an Environment,
  • a providerNetworkRef that names no existing provider network, or an ambiguous one,
  • a SubnetClaim CIDR overlapping another subnet in the same VPC.