Networking Tasks¶
Two halves, and they meet at the provider network.
The underlay is yours: spec.networking on the KMetal object, where the provider segments and the address pools are declared.
Inside a tenant Environment there are exactly three objects, and every tenant networking question is one of them:
| Claim | Reconciles into | Per Environment |
|---|---|---|
VpcClaim |
the Environment's Kube-OVN Vpc, its static route and its isolation policy routes |
Exactly one, enforced at admission. |
SubnetClaim |
a Subnet in that VPC, plus the NetworkAttachmentDefinition worker VMs attach through |
One or more — typically one per cluster. |
EipClaim |
one address on the provider segment: the VPC's egress IP, or an address a workload is published on | As many as you grant. |
All three have immutable specs, and all three derive their VPC binding rather than declaring it.
For the surfaces these act on, see Networking Configuration. For the model, see Networking Concepts.
The underlay¶
Every surface in this half is a field of spec.networking on the KMetal object, and every field, default and constraint is in Networking Configuration.
What follows is what changes safely on a running platform, and what to look at afterwards.
Add a provider network¶
Append an entry to spec.networking.providerNetworks and let the operator reconcile:
$ kubectl get providernetworks.network.kmetal.io
NAME INTERFACE READY AGE
provider bond0.206 True 26d
provider2 bond0.210 True 26d
provider3 bond0.211 True 26d
$ kubectl get km <name> -o jsonpath='{range .status.components[*]}{.name}{"\t"}{.phase}{"\n"}{end}' | grep provider-networks
provider-networks Ready
Adding one is additive and safe on a running platform. Its CIDR is blocked for every tenant's egress automatically, so no tenant gains a path to it by its existence.
Two fields fail quietly when they are wrong, so re-read them before applying: interface must carry no node IP, and vlanID stays 0 on a kernel VLAN sub-interface that already tags.
Both are spelled out under Provider networks.
Block a range tenants must not reach¶
The provider subnets are blocked automatically.
extraBlockedEgressCIDRs covers everything else reachable from the provider segment that tenants must not touch — a corporate range, a metadata endpoint.
Additive, and applied on the next reconcile. It becomes a drop rule on every tenant VPC, so it takes effect for tenants whose VPC was created long before you added the range.
Expose tenant control planes¶
Tenant control planes are ordinary LoadBalancer Services in the tenant's Environment, owned by Kamaji's TenantControlPlane, and MetalLB answers for them — not Kube-OVN.
That is the split to hold onto: MetalLB serves the under cluster, which includes every tenant's api-server endpoint; Kube-OVN serves tenant workloads, through the EipClaims further down this page.
Two pools, declared together under spec.networking.loadBalancer:
| Pool | Object name | Draws from it |
|---|---|---|
tenantControlPlanes |
tenant-cp-pool |
The Service in front of each tenant's api-server and konnectivity endpoint. Required. |
management |
mgmt-pool |
Operator-facing Services — the console, a storage dashboard. Optional: leave it out and no pool is created, so a Service asking for one stays Pending. |
The names are fixed, because Services reference them by name.
Each pool is rendered as an IPAddressPool plus an L2Advertisement scoped to that pool, both in kmetal-metallb, and always as a pair — a pool without an advertisement assigns addresses that nothing on the network answers ARP for.
$ kubectl get ipaddresspools,l2advertisements -n kmetal-metallb
NAME AUTO ASSIGN ADDRESSES
ipaddresspool.metallb.io/mgmt-pool true ["172.21.204.200-172.21.204.250"]
ipaddresspool.metallb.io/tenant-cp-pool true ["172.21.208.100-172.21.208.250"]
NAME IPADDRESSPOOLS
l2advertisement.metallb.io/mgmt-pool ["mgmt-pool"]
l2advertisement.metallb.io/tenant-cp-pool ["tenant-cp-pool"]
Layer 2 is the only mode: MetalLB answers ARP from whichever node holds the address, so the ranges have to be on segments the under-cluster nodes are on, and there is no BGP session to configure.
The advertisement is scoped by pool name on purpose — an unscoped one advertises every pool, which would put tenant control-plane VIPs on the management segment and the reverse.
The speaker runs on every node, and spec.placement deliberately does not reach it: any node may have to answer for a VIP.
Both ranges have to stay off the provider segment, and off each other — the reasons are under Address pools.
Pin or auto-assign a control-plane address¶
A tenant cluster that names its control-plane address gets it: the Cluster's control-plane service address becomes a metallb.io/loadBalancerIPs annotation on the Service, and MetalLB honours it as long as the address falls inside a pool.
It then reports which pool it came from:
$ kubectl get svc -A --field-selector spec.type=LoadBalancer \
-o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,IP:.status.loadBalancer.ingress[0].ip,POOL:.metadata.annotations.metallb\.io/ip-allocated-from-pool'
NS NAME IP POOL
dynamo-prod solar 172.21.208.100 tenant-cp-pool
photon-prod quantum 172.21.208.111 tenant-cp-pool
kmetal-console kmetal-console-headlamp 172.21.204.200 mgmt-pool
Both pools are auto-assignable, so a Service that pins neither an address nor a pool can be served out of either one.
Pin the address on anything whose endpoint is written into a kubeconfig, a certificate SAN, or a DNS record — which is every tenant control plane — or name the pool with the metallb.universe.tf/address-pool annotation, which is how the platform's own Services do it.
Two consequences worth knowing before handing addresses out:
- The address is in the kubeconfig and in the api-server certificate. Moving a control plane to another address means re-issuing SANs and re-distributing kubeconfigs, so treat an assignment as permanent.
- A public endpoint is a separate concern. Where tenants reach their control plane on a routable address, that is DNAT at the edge onto this VIP, and the public address has to be in the certificate's SANs. See Control Plane External Access.
Grow an address pool¶
Add the range to that pool's addresses.
Growing is routine, and a second range in the same pool needs nothing else — the advertisement follows the pool, not the range.
Shrinking one below what is already allocated is not — MetalLB will not reclaim an address in use, and the pool ends up inconsistent with what Services hold.
kubectl get ipaddresspools -n kmetal-metallb
kubectl get svc -A --field-selector spec.type=LoadBalancer
Check the second command before shrinking anything.
A Service that will not get an address¶
kubectl describe svc <name> -n <environment> | tail -20
kubectl logs -n kmetal-metallb -l app.kubernetes.io/component=controller --tail=50
| What you see | Usually |
|---|---|
Pending, no event |
The pool does not exist — management left out of the spec, or the pools component has not reconciled. |
AllocationFailed |
The pinned address is outside every pool, or the pool is exhausted. |
| An address assigned, but nothing answers it | The advertisement is missing for that pool, or no speaker runs on a node on that segment. Check the speaker DaemonSet. |
VpcClaim: the Environment's routing domain¶
One per Environment, and the Vpc it creates takes the namespace's name.
Its spec carries intent and nothing else: whether the tenant may egress, and optionally which segment they egress through.
apiVersion: network.kmetal.io/v1alpha1
kind: VpcClaim
metadata:
name: main
namespace: photon-prod
spec:
externalAccess: true
The static route that makes egress work, and the policy routes that keep this VPC off the provider segment and away from its neighbours, are injected by the controller from your configuration. Neither is writable from the claim, which is the point: a tenant cannot ask for a route.
A second VpcClaim in the same namespace is rejected at admission — only one VpcClaim is allowed per namespace.
Give a tenant its own provider network¶
A VpcClaim that names no provider network egresses through defaultProviderSubnet from spec.networking.
That is a fallback, not a segment tenants are placed on together — nothing is shared out by default.
To put a tenant on a segment of its own, its VpcClaim names the provider network:
That tenant's VPC then egresses through provider2 alone, and their EipClaims draw from that segment's subnet instead of the default one.
Three constraints:
providerNetworkRefrequiresexternalAccess: true. A claim with one and not the other is rejected.VpcClaim.specis immutable. An existing tenant cannot be moved to a different provider network by editing their claim — the claim has to be deleted and re-created, which takes the VPC and everything attached to it. Decide at onboarding, not after.- If the provider network declares more than one subnet, the claim must also carry the
network.kmetal.io/provider-network-subnetannotation naming which one to use. Admission rejects the claim otherwise, naming how many subnets it found.
$ kubectl get vpcclaim -n photon-prod
NAME VPC EXTERNAL PROVIDERNETWORK READY AGE
main photon-prod true provider2 True 44h
Delete one¶
Deletion is held by a finalizer while SubnetClaims still reference the VPC, and the claim reports SubnetClaimsPresent.
Delete the subnets first — which means deleting the clusters attached to them first.
SubnetClaim: IP space inside the VPC¶
Each claim reconciles into one cluster-scoped Subnet named <namespace>-<claim>, and, unless you turn it off, a NetworkAttachmentDefinition named <claim>-net in the Environment.
Worker VMs attach through that NAD, so the subnet is what a tenant cluster actually lands in.
An Environment may hold several. That is how one tenant gets separate dev and prod IP planes inside a single isolated VPC.
apiVersion: network.kmetal.io/v1alpha1
kind: SubnetClaim
metadata:
name: production
namespace: photon-prod
spec:
cidrBlock: 10.100.0.0/24
gateway: 10.100.0.1 # defaults to the first usable address
protocol: IPv4 # IPv4 | IPv6 | Dual
excludeIps:
- 10.100.0.1
enableSnatRule: true
| Field | Notes |
|---|---|
cidrBlock |
Required. Must not overlap another subnet in the same VPC. Tenants may overlap each other freely — isolation is the routing domain, not a global allocation. |
gateway |
Defaults to the first usable host in cidrBlock, and must be inside it. |
protocol |
IPv4 by default. |
excludeIps |
Addresses held back from the subnet's IPAM, up to 256. The gateway belongs here. |
attachment.managed |
On by default: the controller creates the NAD and sets the Subnet's provider to <nad>.<namespace>.ovn. Turn it off only when something else owns the attachment. |
attachment.name |
Overrides the generated <claim>-net NAD name. |
enableSnatRule |
SNATs this subnet's egress through the VPC's router-port EIP. Requires externalAccess: true on the Environment's VpcClaim, since that is what creates the EIP — set on a VPC without it, the rule points at an address that does not exist. |
The VPC and the namespaces the subnet serves are derived from the Environment's VpcClaim and are deliberately absent from the spec, so a subnet cannot be attached to a VPC that is not the tenant's.
Size it before applying it¶
SubnetClaim.spec is immutable, so the CIDR is a one-time decision: re-addressing means deleting the claim, which takes the Subnet and every workload interface on it.
Size for the worker VMs the tenant will run, plus what is held back: a /24 leaves around 250 usable once the network and broadcast addresses, the gateway and any excludeIps are out.
Apply the VpcClaim first and let it go Ready.
The overlap check resolves the VPC from that claim's status, and is skipped while the Environment has no VPC — so two overlapping SubnetClaims applied into an empty Environment are both admitted, and the collision only surfaces later, in Kube-OVN.
Once a VPC is there, an overlap is refused at admission with the range it collided with:
$ kubectl apply -f second-subnet.yaml
Error from server: admission webhook "vsubnetclaim.kb.io" denied the request:
cidrBlock "10.100.0.0/25" overlaps SubnetClaim "production" (10.100.0.0/24) in namespace "photon-prod"
Add a subnet to an existing tenant¶
Additive, and safe: a new claim in the same Environment gets its own Subnet and its own NAD, and nothing about the existing subnets changes.
$ kubectl get subnetclaims -n photon-prod
NAME SUBNET VPC CIDR READY AGE
production photon-prod-production photon-prod 10.100.0.0/24 True 45h
staging photon-prod-staging photon-prod 10.101.0.0/24 True 2m
A tenant Cluster joins one by the claim's own name, through the ClusterClass network.subnet variable — staging.
The Kube-OVN Subnet name photon-prod-staging is printed for troubleshooting purposes, and must not be surfaced anywhere else.
Watch capacity¶
The claim mirrors the Subnet's counters, so you can see a tenant running out without reading Kube-OVN objects:
$ kubectl get subnetclaims -A \
-o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,CIDR:.spec.cidrBlock,FREE:.status.v4AvailableIPs,USED:.status.v4UsingIPs'
NS NAME CIDR FREE USED
dynamo-prod production 10.100.0.0/24 250 3
photon-prod production 10.100.0.0/24 243 10
photon-prod staging 10.101.0.0/24 251 2
A subnet close to exhaustion cannot be grown — the spec is immutable.
The answer is a second SubnetClaim and new clusters placed on it, which is why it is worth watching before a tenant hits the wall.
The Ready condition is Kube-OVN's own, copied onto the claim, so its reasons come from there and read a little oddly — ResetLogicalSwitchAclSuccess on a healthy subnet is normal.
WaitingForVpc is kMetal's, and means what it says.
EipClaim: addresses on the provider segment¶
An EipClaim is one address on the provider segment, and it is the unit of grant.
The claims sitting in an Environment are the addresses that tenant has — a cluster with no free claim to bind cannot publish a LoadBalancer Service at all.
Two kinds, told apart by spec.type:
spec.type |
What consumes it |
|---|---|
lrp |
The tenant's VPC, as the egress/SNAT address on its router port. The one the VPC uses is the well-known claim named external; any other lrp claim is inert. |
nat |
A LoadBalancer Service in one of the tenant's clusters, DNATed straight to the workload. One Service at a time. |
Addresses come from the VPC's external subnet — the provider subnet the VpcClaim resolved to, or defaultProviderSubnet when it named none.
Putting a tenant on a provider network of their own therefore also moves where their addresses come from.
Grant one¶
apiVersion: network.kmetal.io/v1alpha1
kind: EipClaim
metadata:
name: web
namespace: photon-prod
spec:
type: nat
v4Ip: 198.51.100.24 # optional
| Field | Notes |
|---|---|
type |
nat (the default) or lrp. |
v4Ip |
Pins a specific address. Left out, Kube-OVN allocates one from the external subnet, minus its excludeIPs. Pin the addresses the edge router or DNS already knows about; let the rest float. |
clusterSelector |
Restricts which of the Environment's Clusters may bind a Service to this claim. Empty or absent matches every cluster in the Environment. |
EipClaim.spec is immutable, like the other claims.
Re-addressing means deleting the claim and re-creating it, which takes the OVN NAT rules with it.
The external claim¶
A VpcClaim with externalAccess: true is gated on the claim named external, of type lrp, being Ready.
The controller auto-creates it, owned by the VpcClaim, so neither you nor the tenant has to write it.
Pre-create it for one reason only: to pin the tenant's egress address with v4Ip before the VpcClaim gets there.
Give it the wrong type and the VPC never becomes Ready — the condition reads EipClaim "external" must have spec.type: lrp.
How a tenant Service picks one up¶
Their side of it, worth knowing because your grant is what makes it work:
- The Service carries
spec.loadBalancerClass: kmetal. kMetal reconciles nothing else, which leaves any otherLoadBalancerimplementation in the tenant cluster alone. - With the annotation
network.kmetal.io/eip-claim: web, it binds to that claim by name. - Without it, it binds to the first free
Readynatclaim in the Environment, in name order. - Bindings are exclusive. A Service naming a claim another Service holds waits, and a warning event lands on the claim; a Service that finds nothing free waits too, with
no free EipClaim available in namespace "photon-prod"on the Service.
An lrp claim is never picked up implicitly, so the external address cannot be taken by a workload.
See who holds what¶
status.binding names the consumer, which makes the Environment's claims an allocation table:
$ kubectl get eipclaims -n photon-prod
NAME EIP V4 BOUND READY AGE
external photon-prod-external 198.51.100.22 true True 47h
web photon-prod-web 198.51.100.24 true True 47h
$ kubectl get eipclaim external -n photon-prod -o jsonpath='{.status.binding.type}{"\n"}'
Vpc
$ kubectl get eipclaim web -n photon-prod \
-o jsonpath='{.status.binding.service.cluster.name}{" "}{.status.binding.service.service.namespace}{"/"}{.status.binding.service.service.name}{"\n"}'
solar demo/whoami
That is the cluster and the Service inside it, reported from the tenant cluster by the binding controller — you do not need the tenant's kubeconfig to see who is on which address.
The EIP column is the cluster-scoped OvnEip behind the claim, named <namespace>-<claim>.
The external claim is the exception: its EIP is <vpc>-<external subnet>, matching Kube-OVN's own convention so it adopts the object as the VPC's gateway EIP.
Revoke one¶
Deleting a bound claim is rejected — the Service has to be deleted or unannotated first, and a claim attached to a VpcClaim cannot be deleted while that VPC exists.
The exception is a binding to a cluster that is already gone, which is not enforced.
If you also let tenants create their own EipClaims, bound them with the count/eipclaims.network.kmetal.io quota rather than by hand — see Multi-tenancy.
Either way, your remaining part is the edge router: the provider segment has to be routed to the under cluster for any of these addresses to answer.
See Load Balancer for the two integration patterns.
Diagnose a claim that will not become Ready¶
kubectl get vpcclaims,subnetclaims,eipclaims -A
kubectl describe subnetclaim <name> -n <environment>
| Condition reason | On | Means |
|---|---|---|
WaitingForVpc |
SubnetClaim |
The Environment has no Ready VpcClaim yet. Fix the VPC first. |
WaitingForEip |
VpcClaim |
External access is waiting on the external EipClaim — check it is type lrp and Ready. |
ExternalGatewayNotReady |
VpcClaim |
The VPC exists, but its external gateway is not up. Look at the provider network and its subnet. |
SubnetClaimsPresent |
VpcClaim |
Deletion is blocked because subnets still reference the VPC. Delete those, and the clusters on them, first. |
EipNotReady |
EipClaim |
Kube-OVN has not assigned the address. Usually a v4Ip outside the external subnet or in its excludeIPs, or an exhausted subnet. |
EipInUseByVpc |
EipClaim |
Deletion is blocked: the claim is the VPC's egress address. The VpcClaim goes first. |
Three failures happen at admission instead, so they never reach a condition — the kubectl apply is what carries the message:
- a second
VpcClaimin an Environment, - a
providerNetworkRefthat names no existing provider network, or an ambiguous one, - a
SubnetClaimCIDR overlapping another subnet in the same VPC.