Skip to content

Networking Configuration

Everything the platform needs to know about your network lives in spec.networking on the KMetal object. The operator turns it into the Kube-OVN, MetalLB and kMetal Networking objects that implement it — none of which you create by hand.

For the model behind it, see Networking Concepts. For tenant-facing service exposure, see Load Balancer. For every field and its constraints, see Platform Configuration Reference.

Before you configure anything

Two things must be true of the under cluster first, and neither is kMetal's to do:

A primary CNI is installed and working. Kube-OVN runs as a non-primary CNI and never takes over pod eth0. Its overlay subnet must not overlap your primary CNI's pod network.

The nodes carrying the overlay are labelled.

kubectl label node worker-01 kmetal.io/ovn=

Kube-OVN's per-node DaemonSets run only where that label is present. Every node carrying it must have the tunnel interface configured — a node labelled without one gets DaemonSets that crash-loop, since the chart tolerates every taint and would otherwise place them anywhere.

The overlay

spec:
  networking:
    kubeOVN:
      podCIDR: 10.16.0.0/16
      podGateway: 10.16.0.1          # optional; defaults to the first usable address
      tunnelInterface: bond0.205
      tunnelType: geneve             # geneve | vxlan | stt
      centralNodes:
        - worker-01
        - worker-02
        - worker-03
    serviceCIDR: 10.96.0.0/16
Field Notes
podCIDR The ovn-default subnet — not the cluster pod network, which the primary CNI owns. The two must not overlap.
tunnelInterface The node interface carrying the overlay, typically a kernel VLAN sub-interface. No default is possible; kMetal cannot guess your NIC.
centralNodes The nodes running ovn-central and kube-ovn-controller, forming the OVSDB Raft cluster.
serviceCIDR Must match the --service-cluster-ip-range the control plane was bootstrapped with. Kube-OVN needs it to keep service traffic off the overlay.

Use an odd number of centralNodes so the Raft cluster keeps quorum on node loss — three tolerates one failure. Every node listed must carry tunnelInterface, which usually rules out the control plane.

The operator resolves each name to that Node's InternalIP and passes the list explicitly rather than letting the chart discover nodes by label: discovery races node labelling during a cold bootstrap and yields an empty cluster.

Jumbo frames (MTU 9100) on the transport carrying the overlay are worth having. Geneve encapsulation costs header space, and a 1500-byte transport leaves tenant workloads with an awkward effective MTU.

The invariants are not fields

NON_PRIMARY_CNI, and the CNI config priority that stops Kube-OVN becoming Multus' default delegate, are pinned by the release and absent from the API. They are not site configuration — getting them wrong takes the cluster's pod network down.

Provider networks

A provider network is an underlay segment tenant VPCs egress through. At least one is required; without it tenants have no external path.

spec:
  networking:
    providerNetworks:
      - name: provider
        interface: bond0.206
        excludeNodes:
          - controller-01
        subnets:
          - name: external
            vlanID: 0
            cidrBlock: 198.51.100.0/27
            gateway: 198.51.100.1
            excludeIPs:
              - 198.51.100.0                  # network address
              - 198.51.100.1                  # gateway
              - 198.51.100.2                  # under-cluster management
              - 198.51.100.10..198.51.100.20  # reserved for DNAT to tenant control planes
              - 198.51.100.31                 # broadcast
    defaultProviderSubnet: external

The operator creates a ProviderNetwork per entry, and the kMetal Networking controllers reconcile the Kube-OVN objects behind it.

Field Notes
interface OVN takes this wholesale into its provider bridge, so it must carry no node IP. Give it a dedicated NIC or VLAN sub-interface, never the one the node is managed through.
autoCreateVLANSubinterface Has the kMetal networking daemon create interface on each participating node, instead of expecting the host to have it already.
excludeNodes Every node without the physical NIC behind interface — typically the control plane. A node left in without the NIC gets an OVN bridge with nothing behind it.
defaultProviderSubnet A Kube-OVN Subnet name declared above, not a provider network name. It is the subnet a tenant VPC egresses through when it names none.

Leave vlanID at 0 for a kernel sub-interface

If interface is a kernel VLAN sub-interface such as bond0.206, the host already tags. Setting vlanID non-zero tags a second time, which double-tags the frame at the OVN localnet port and the traffic is silently dropped.

Set it only when handing OVN an untagged interface and asking it to do the tagging.

Be exhaustive with excludeIPs

These addresses are held back from OVN's IPAM. The gateway belongs there, as does anything already assigned on the segment. An address in use elsewhere that is not excluded will be handed to a tenant, taking out whatever was using it.

Entries are single addresses or inclusive first..last ranges, up to 32 per subnet.

Kube-OVN additionally requires a subnet literally named external to exist before it will reconcile external access on tenant VPCs.

Per-tenant egress isolation

Declaring more than one provider network is how a tenant gets a segment of its own. The tenant's VpcClaim names it, and that tenant's egress leaves on that segment alone:

spec:
  networking:
    providerNetworks:
      - name: provider                    # carries the defaultProviderSubnet
        interface: bond0.206
        subnets: [ ... ]
      - name: provider2                   # one tenant's own
        interface: bond0.210
        autoCreateVLANSubinterface: true
        excludeNodes:
          - controller-01
        subnets:
          - name: external-provider2
            vlanID: 0
            cidrBlock: 172.21.210.0/24
            gateway: 172.21.210.1
            excludeIPs:
              - 172.21.210.1

Address pools

spec:
  networking:
    loadBalancer:
      tenantControlPlanes:
        addresses:
          - 172.21.208.100-172.21.208.250
      management:
        addresses:
          - 172.21.204.200-172.21.204.250

The operator creates the MetalLB pools from these. Their names are fixed — tenant-cp-pool and mgmt-pool — because a Service opts into a pool by annotation, and renaming one would strand every Service already pointing at it:

metadata:
  annotations:
    metallb.universe.tf/address-pool: tenant-cp-pool

Each entry is a CIDR (192.0.2.0/24) or an inclusive range (192.0.2.10-192.0.2.20). management is optional; without it no management pool is created, and any Service requesting one stays Pending — including the console.

Keep the pools off the provider segment

The two pools are separate because they sit on different segments with different exposure, and collapsing them lets tenant control-plane VIPs and operator VIPs consume each other's space.

Putting the tenant control-plane range on the provider segment — the one carrying OVN's external gateway — would give tenants a path to SNAT egress and to each other's traffic.

Tenant workload LoadBalancer services do not come from MetalLB at all. They are fulfilled by Kube-OVN from the provider network, through an EipClaim — see Load Balancer.

Blocking additional egress

The provider subnets are blocked for tenant egress automatically — that list is exactly their CIDRs, so writing them out again only created a way to forget one and silently open a path between tenants.

Use extraBlockedEgressCIDRs for ranges that are not provider networks but that tenants still must not reach:

spec:
  networking:
    extraBlockedEgressCIDRs:
      - 10.0.0.0/8          # corporate network reachable from the provider segment

Per-tenant networking

You do not create tenant VPCs, subnets or SNAT rules. Tenants declare them as claims in their own Environment, and the kMetal Networking controllers reconcile the cluster-scoped Kube-OVN objects — injecting the isolation policy routes and the external static route, which are never the tenant's to write.

See Create First Environment for the manifests, and Networking Concepts for why the indirection exists.

Verifying

# The networking components reconciled
kubectl get km kmetal -o jsonpath='{range .status.components[*]}{.name}{"\t"}{.phase}{"\n"}{end}' \
  | grep -E 'kube-ovn|multus|metallb|kmetal-networking|provider-networks'

# Provider networks, as the operator created them
kubectl get providernetworks

# Address pools, with the fixed names
kubectl get ipaddresspool -A

# Nodes eligible for Kube-OVN's per-node workloads
kubectl get nodes -l kmetal.io/ovn

# Smoke-test an address from the management pool
kubectl create service loadbalancer test-lb --tcp=80:80 \
  --dry-run=client -o yaml | kubectl annotate -f- --local -o yaml \
  metallb.universe.tf/address-pool=mgmt-pool | kubectl apply -f-
kubectl get service test-lb -w
kubectl delete service test-lb

DNS

The under cluster uses whatever DNS your Kubernetes installer deployed; kMetal does not manage it. Tenant clusters get their own CoreDNS with their control plane.

External DNS records for tenant services are outside kMetal's scope — run ExternalDNS or your provider's automation against your zone.

Troubleshooting

# Kube-OVN
kubectl get pods -n kube-system -l app=ovs
kubectl get pods -n kube-system -l app=kube-ovn-controller
kubectl get pods -n kube-system -l app=kube-ovn-cni
kubectl get subnets

# Multus
kubectl get pods -n kube-system -l app=multus

# MetalLB
kubectl get pods -n kmetal-metallb
kubectl logs -n kmetal-metallb -l app.kubernetes.io/component=controller --tail=100
kubectl logs -n kmetal-metallb -l app.kubernetes.io/component=speaker    --tail=100

# The claims layer
kubectl get vpcclaims,subnetclaims,eipclaims -A

See Troubleshooting: Networking for diagnostic flows.