On a bare-metal Kubernetes cluster, how does MetalLB give LoadBalancer Services an external IP, and how do you choose between its Layer 2 and BGP modes?
answer
- allocate, then announce
- controller plus speaker DaemonSet
- one ARP owner per IP
- routers do ECMP
- rehash resets some flows
basics
~20 sMetalLB's controller assigns an address from an IPAddressPool to each LoadBalancer Service. Its speakers then attract traffic for that address. In Layer 2 mode one node answers ARP/NDP for it; in BGP mode nodes advertise it to routers, which spread traffic across nodes with ECMP.
solid answer
~50 sBare metal has no cloud-controller-manager that can build balancers, so MetalLB fills the gap and does two jobs. Its **controller** assigns an address from an `IPAddressPool` and writes it into the Service status. Its **speaker** DaemonSet makes the network deliver that address to a node, and kube-proxy takes over from there. In **Layer 2** mode, one elected node answers ARP (IPv4) or NDP (IPv6) for the address. That is simple and needs only a shared subnet, but all traffic for the address enters through one node, and failover waits for neighbours to learn the new owner. In **BGP** mode, speakers peer with your routers and advertise the address from several nodes, so routers spread flows with ECMP. You get real multi-node spreading, but you need router cooperation, and ECMP rehashing can reset flows when nodes change. Choose L2 for small flat networks and BGP when you control routers and need throughput.
code
yaml · 21 linesapiVersion: metallb.io/v1beta2
kind: BGPPeer
metadata:
name: tor-rack-7
namespace: metallb-system
spec:
myASN: 64512
peerASN: 64513
peerAddress: 10.18.7.1
---
apiVersion: metallb.io/v1beta1
kind: BGPAdvertisement
metadata:
name: billing-bgp
namespace: metallb-system
spec:
ipAddressPools:
- billing-pool
nodeSelectors:
- matchLabels:
node-role.example.com/edge: "true"go deeper
Recall that bare metal has no cloud balancer, so MetalLB assigns an IP from a pool and makes a node answer for it.
Explain the controller's allocation from an IPAddressPool and the speaker's announcement, and contrast one ARP owner with BGP advertisement plus ECMP.
Show operating judgment: the L2 single-node bottleneck and failover through gratuitous ARP, BGP rehash resets, limiting advertising nodes, strictARP with IPVS, and per-tenant pools.
Decide how the platform exposes Services on premises: router peering ownership with the network team, pool allocation across tenants, and when a hardware or dedicated edge tier beats in-cluster announcement.
## Why bare metal needs something extra A `type: LoadBalancer` Service needs an implementation to assign an address and make it reachable. On bare metal no cloud-controller-manager can build a balancer, so the Service stays at `EXTERNAL-IP <pending>`. **MetalLB** is the common in-cluster implementation. It does not add a separate balancer device. It makes the **nodes themselves** reachable at the Service's address, and kube-proxy (or another Service dataplane) forwards from the node to the pods. ## MetalLB's two jobs 1. **Address allocation**: the `controller` Deployment watches LoadBalancer Services, picks an address from an `IPAddressPool` (`spec.addresses`, plus options such as `autoAssign` and `avoidBuggyIPs`), and writes it into `status.loadBalancer.ingress`. A Service can request a specific address with the `metallb.io/loadBalancerIPs` annotation, or a specific pool with `metallb.io/address-pool`. 2. **Announcement**: the `speaker` DaemonSet makes the physical network deliver packets for that address to a node. An `L2Advertisement` or `BGPAdvertisement` object decides which pools are announced, how, and from which nodes (`nodeSelectors`). MetalLB can be limited to Services with a matching `spec.loadBalancerClass` using its `--lb-class` flag, which lets it coexist with another implementation. Because traffic arrives addressed to the balancer IP itself, MetalLB does not need NodePorts, so `allocateLoadBalancerNodePorts: false` works with it. ## Layer 2 versus BGP | Aspect | Layer 2 mode | BGP mode | |---|---|---| | How traffic is attracted | one elected node answers ARP/NDP for the IP | speakers advertise the IP to BGP peers | | Nodes receiving traffic for one IP | one at a time | every advertising node (router ECMP) | | Network requirement | clients or router on the same L2 segment | routers that peer and accept the routes | | Failover | new owner sends gratuitous ARP/NDP; clients and switches must update caches | peers withdraw the route when the session drops; speed depends on hold time or BFD | | Main limitation | single-node bottleneck for that IP | ECMP rehash on node changes can reset flows | | Setup effort | low | needs `BGPPeer` (`myASN`, `peerASN`, `peerAddress`) and router config | ### Layer 2 details - Each speaker takes part in an election for each address, and one node **owns** it. - With `externalTrafficPolicy: Local`, MetalLB only considers nodes that have a local ready endpoint. That keeps the client IP without an extra hop. - It is not true load balancing across nodes: one node's NIC carries all traffic for that address. You spread load by using several addresses. - If kube-proxy runs in IPVS mode (deprecated in Kubernetes 1.37), set `strictARP: true` in its configuration so nodes do not answer ARP for addresses they do not own. ### BGP details - Each eligible speaker opens a BGP session to the configured routers and advertises the service address as a host route. - Routers install several equal-cost next hops and hash each flow to one node. - When a node joins or leaves, the router **rehashes**. Some existing flows land on a node that has no connection-tracking state for them and get reset. Keeping the advertising node set small and stable, for example dedicated edge nodes selected with `nodeSelectors`, reduces this churn. - A shorter `holdTime` or BFD makes failure detection faster. ## Choosing on a real cluster For the PDF-invoice renderer on a 210-node multi-tenant platform cluster in a data centre with top-of-rack routers: - **BGP** is the natural choice. Throughput spreads across nodes, and advertisements can be limited to a labelled set of edge nodes, so the routers do not juggle 210 next hops. - **Layer 2** suits a small lab, a single-rack cluster, or a network team that will not peer with Kubernetes. - Either way, give tenants separate pools (for example a dedicated `/28` for billing) and control which namespaces may request which pool. ```yaml apiVersion: metallb.io/v1beta1 kind: IPAddressPool metadata: name: billing-pool namespace: metallb-system spec: addresses: - 198.51.100.192/28 autoAssign: false --- apiVersion: metallb.io/v1beta1 kind: L2Advertisement metadata: name: billing-l2 namespace: metallb-system spec: ipAddressPools: - billing-pool ```
- In MetalLB Layer 2 mode, what does a node failure look like to clients?Another speaker is elected owner and sends gratuitous ARP (or unsolicited NDP) so neighbours update their caches. Connections that were open on the failed node are lost. New connections recover once clients and switches accept the new mapping, which can take noticeably longer if a device ignores gratuitous ARP. Throughput stays limited to one node's link afterwards.
- Why should you limit which nodes advertise addresses in MetalLB BGP mode?Every advertising node is an ECMP next hop, and every change to that set makes routers rehash flows, which can reset some connections. A small, stable set of edge nodes means less churn, fewer routes for the routers to hold, and a clear network boundary. `nodeSelectors` on the `BGPAdvertisement` or `L2Advertisement` set it.
Layer 2 mode is one receptionist who answers the phone for a number, and a new one takes over only after callers learn to reach them. BGP mode publishes the number in several directories at once, so calls spread across several desks.
saying these in an interview costs you the question
- MetalLB Layer 2 mode spreads one IP across all nodes
- BGP mode works without any router configuration
- MetalLB runs a separate proxy that terminates client connections
- MetalLB needs a cloud-controller-manager to assign addresses
- Node changes never affect existing flows in BGP mode