In an XDP-based L4 load balancer, what is the difference between returning XDP_TX and returning XDP_REDIRECT for a forwarded packet, and why do such designs typically encapsulate the packet and use direct server return rather than plain NAT?
answer
- one exit is the same NIC only
- the other needs a redirect map
- the reply should skip the balancer
- wrap it, do not rewrite it
- hash plus a flow table for stickiness
basics
~20 sXDP_TX resends the packet out the interface it arrived on; XDP_REDIRECT hands it to a different target through a redirect map. Encapsulation with direct server return lets backends reply straight to the client, so only the request half crosses the load balancer.
solid answer
~50 s`XDP_TX` puts the frame back on the wire through the same NIC, which means the load balancer must rewrite the Ethernet header itself and cannot choose an egress device. `XDP_REDIRECT` is the general form: with `bpf_redirect_map()` it can send the packet to another netdev, to another CPU, or into an AF_XDP socket, and a redirect to another device needs that device's driver to support XDP transmit. As for the forwarding scheme: an XDP balancer typically hashes the flow, looks up a backend, wraps the original packet in an outer IP header — IPIP or a similar encapsulation — and sends it on. The backend decapsulates and replies **directly to the client**, so return traffic never touches the balancer. That halves the traffic it carries, avoids keeping return-path state, and fits `XDP_TX`, which can only put packets back out the interface they came in on.
go deeper
Know that an XDP program can forward a packet as well as drop it, and that the load balancer decides which backend a packet goes to before the kernel network stack sees it.
Be able to distinguish the two forwarding verdicts precisely — same interface versus a redirect-map target — and explain that XDP_TX means writing the Ethernet header yourself.
Explain why encapsulation with direct server return is the standard shape: the balancer carries requests only, keeps no return-path state, preserves the client's addresses, and matches XDP_TX's same-interface constraint. Mention flow stickiness and MTU without being asked.
Own the tradeoff against an L7 proxy tier: what you gain in cost per packet you lose in request-level control, and the design pushes real requirements onto every backend, onto the MTU of the fabric, and onto how draining and health checking are done.
## The two ways out An XDP program that has decided to forward a packet has two exits. `XDP_TX` is the simplest: the driver transmits the buffer back out the interface it arrived on. There is no routing lookup and no neighbour resolution, because those live far above this hook, so the program is responsible for the Ethernet header. If you are relaying to a next hop, you write the source and destination MAC addresses yourself. The constraint that matters architecturally is the interface: `XDP_TX` cannot choose a different NIC. `XDP_REDIRECT` is the general mechanism. The program calls `bpf_redirect_map()` (or `bpf_redirect()`) to name a target and returns `XDP_REDIRECT`; the driver completes the redirect at the end of its poll cycle, batching the transmits. Targets fall into three families: another network device, another CPU for heavier processing, and an AF_XDP socket for delivery to a userspace program with a zero-copy path. Redirecting to another device requires the *egress* driver to implement the XDP transmit entry point, so this can fail on a NIC that supports XDP for receive only. ## Why encapsulate The naive design is NAT: rewrite the destination address to a backend, rewrite the source on the way back, keep a translation table. It works, and it has two problems at scale. First, the return path. If the balancer rewrote the source address, the backend replies to the balancer, which must translate again and forward to the client. Now every byte of a response — and responses are usually much larger than requests — crosses the balancer, and the balancer must hold per-connection state that both directions consult. Its capacity is bounded by the total traffic of the service, not by the request traffic. Second, `XDP_TX` can only transmit back out the ingress interface, which suits a design where packets arrive and leave on the same link, and suits it much better than one that needs asymmetric routing decisions. Direct server return inverts this. The balancer leaves the client's original packet intact and wraps it in an outer IP header addressed to the chosen backend — IPIP, IPv6-in-IPv4, or a scheme like GUE that carries extra metadata. The backend strips the outer header and finds a packet addressed to the service's virtual IP, which it holds on a loopback address, so it answers as though it had received the packet directly. Its reply goes straight to the client, following normal routing, and never returns through the balancer. The balancer therefore carries request traffic only, keeps no return-path state, and — crucially for consistency — has not touched the client's addresses at all, so the connection the client sees is end-to-end with the virtual IP throughout. ```c /* sketch: the forwarding decision, not a complete program */ struct backend *b = bpf_map_lookup_elem(&backends, &hash); if (!b) return XDP_DROP; if (encap_ipip(ctx, b->addr) < 0) return XDP_DROP; return XDP_TX; ``` ## Keeping flows on one backend Encapsulation solves the path; it does not by itself solve stickiness. Every packet of a TCP connection must reach the same backend, and the balancer is stateless per packet. Two mechanisms do the work together: a hash over the flow's five-tuple to pick a backend deterministically, chosen so that adding or removing one backend disturbs as few existing flows as possible, and a small lookup table of recently seen flows so that a change in the backend set does not rehash connections that are already established. That table lives in a BPF map and is consulted before the hash. ## What the design gives up Direct server return is an L4 technique. The balancer never terminates the connection, so it cannot inspect or route on anything above the transport header — no host-header routing, no TLS termination, no request-level retries. It requires cooperation from every backend, which must decapsulate and must be configured with the virtual IP. Encapsulation adds header bytes, so the effective MTU shrinks and path MTU handling becomes something you must actually think about. And health checking, capacity planning and drain behaviour all have to be built around a datapath that keeps no connection state of its own. Those are real costs, and they are why this design shows up in front of large stateless fleets — Meta's open-source Katran and Cilium's XDP balancer both work this way — rather than as a general-purpose replacement for an L7 proxy.
- With direct server return, how does the backend end up answering for an address it does not own on any physical interface?It does own it, just not visibly on the wire: the virtual IP is configured on a loopback interface so the host accepts packets addressed to it and sources replies from it, while never answering ARP for it. After decapsulation the inner packet is addressed to that VIP, so the normal stack processes it and replies with the VIP as source — exactly what the client expects.
- Why does an XDP load balancer keep a flow table when a consistent hash already picks a backend deterministically?Because the hash's input set changes. When a backend is added or removed, a consistent hash disturbs only a fraction of flows — but a fraction of established TCP connections still breaks. A lookup table of recently seen five-tuples pins those connections to the backend they started on, so the hash only decides for flows that are genuinely new.
- What breaks first when you add an encapsulation header to forwarded packets?Path MTU. The outer header eats bytes, so a full-size client packet plus encapsulation exceeds the link MTU somewhere. Either the fabric between balancer and backends runs a larger MTU, or the balancer clamps the advertised MSS on the way through, or you end up depending on fragmentation and ICMP messages that are frequently filtered.
- Can an XDP-based L4 load balancer route requests based on the HTTP Host header?No. It forwards packets without terminating the connection, so it never sees a reassembled byte stream, has no TLS keys, and cannot rely on a header being present in the first segment. Host-based routing requires a proxy that terminates the connection; the XDP layer's job is to get packets to a pool of such proxies cheaply.
saying these in an interview costs you the question
- Thinking XDP_TX can transmit out any interface
- Expecting the stack to rewrite MACs for XDP_TX
- Believing DSR still routes replies via the balancer
- Assuming an L4 XDP balancer can read HTTP headers
- Ignoring the MTU cost of encapsulation