skip to content

Two 10 Gb NICs on a Linux server are bonded together with LACP (802.3ad), yet a single large TCP transfer never goes faster than 10 Gb/s. Why, and what would actually make use of both links?

level: seniorimportance: nice to knowfreq 32%

answer

  1. capacity across, not within
  2. the same fields every packet
  3. in-order delivery has a price
  4. everything behind one router looks identical
  5. the switch decides the other direction

basics

~20 s

A bond chooses an outgoing member by hashing packet header fields, and a single TCP connection has constant header values, so every packet takes the same link. Aggregation adds capacity across many flows, never within one.

solid answer

~50 s

An 802.3ad bond does not stripe packets across its members. For each outgoing packet it computes a hash over selected header fields and uses that to pick one member, which guarantees that all packets of a conversation take the same link and therefore arrive in order. A single TCP connection has a fixed source and destination for every field the hash considers, so it hashes to the same member every time and is capped at that member's speed. Incoming traffic is distributed by the switch, using its own independent hash, so you cannot control it from the host. To use both links you need multiple simultaneous flows — a multi-stream copy, several client connections, many peers. The hash policy matters too: the default keys on MAC addresses, which collapses to one link when everything is behind the same router, and switching to a policy that also keys on IP addresses and ports spreads such traffic properly.

go deeper

for a junior

Know that bonding several links gives more total bandwidth and redundancy, not a faster single connection, and that a bond appears to the system as one interface.

for a middle

Explain member selection by hashing over header fields and why that preserves packet ordering, and contrast active-backup with LACP in terms of what the switch must be configured to do.

for a senior

Diagnose the real-world case: the default layer-2 hash collapsing all off-subnet traffic onto one member, the asymmetry between egress hashing and the switch's own inbound choice, and the reordering trade-off of layer-3+4 hashing.

for a principal

Own the capacity decision — when aggregation genuinely buys headroom for a many-flow workload versus when the requirement is single-flow throughput and the answer is a faster interface, plus the redundancy topology across switches that bonding is often really bought for.

## What link aggregation actually promises Link aggregation combines several physical links into one logical interface. The promise is *aggregate* bandwidth and redundancy — not a faster single connection. Interviewers ask this because the gap between those two things is where expensive misunderstandings live: someone buys a second NIC to make a backup job faster and it changes nothing. ## Why hashing, not striping The obvious implementation would be to send packet 1 on link A, packet 2 on link B and so on. Linux can do exactly that — it is the round-robin bonding mode — and it is almost never what you want. The two links have independent queues and slightly different latencies, so packets arrive out of order. TCP interprets out-of-order arrival as loss, triggering duplicate acknowledgements, spurious fast retransmits and a collapsing congestion window. The result is frequently *slower* than one link, and it wreaks havoc on anything sensitive to reordering. So every serious mode instead selects a member by hashing. The bond computes a hash over chosen header fields and maps the result to one of the active members. Because a conversation's header fields do not change, every packet of that conversation traverses the same member, and in-order delivery is preserved. The cost is exactly the constraint in the question: one conversation, one link, one link's worth of bandwidth. ## The hash policy is the interesting knob The transmit hash policy decides which fields feed the hash, and getting it wrong is a very common production disappointment: - **Layer 2** — hashes source and destination MAC addresses. This is the conservative default and it is strictly 802.3ad compliant. It also fails badly in the most common server topology: if the server talks to clients on other subnets, every packet has the *router's* MAC as its destination, so every flow to the entire outside world hashes identically and lands on one member. The bond looks broken; it is behaving exactly as configured. - **Layer 2 + 3** — adds the IP addresses, so traffic to distinct remote hosts spreads even through a single router. - **Layer 3 + 4** — adds the transport ports, so even multiple connections between the same pair of hosts spread across members. This is usually what people want for server workloads. The caveat, documented in the kernel's bonding documentation, is that it is not fully 802.3ad compliant, because fragmented packets lack port information and can therefore land on a different member than the rest of their flow, admitting reordering in that corner case. Note that this policy governs **egress only**. The switch on the other side runs its own hash, with its own policy that you may not control, to decide which member inbound traffic uses. Asymmetric utilisation between transmit and receive is normal and is not a fault. ## Mode choice Linux bonding offers several modes; the ones that matter in an interview are: - **802.3ad (LACP)** — negotiates the aggregation with the switch using the Link Aggregation Control Protocol. It requires matching configuration on the switch side: the ports must be configured as a port-channel/aggregation group. If the switch is not configured, the bond either falls back to a single link or, worse, creates a loop and duplicate frames. The advantage over static aggregation is that LACP verifies both ends agree before forwarding, and detects a link that is physically up but logically broken. - **active-backup** — one member carries all traffic, another takes over on failure. No switch configuration required, works across two independent switches, gives redundancy but zero extra bandwidth. For many servers this is the correct choice, and saying so is a mark of judgment rather than ignorance. - **round-robin** — true per-packet striping with the reordering consequences described above. Occasionally justified for a controlled point-to-point path with a reordering-tolerant workload; rarely elsewhere. Bond state — the mode in effect, hash policy, per-member link status and LACP negotiation details — is reported by the kernel under `/proc/net/bonding/`, one file per bond, which is the first thing to read when a bond is behaving oddly. Member link supervision is configured with a monitoring interval, so that a failed member is removed from the aggregate promptly rather than continuing to black-hole its share of the hash space. ## Making both links count Given all of that, the honest answers to "how do I use both links for this transfer" are: 1. **Use more flows.** Multi-stream transfer tools, parallel connections, or simply many clients. This is what the design is for and it works. 2. **Fix the hash policy** so those flows actually spread — the layer-2 default is frequently the real culprit rather than any inherent limit. 3. **Buy a faster link.** If one conversation genuinely needs more than a single member's speed, aggregation is the wrong tool and a higher-speed NIC is the right one. What does *not* work is expecting a single TCP stream to exceed one member. Any configuration that appears to achieve it is either striping and paying for it in reordering, or measuring something other than a single stream.

  • Why does round-robin striping across bond members usually make TCP throughput worse rather than better?
    Because the members have independent queues, packets arrive out of order. TCP treats out-of-order arrival as a loss signal: it emits duplicate acknowledgements, triggers spurious fast retransmits and shrinks the congestion window. The sender ends up transmitting less than a single link could carry, plus wasted retransmissions. Only a workload genuinely tolerant of reordering, over a tightly controlled point-to-point path, justifies the mode.
  • What is the risk of configuring an LACP bond on the host while the switch ports are left as ordinary access ports?
    The aggregation never forms properly. In the best case LACP refuses to bring members into the aggregate and you run on one link. In the worse case, with static aggregation rather than negotiated LACP, the switch sees the same MAC address appearing on two ports and you get MAC flapping, duplicate frames and unpredictable loss. That mutual-verification property is the main reason to prefer negotiated LACP over static aggregation.
  • When is a bond the wrong answer to a throughput requirement?
    When the requirement is single-flow throughput. Aggregation adds capacity across conversations and cannot exceed one member for any single one, so a backup job, a replication stream or a large single copy sees no benefit at all. The correct remedies are to parallelise the workload into multiple streams, or to buy an interface with the speed the flow actually needs. Bonding remains valuable in that scenario for redundancy, but it should be sold on that, not on speed.

saying these in an interview costs you the question

  • Expects one TCP connection to reach the sum of both links
  • Thinks LACP stripes packets across members by default
  • Configures a bond without matching switch port-channel configuration
  • Ignores that inbound distribution is chosen by the switch
  • Believes round-robin striping is a free way to double throughput

context