Two HAProxy nodes sit behind a round-robin DNS record and each enforces a per-source-IP request limit using its own stick table. What goes wrong, and what does adding an HAProxy `peers` section change?
answer
- tables are per-process memory
- limit multiplied by node count
- peers push updates asynchronously
- local peer matched by hostname
- best effort, not consensus
basics
~20 sEach node counts only the traffic it sees, so a client spread across both nodes gets roughly double the intended limit, and a client that moves nodes loses its stickiness. A peers section replicates table entries between nodes asynchronously, giving one shared view rather than two independent ones.
solid answer
~60 sStick tables are per-process memory. With two nodes and no replication, a client whose requests land on both is counted twice over, so a 100-requests-per-10-seconds limit effectively becomes 200, and a client that lands on the other node finds no entry, so its sticky server assignment is gone and it is rebalanced. A `peers` section fixes the split: you declare the peer nodes and their sync addresses, add `peers <name>` to the stick-table declaration, and each node pushes entry updates to the others as they change. Two caveats matter. Replication is asynchronous best-effort push, not a consensus store — there is a propagation delay, so a fast burst can briefly overshoot on both nodes, and during a partition each side counts alone. And the local peer is identified by matching the node's hostname to a `peer` entry, so a hostname mismatch leaves a node replicating nothing while looking correctly configured. A useful side effect: on reload the outgoing process hands its tables to the new one over the same mechanism, so counters and pins survive.
code
ini · 8 linespeers lb_cluster
peer lb1 10.0.0.1:10000
peer lb2 10.0.0.2:10000
frontend fe_public
bind :80
stick-table type ip size 100k expire 10m store http_req_rate(10s) peers lb_cluster
http-request track-sc0 srcgo deeper
Know that a stick table belongs to one HAProxy process, so two load balancers keep two separate tables unless they are explicitly told to share.
Explain the arithmetic — a per-node limit is effectively multiplied by the number of nodes — and describe what a peers section adds to the stick-table declaration.
Diagnose the whole failure: doubled limits, lost pins, blocklists known to one node only. Characterise peers replication honestly as asynchronous best-effort, name the hostname-matching trap, and mention that the same mechanism carries tables across reloads.
Decide whether the state should be shared at all: prefer stateless cookie persistence across a proxy fleet, accept approximate replicated counters for abuse shaping, and push genuinely exact quotas into a store designed for consistency rather than into the proxy.
## Why two nodes are two worlds A stick table lives in one HAProxy process's memory. Nothing about the traffic being load-balanced across nodes makes those tables converge. The consequences show up differently depending on what the table is for: - **Rate limiting.** Each node sees roughly half the client's traffic and counts only that half. Your threshold is effectively multiplied by the node count. Worse, it is multiplied unevenly — a client that happens to pin to one node hits the real limit, while one spread across both never does. - **Stickiness.** An entry created on node A does not exist on node B. When a client's next connection resolves to B, no entry matches, the balance algorithm runs, and the client can land on a different server mid-session. If the application keeps session state locally, that reads to the user as being logged out at random. - **Blocklists.** A client marked abusive on one node is unknown on the other, so half the abusive traffic sails through. ## The peers section ``` peers lb_cluster peer lb1 10.0.0.1:10000 peer lb2 10.0.0.2:10000 frontend fe_public bind :80 stick-table type ip size 100k expire 10m store http_req_rate(10s) peers lb_cluster http-request track-sc0 src http-request deny deny_status 429 if { sc_http_req_rate(0) gt 100 } ``` The same block is deployed to both nodes. Each connects to the others and pushes entry updates as they happen; a node that joins pulls the current contents from its peers so it does not start blind. Newer HAProxy versions also let you declare the tables inside the `peers` section itself, which keeps a shared table definition in one place. ## Which peer am I? HAProxy works out which `peer` line refers to itself by matching the node's hostname, which can be overridden on the command line with `-L`. This is the classic operational trap: deploy an identical config to a host whose hostname matches none of the peer names and the node quietly participates in nothing. It starts, it serves traffic, its table simply never syncs. Verifying that each node knows its local peer identity belongs in the deploy check, not in the incident review. The table definitions must also agree across peers — same type, same stored data types — or updates cannot be applied on the receiving side. ## What replication does and does not guarantee The peers protocol is asynchronous replication over a TCP connection between nodes. It is not a quorum, not a lock service, and not a transaction. Practically: - **There is a propagation delay.** A burst arriving simultaneously on both nodes can push both past the threshold before either learns about the other's increments. For rate limiting that is acceptable — you are shaping abuse, not enforcing a billing quota. For anything that must be exact, a proxy-local table is the wrong tool. - **A partition splits the state.** Each side keeps counting on its own and reconverges afterwards; there is no fencing and no elected owner. - **Counters merge by update, not by arithmetic.** Do not reason about the result as though the two nodes' request counts were being summed into an exact global figure. ## The reload benefit The same mechanism carries tables across a configuration reload: the outgoing process transfers its stick tables to the incoming one, so accumulated rates, pins and blocklist entries survive rather than resetting to empty. On a fleet that reloads on every config change, that difference is the gap between a limit that works continuously and one that is reset several times a day. ## The design alternative worth naming Before reaching for replication, ask whether the state needs to be shared at all. Cookie-based persistence carries the routing decision in the client's request, so any node can honour it with no table and no peers — which is why it scales across a proxy fleet more gracefully than `stick on src`. Replication is unavoidable for counters, since a counter is inherently server-side state, but it is often avoidable for persistence. And if the requirement is a genuinely exact, globally consistent quota across nodes, that belongs in a shared store designed for it rather than in the proxy's in-memory table.
- Why can a client still briefly exceed the limit even with peers configured?Replication is asynchronous. A burst hitting both nodes at once is counted locally before either node's updates reach the other, so both can admit traffic up to the threshold before converging. That is a shaping mechanism behaving as designed, not a bug — anything requiring an exact global count needs a store built for it.
- A node was added to the cluster and its table never syncs, though the config is identical. Where do you look first?The local peer identity. HAProxy matches the node's hostname against the `peer` names to decide which entry is itself; a host whose name matches none of them participates in nothing while starting cleanly. Confirm the hostname or set it explicitly with `-L`, and check the peer definitions are reachable on the sync port.
- Would cookie-based persistence have needed peers at all?No. The routing decision travels in the client's cookie, so any node can decode it without shared state. That makes cookie persistence the more scalable choice across a proxy fleet. Replication remains necessary for counters, because a request rate is inherently state the server has to accumulate.
saying these in an interview costs you the question
- Assuming stick tables are shared across nodes by default
- Calling peers a consistent or quorum-based store
- Expecting exact global counts across replicated nodes
- Ignoring that the local peer is matched by hostname
- Thinking a reload always resets the tables regardless