Linux's SO_REUSEPORT lets several processes hold listening sockets on the same TCP port. How would you decide whether to scale a service that way rather than having one process accept connections and hand them to workers?
answer
- one queue each, not one shared
- the kernel hashes the four-tuple
- hashing is not balancing
- what happens to a closing worker's queue
- set before bind, same uid
basics
~20 sSO_REUSEPORT gives every worker its own accept queue and lets the kernel spread new connections across them by hashing the connection's four-tuple, removing the single-acceptor bottleneck. The tradeoffs are hashed rather than balanced load, and connections lost from a worker's queue when it exits.
solid answer
~60 sWith `SO_REUSEPORT` each worker creates its own listening socket on the same address and port, and the kernel picks one of them per incoming connection by hashing the four-tuple. That gives each worker a private accept queue, removes the contention and thundering-herd effects of many processes accepting on one shared socket, and scales cleanly across cores — it is why nginx offers `reuseport` on its listen directive. The costs are the reason it is a decision rather than a default. Distribution is a hash, not a balancer, so long-lived or unevenly expensive connections skew across workers with no rebalancing. More importantly, when a socket leaves the group, connections already sitting in that worker's accept queue are dropped, so naive rolling restarts lose in-flight connections. A single acceptor that distributes work, or a load balancer in front, gives you control over placement and draining at the cost of a central bottleneck. I would reach for `SO_REUSEPORT` on connection-rate-bound services where per-core scaling dominates, and keep a single acceptor where fair placement or clean draining matters more.
go deeper
Know that SO_REUSEPORT is what allows several processes to listen on one port at once, and that it is a different option from SO_REUSEADDR with a different purpose.
Explain the mechanism: every socket sets the option before bind, the kernel groups them and picks one per connection by hashing the four-tuple, and each worker drains its own accept queue.
Show the operational consequences — hashed distribution skews with long-lived or uneven connections, and closing a worker discards the connections queued on its socket, which breaks naive rolling restarts.
Own the decision. Justify it only where connection-establishment cost is the measured bottleneck and connections are short and uniform, and weigh it against a single acceptor, a shared listening socket or a front-end balancer on the axes of placement control and graceful draining.
## The problem it solves A multi-process server historically had two options. Either one process holds the listening socket and every worker calls `accept()` on the *same* inherited descriptor, or one process accepts and hands the connection to a worker. The first suffers contention: many workers blocked in `accept()` on one socket, one shared accept queue, and historically a thundering herd as they were woken. The second serialises every new connection through a single process, which becomes the ceiling on connection rate. `SO_REUSEPORT`, added in Linux 3.9, changes the shape. Each worker creates its **own** socket, sets the option before `bind()`, and binds the same address and port. The kernel keeps them as a group and, for each incoming connection, selects one member by hashing the connection's four-tuple. Each socket then has its own SYN queue and its own accept queue, and each worker accepts only from its own. Two requirements are worth stating precisely: every socket in the group must set the option *before* binding, and for security the sockets must be created by the same effective UID, so an unprivileged user cannot slide a listener into a group serving a privileged service's port. Note also that a UDP variant exists with the same option, spreading datagrams across sockets by the same hashing idea. ## What you actually get **Per-core scaling.** No shared queue means no cross-CPU contention on the hot path of connection establishment. On connection-rate-bound workloads — short-lived connections, high churn — this is the difference between saturating one core and using all of them. **No single acceptor bottleneck.** Nothing has to touch every connection before a worker sees it. **Simple operation.** Workers are symmetric and independent. nginx exposes it as `listen 80 reuseport;`, and several runtimes expose it directly on their listener setup. ## What it costs **Hashing is not balancing.** Selection is by four-tuple hash, so distribution is statistically even over *connections*, not over *work*. Long-lived connections, connections with wildly different cost, or a small number of heavy clients all skew load, and the kernel never rebalances an established connection. A busy worker stays busy. If your connections are long-lived and expensive, this is the argument against. **Draining is genuinely hard.** This is the tradeoff most candidates miss. When a socket is removed from the group — the worker exits, or closes its listener — connections that already completed the handshake and are waiting in *that socket's* accept queue have no worker to accept them and are lost. A rolling restart that stops workers one by one therefore drops a slice of in-flight connections each time. Doing it safely means stopping new work to that worker, draining its accept queue, and only then closing, which is precisely the coordination `SO_REUSEPORT`'s independence was meant to avoid. A single-acceptor design, or a load balancer that can be told to stop sending to an instance, gives you a drain point. **Less control over placement.** You cannot say "send this connection to the idle worker". Linux does offer `SO_ATTACH_REUSEPORT_CBPF` and `SO_ATTACH_REUSEPORT_EBPF` (from Linux 4.5) to attach a classic or extended BPF program that chooses the socket — commonly used to steer a connection to a socket pinned to the CPU where its packets arrive, improving locality. That is a real lever, but it is an expert one and it does not turn hashing into load balancing. ## The alternatives on the table **Single acceptor with descriptor passing.** One process accepts and passes the connected descriptor to a chosen worker over a Unix domain socket using `SCM_RIGHTS`. You get exact placement, real load awareness and a clean drain, at the cost of a component that must handle every connection. **Shared listening socket.** Workers inherit one descriptor and all call `accept()`. Simple and still perfectly adequate at moderate rates; the shared queue self-balances (whoever is free accepts next), which is an advantage `SO_REUSEPORT` gives up. **A layer above.** A load balancer or proxy in front makes the whole question moot for placement and draining, at the price of another hop and another thing to operate. ## How to decide The honest framing is that `SO_REUSEPORT` optimises the connection-establishment path and pays for it in placement control and lifecycle complexity. Choose it when connection rate is the measured bottleneck and connections are short-lived and roughly uniform, so that hashing approximates balancing. Prefer a shared socket or a single acceptor when connections are long-lived or heterogeneous, when you need graceful draining for deploys, or when you simply have not measured accept-path contention as the limiting factor. And whichever you choose, make the deployment story explicit first: an architecture that cannot drain is a production problem long before accept contention is.
- What exactly goes wrong during a rolling restart of SO_REUSEPORT workers?Each worker owns a private accept queue. When one closes its listening socket, connections already established and waiting in that queue have no one left to accept them and are dropped, so clients see resets. Safe restarts require draining that worker's queue before closing, or fronting the group with something that can be told to stop directing work at it.
- If the kernel hashes connections across the group, why can load still end up badly uneven?Because the hash distributes connections, not work. Long-lived connections stay pinned to the socket they landed on, and a few expensive clients can concentrate on one worker with no rebalancing. Uniform, short-lived connections make hashing a good approximation of balancing; heterogeneous or persistent ones make it a poor one.
- What does attaching a BPF program with SO_ATTACH_REUSEPORT_EBPF let you change?It replaces the default four-tuple hash with your own socket-selection logic for the group. The common use is CPU locality — steering a connection to the socket owned by the worker pinned to the CPU that received the packets, cutting cross-core traffic. It gives control over placement, but it still selects at connect time and cannot rebalance established connections.
saying these in an interview costs you the question
- SO_REUSEPORT round-robins connections evenly
- It is the same thing as SO_REUSEADDR
- Only the first socket needs the option set
- Any user's process can join the group
- Restarting workers one by one is automatically safe