A single event loop saturates one CPU core. Compare running one independent loop per core with shared-nothing state against a single loop feeding a shared worker pool, and say when you would choose each.
answer
- loop per core = shard connections + state, no locks
- no work stealing: hot shard pins one core
- handoff = queue + wakeup + cold cache, paid twice
- one loop still saturates one core at high event rates
- shared-nothing converts concurrency into partitioning
basics
~20 sLoop-per-core shards connections and state across cores with no shared mutable data, so it scales near-linearly and avoids locks and cache-line contention — but it cannot rebalance a hot shard. One loop plus a worker pool balances load automatically but pays handoff latency, cache misses, and synchronisation on shared state.
solid answer
~60 s**Loop per core (shared-nothing).** N loops, each pinned to a core, each owning a disjoint set of connections and a disjoint shard of state. No locks, no cache-line ping-pong, no cross-core coordination on the hot path; throughput scales close to linearly with cores. Costs: load imbalance is unfixable at runtime — one hot connection or hot shard pins one core while others idle, since there is no work stealing. Cross-shard operations become explicit message passing with latency and complexity. Per-core caches and pools duplicate memory. **One loop plus worker pool.** The loop does I/O and dispatch; heavy work goes to workers. Load balances itself, and long tasks cannot stall the loop. Costs: every handoff is a queue operation plus a cache-cold restart on another core; shared state now needs synchronisation; and the single loop becomes the bottleneck at high event rates. **Choosing.** Uniform, partitionable, latency-critical, very high event rate favours loop-per-core. Heterogeneous or unpredictable work with genuinely shared state favours the pool. A common hybrid is per-core loops plus a small shared pool for rare heavy tasks.
go deeper
Recall that one loop uses one core, so scaling means either several loops or a worker pool, and that several loops means splitting the data between them.
Contrast the two: shared-nothing avoids locks but cannot rebalance; a pool balances automatically but pays handoff cost and needs synchronised shared state.
Discuss connection distribution, cross-shard messaging, skew, and the metrics that decide — per-core utilisation variance, cross-shard message rate, and handoff share of latency.
Frame the choice as a state-architecture decision: shared-nothing converts a concurrency problem into a partitioning problem, which is only a win with a good key, bounded skew, an answer for cross-shard work, and a shard-migration story.
## The starting constraint One loop is one thread and therefore one core. To use a machine with many cores you must add either more loops or more workers. The choice determines your entire state model. ## Design A: loop per core, shared nothing Run one loop per core, each pinned to its core. Each loop owns: - a disjoint subset of connections, - a disjoint shard of application state (typically hash-partitioned by key), - its own buffers, caches, and allocation pools. Nothing mutable is shared, so there are no locks, no atomic contention, and — importantly — no cache lines bouncing between cores. On modern hardware that last point often dominates: a contended line ping-ponging between cores can cost far more than the logical work protecting it. Getting connections onto the right core happens either at accept time (each loop accepts independently, with the kernel distributing incoming connections across listeners) or via a dedicated acceptor that hands the connection to a chosen loop. Requests that need a shard owned by another core are forwarded as a message, ideally over a lock-free single-producer single-consumer queue per core pair. **Strengths.** Near-linear scaling; predictable per-core latency; no synchronisation on the hot path; failure and slowdown are contained to one core's shard. **Weaknesses.** - *Imbalance.* No work stealing means a hot shard or a single very active connection saturates one core while others idle. Real traffic is rarely uniform (celebrity keys, one enormous tenant), so you need enough shards per core to average out, and a way to split or migrate a hot shard. - *Cross-shard operations.* Anything spanning shards becomes a distributed problem inside one process: multi-key transactions, global aggregates, and joins now require coordination protocols. - *Memory duplication.* Per-core caches and pools multiply footprint by core count. - *Long tasks still stall a loop* — sharding does not fix per-event service time, it only limits the blast radius to one core's connections. ## Design B: one loop, shared worker pool The loop handles I/O and dispatch, offloading heavy or blocking work to a pool. The pool can steal work, so imbalance self-corrects, and a long task never stalls the loop. **Weaknesses.** - *Handoff cost.* Each dispatch is a queue enqueue plus a wakeup and a cache-cold start elsewhere — often several microseconds, which is enormous relative to a small request and is paid twice (out and back). - *Shared state needs synchronisation*, reintroducing the locking and contention problems the single-threaded model removed. - *The loop is still a single core.* At very high event rates the loop itself saturates before the pool does, so this design scales work but not I/O. ## Choosing Prefer **loop per core** when: the workload is naturally partitionable by key or connection; per-request work is small and uniform; event rate is high enough that dispatch cost matters; and tail latency is a first-class requirement. This is the shape used by high-throughput proxies, packet processors, and thread-per-core databases. Prefer **loop plus pool** when: work is heterogeneous and unpredictable in cost; state is genuinely global and hard to partition; or you must call libraries that block and cannot be made non-blocking. **Hybrid** — per-core loops for the common path plus a small shared pool for rare expensive operations — is often the right production answer, provided the rare path really is rare, because every use of it pays the handoff and reintroduces sharing. ## What to measure before deciding - **Per-core loop utilisation variance.** High variance means sharding is not balancing; either re-key or add work stealing. - **Cross-shard message rate.** If a large fraction of requests need another core's shard, shared-nothing is paying coordination cost for no benefit; the partition key is wrong. - **Handoff share of latency.** If offloading costs more than the work offloaded, keep it on the loop. - **Cache-miss and coherence traffic.** Rising coherence traffic with core count is the signature of accidental sharing that shared-nothing was supposed to remove. ## The framing that lands Shared-nothing does not remove the concurrency problem; it *converts it into a partitioning problem*. That is usually a good trade — partitioning is easier to reason about and to observe than lock contention — but only if the data has a natural key, the key distribution is not badly skewed, and you have an answer for cross-shard work and for a hot shard.
- With one loop per core and no work stealing, how do you handle a single hot partition?Increase the number of shards well beyond the number of cores so ownership can be reassigned at a finer granularity, then migrate shards between cores when utilisation diverges. If the skew is a single key rather than a range, replicate that key read-only across cores or promote it to a per-core cache, so reads stay local and only writes coordinate.
- Why can offloading small tasks to a worker pool make latency worse rather than better?Each handoff costs a queue operation, a thread wakeup, and a cache-cold restart on another core, and you pay it again to return the result — often several microseconds. If the task itself is shorter than that, offloading strictly adds latency and coherence traffic while also giving up the lock-free property of loop-owned state.
Loop-per-core is a bank with separate tellers, each owning specific account ranges: no coordination, but a rush on one range leaves other tellers idle. A shared pool is one queue feeding all tellers: perfectly balanced, but everyone pays the walk to whichever window opens.
saying these in an interview costs you the question
- Assuming a worker pool automatically fixes latency, ignoring handoff and cache costs
- Believing shared-nothing removes the need to think about hot keys and skew
- Forgetting that a single dispatch loop still caps I/O throughput at one core
- Sharding by an attribute with heavy skew and expecting even core utilisation
- Treating cross-shard operations as free rather than as in-process message passing with latency and failure modes