skip to content

An actor receives messages faster than it can process them and its mailbox keeps growing. What strategies exist for handling mailbox overflow, and how do you choose between them?

level: seniorimportance: should knowfreq 34%

answer

  1. Little's law: rate mismatch → unbounded queue
  2. unbounded mailbox = process-wide OOM
  3. drop newest / drop oldest / priority / reject
  4. credit-based = real backpressure, no blocking
  5. never block the sender; shard the hot actor

basics

~20 s

Options: unbounded mailbox (risks memory exhaustion), bounded with drop-newest, drop-oldest or priority-based dropping, bounded with rejection back to the sender, or credit-based flow control where senders only send what the consumer has requested. Choose by whether the messages are droppable, and fix the throughput mismatch as well.

solid answer

~60 s

An unbounded mailbox does not remove the problem — it converts a rate mismatch into unbounded memory growth and rising latency, ending in an out-of-memory crash that takes the whole process down. So bound it, and decide what happens when it is full: - **Drop newest** — cheapest; fine for telemetry or refreshable state. - **Drop oldest** — right when only the latest value matters (sensor readings, price ticks, UI state). - **Priority / selective drop** — keep commands, shed notifications; needs message classification. - **Reject and notify the sender** — the sender decides: retry with backoff, fail the request, degrade. - **Credit-based flow control** — the consumer grants the producer N messages at a time; the producer never sends more than granted. This is the only option that actually applies backpressure without blocking. Avoid making the sender *block* on a full mailbox: senders are themselves actors on a shared pool, so blocking propagates stalls and can deadlock a cycle. All of this is triage — also fix the mismatch: shard the actor by key, batch work, move blocking I/O out of the handler, or add a dead-letter path for what you drop.

code

text · 6 lines
text
Consumer -> Producer: Credit(100)
Producer sends at most 100 messages, then waits
Consumer processes 60, then -> Producer: Credit(60)

result: producer rate is bounded by consumer demand;
        no sender ever blocks a thread.

go deeper

for a junior

Know that mailboxes should be bounded and name the basic options: drop, reject, or slow the producer down.

for a middle

Explain why unbounded means an out-of-memory crash for the whole process, and match each strategy to message semantics — latest-value versus every-item.

for a senior

Add credit-based flow control as the real backpressure mechanism, explain why blocking the sender is dangerous on a shared pool, and address the throughput fix: sharding, non-blocking handlers, batching, conflation.

for a principal

Frame it as a capacity and degradation policy: which message classes are droppable, where admission control belongs, what the system promises under overload, and how depth and drop metrics feed alerting and capacity planning.

## Why the queue grows A mailbox is a queue between an arrival process and a service process. Little's law makes the consequence concrete: the average number of messages in the system equals arrival rate multiplied by average time in the system. If arrival rate exceeds service rate for a sustained period, queue length has no equilibrium — it grows without bound, and so does latency, since every new message waits behind the entire backlog. No amount of memory fixes a rate mismatch; it only postpones the failure. ## Why an unbounded mailbox is a decision, not a default An unbounded mailbox is often the default, and it silently chooses: 'I would rather crash the whole process later than lose a message now.' The failure is also badly shaped — memory grows, garbage collection pressure rises, *unrelated* actors slow down, and finally the process dies, losing the entire backlog anyway. The bound is what turns an unbounded, global, unpredictable failure into a local, immediate, designed one. ## The strategies **Drop newest (reject on arrival).** Constant cost, keeps the oldest work. Correct when the value of a message decays or is replaceable: metrics samples, log events, cache hints. **Drop oldest.** Keeps the freshest data. Correct when messages represent *current value* rather than *events to process*: a sensor reading, a price tick, a UI state update. Wrong for commands, since you would silently discard requested work. **Priority / selective dropping.** Classify messages and shed the low-value classes first — keep control commands and heartbeats, drop bulk notifications. Costs a priority structure and the discipline to classify every message type honestly; a system where everything is high priority has no strategy. **Reject and inform the sender.** The mailbox refuses the message and the sender learns about it, so the *sender* decides: retry with backoff, fail the user request fast, or degrade to a cheaper path. This preserves end-to-end correctness because nothing is silently lost, and it converts overload into a visible, measurable signal. It is the right default for command-shaped messages. **Credit-based flow control (application-level backpressure).** The consumer grants the producer a credit of N messages; the producer sends at most N and waits for more credit, which the consumer issues as it drains. The rate mismatch is now impossible by construction — the producer is demand-driven — without anyone blocking a thread. The cost is a protocol on both sides and only works when the producer is under your control (not, say, inbound network traffic). **Blocking the sender — the trap.** Making send block until space frees looks like backpressure, but in an actor runtime the sender is another actor on a shared thread pool. A blocked sender stalls its own mailbox, holds a pooled thread, and propagates the stall upstream; with a cycle in the message graph it deadlocks outright. Rejection or credit is what you want, not blocking. ## Choosing Ask three questions: 1. **Is the message droppable?** Telemetry: yes, drop. A customer's order: no — reject visibly or apply backpressure, never silently discard. 2. **Is the value in the latest item or in every item?** Latest-value semantics point to drop-oldest or to conflation (collapse queued updates for the same key into one). Every-item semantics forbid dropping at all. 3. **Do you control the producer?** If yes, credit-based flow control is the strongest answer. If not, you must drop or reject at the boundary, and ideally do so as early as possible — admission control at the edge is cheaper than deep in the system. ## Do not stop at triage Overflow handling limits damage; it does not create throughput. Address the mismatch itself: - **Shard the actor.** Per-actor processing is serial, so one actor is capped at one core. Route by key to N actors and the ceiling multiplies. - **Get blocking work out of the handler.** A handler doing synchronous I/O has a service rate set by network latency, not CPU. - **Batch.** Process or persist in groups so per-message overhead falls. - **Conflate.** For latest-value data, replace queued duplicates for the same key. ## Operate it Make mailbox depth a first-class metric with an alert, count and dead-letter every dropped or rejected message rather than discarding it silently, and load-test the overflow path deliberately — the failure mode you never exercise is the one that behaves unexpectedly in production. ## Interview delivery Start with 'unbounded is a choice to fail globally later', list the bounded strategies with the semantics that select each, explain why blocking the sender is dangerous in an actor runtime specifically, then pivot to fixing the mismatch (sharding, non-blocking handlers, batching) and to observability of depth and drops.

  • Why not simply block the sender when the mailbox is full — isn't that classic backpressure?
    In an actor runtime the sender is itself an actor running on a shared thread pool. Blocking it stalls its own mailbox, removes a worker thread from the whole system, and propagates the stall to everyone upstream; if the message graph has a cycle it deadlocks. The same intent is achieved safely by rejecting the message or by credit-based flow control, where the producer simply does not send more than it has been granted.
  • The mailbox keeps growing even though the handler is fast. What would you check?
    Whether the handler is really fast in production rather than in a benchmark — a synchronous call to a slow dependency sets the service rate to that dependency's latency. Also check whether one actor is a hot key that should be sharded, since processing is serial per actor and one actor cannot exceed one core. Finally, check whether the producer's rate has legitimately grown, in which case the capacity plan, not the mailbox, is the problem.
  • Where should dropped messages go?
    To a dead-letter destination with enough context to diagnose them: message type, sender, and the reason for the drop. Silent discards make overload invisible and destroy the audit trail for work that was requested but never done. Counting drops per type also tells you whether the classification behind a priority-drop strategy is actually correct.

A mailbox is a restaurant's order queue. Unlimited tickets means the kitchen falls hours behind and the whole restaurant eventually collapses; a bounded rail means you either stop taking orders, discard the stalest, or tell the host how many the kitchen can accept.

saying these in an interview costs you the question

  • Treating an unbounded mailbox as a safe default
  • Blocking the sending actor when the mailbox is full
  • Dropping command-shaped messages silently with no dead-letter or sender notification
  • Believing more memory or a bigger bound solves a sustained rate mismatch
  • Ignoring that per-actor processing is serial, so a hot actor needs sharding rather than a bigger queue

context