skip to content

Reactive Systems

The system-level claim behind the paradigm: responsive, resilient, elastic services built on asynchronous message passing. Interviewers use it to see whether you can say when not to go this way.

on this pageshow

explore

questions

14

In a fulfilment pipeline, picking sends packing an asynchronous message instead of calling it — what does picking wait for?

level: juniorimportance: must knowfreq 60%

answer

  1. the send ends at the boundary
  2. no thread, no frame, no result
  3. acceptance is not processing
  4. faults do not unwind across it
  5. answers come back as messages

basics

~10 s

Only for the message to be accepted for delivery, not for packing to run. Picking holds no thread and no result, and hears about the outcome only if a later message tells it.

solid answer

~40 s

The send finishes at the boundary. Picking builds a self-contained message, hands it toward packing's address and returns immediately — it keeps no thread, no call frame and no pending result, and it has no evidence that packing has started, only that the message was accepted. Because no call frame connects the two, a fault inside packing never unwinds into picking; picking's error handling covers sending, not handling. If picking needs an answer it arrives as another message, which means picking needs a reply address, a correlation value such as the order identifier, and somewhere to keep the in-flight order until the answer comes or is given up on. Latency per order does not improve — picking simply stops spending it.

code

pseudocode · 11 lines
pseudocode
// direct call: picking holds the outcome and the failure
function pick(order)
    items = reserve(order)
    label = packing.pack(items)     // picking waits right here
    return label

// message-driven: picking hands the work on and moves along
function pick(order)
    items = reserve(order)
    send(to: packingAddress, PackOrder(order.id, items))
    return                          // no label, nothing waited for

go deeper

for a junior

Recall that the send returns once the message is accepted for delivery, and that the sender gets no result and no error from the recipient's work. Any answer arrives later, as a separate message.

for a middle

Explain the two decouplings — in time and in fate — and what replaces the call frame: a reply address, a correlation value, and explicit state for work still in flight.

for a senior

Show that you plan for the quiet failure this creates. Absent replies, not raised errors, are the signal, so the sender needs its own notion of too long and the boundary needs measurement.

for a principal

Frame the trade honestly: per-order latency usually rises while the system stops propagating one component's slowness everywhere. Decide which boundaries are worth that bookkeeping and which should stay direct calls.

## Two components, two ways to hand over work In a fulfilment pipeline, picking gathers the items for an order and packing puts them in a box. There are two ways picking can hand that work over, and the difference is not stylistic: it decides what picking holds while packing works, and what picking finds out when packing fails. A **direct call** puts both components on one thread of execution. Picking's call frame stays on the stack holding the order, the items and the expectation of a return value. Packing's duration becomes part of picking's duration. Packing's failure unwinds into picking's frame. An **asynchronous message send** ends at the boundary. Picking builds a message — a self-contained description of the work to be done — hands it toward packing's address, and returns. The send completes when the message is accepted for delivery, not when packing has done anything with it. ## What picking waits for Only the handoff. Concretely, once the send returns: - picking holds **no thread** parked on packing; - picking holds **no result** — there is no label, no confirmation, nothing to inspect; - picking has **no evidence** that packing is even running, only that the message was accepted toward its address; - picking will **not** receive packing's failure as an exception, because no call frame connects them. That last point surprises people most. On a message boundary the sender's error handling covers *sending*, not *handling*. A fault inside packing belongs to packing, which must record it, retry it, or escalate it to whatever supervises packing. ## The two kinds of decoupling this buys 1. **Decoupling in time.** Picking's rate is no longer tied to packing's. If packing is momentarily slow, picking keeps accepting orders, and the difference in rates accumulates at the boundary instead of travelling backwards as latency. 2. **Decoupling in fate.** A stalled or crashed packing cannot hold picking's resources hostage, because picking lent it nothing — no thread, no frame, no lock. | | Direct call | Asynchronous message | |---|---|---| | what the sender holds | a frame, a thread, a pending result | nothing once the send returns | | when the sender learns of failure | immediately, as a raised error | only if a later message says so | | sender's latency | includes the recipient's work | excludes the recipient's work | | what the sender must know | the recipient's interface | an address and the messages it accepts | ## Getting an answer back Asynchronous does not mean answerless. It means the answer is another message. Where picking genuinely needs one, three things must exist: an **address** for the reply to be sent to, a **correlation value** such as the order identifier so the reply can be matched to the request that caused it, and **somewhere to keep the in-flight state** until the reply arrives or picking decides it never will. That third item is the real cost. A call stack tracked in-flight work for free: the frame *was* the record that an order was half-done. Once the frame is gone, the sender owns that bookkeeping explicitly. Fire-and-forget — where picking genuinely does not care what happens next — is the cheap case, not the only case. ## What this costs - **Per-order latency does not improve.** The work still takes as long, and usually a little longer, because a handoff and a queue were added. What improves is that the sender is not the one spending that time. - **Failure becomes quiet.** Nothing is raised and nothing stops. A recipient that is not processing looks exactly like one that is merely busy, until somebody measures. - **Messages must be self-contained.** A message carrying a reference into the sender's live state is a direct call in disguise; the recipient may be in no position to follow that reference. - **Ordering is narrower than people assume.** Two messages produced by two different senders are not put in order by the boundary just because one was created first. ## Two ways to fake it Moving the same direct call onto a background worker makes the *caller* non-blocking without making the *boundary* message-driven: the caller still names the component, still depends on its availability to get a result, and still receives its failure. Returning a placeholder that the caller resolves later has the same shape — it changes what the wait looks like, not the direction of the dependency. The boundary becomes message-driven when what crosses it is a self-contained message addressed to a recipient, and the sender's next action does not depend on what that recipient does with it.

  • If picking does not wait, how does it ever learn that packing succeeded?
    Through a later message. Packing sends a result to a reply address, carrying the order identifier so picking can match it to the request. Picking must keep enough in-flight state to do that matching, or hand ownership of the outcome to some other component entirely.
  • Does an asynchronous send mean picking can never be slowed by packing?
    No. Decoupling in time is real but not unlimited. If the boundary applies flow control, a full intake eventually makes the send itself wait or be refused, and picking feels it. What is removed is the per-order coupling, not every possible form of pressure.
  • Is running the existing call on a background worker the same thing?
    No. The caller stops blocking, but it still names the component, still needs it available to get a result, and still receives its failure. The dependency and its direction are unchanged; only the shape of the wait moved.

saying these in an interview costs you the question

  • Asynchronous means each order is processed faster
  • The sender catches the recipient's exception where it sent the message
  • A completed send means the recipient accepted and started the work
  • A sender needing a reply keeps no state to match it against
  • Putting the same direct call on another thread makes the boundary message-driven
open as a page

Why does asynchronous data flow buy nothing measurable for an internal admin tool with a dozen concurrent users?

level: middleimportance: must knowfreq 60%

basics

~20 s

Asynchronous data flow buys capacity, not speed: it stops workers being parked while waiting on input and output. A tool with a dozen users never runs out of workers, so there is no parked capacity to reclaim.

open as a page

In a fulfilment system, what distinguishes a message-driven boundary between packing and shipping from an event-driven one?

level: middleimportance: must knowfreq 72%

basics

~10 s

Addressing. A message is directed at a named recipient and expresses intent toward it; an event is a broadcast fact about what already happened, addressed to nobody, which zero or many observers may consume.

open as a page

Which of the four reactive system properties does a service fail to deliver if it uses asynchronous streams internally but calls every dependency with a blocking request?

level: seniorimportance: must knowfreq 55%

basics

~20 s

All four remain unclaimed, because the four properties are claims about the boundary between components and that boundary is still synchronous. Internal streams buy more concurrent conversations per worker, which is a capacity gain, not one of the properties.

open as a page

A ticket site's on-sale dashboard shows a healthy mean response time while buyers report waits - what measurement would actually prove the system is responsive?

level: seniorimportance: must knowfreq 62%

basics

~20 s

A high percentile of one named request, measured over every outcome including timeouts and refusals, at a stated arrival rate, across the spike window. Responsiveness is a bounded tail under load, not a healthy average.

open as a page

When components address each other by logical name rather than network location, what does that transparency make possible?

level: middleimportance: should knowfreq 48%

basics

~20 s

Moving the recipient without touching its senders. A logical address can resolve to the same process, another machine, a restarted instance or many instances, so components can be relocated, restarted and scaled while the sending code stays identical.

open as a page

The Reactive Manifesto names responsive, resilient, elastic and message-driven - how do those four properties depend on one another?

level: middleimportance: should knowfreq 50%

basics

~20 s

Responsiveness is the goal. Resilience and elasticity are the means that keep it true under failure and under changing load. Asynchronous message passing is the foundation both means stand on. They are four layers, not four equal virtues.

open as a page

Why does one dependency with no non-blocking path change whether asynchronous data flow is worth adopting at all?

level: seniorimportance: should knowfreq 52%

basics

~20 s

The benefit is capacity reclaimed by never parking a worker, so it exists only on paths that never park one. A dependency with no non-blocking path either forfeits that saving or forces a dedicated pool sized like the old design.

open as a page

Which measurements taken before an asynchronous rewrite would confirm or kill a claimed capacity gain?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Measure where a request's time goes, how many requests are in flight at peak, what the machine runs out of first, and the memory each connection holds. Parked workers beside idle processors confirm the claim; saturated processors kill it.

open as a page

If the shipping component crashes while packing keeps sending it messages, what does the message boundary actually contain?

level: seniorimportance: should knowfreq 52%

basics

~20 s

The fault, not the work. Shipping's crash cannot unwind into packing, which holds no thread or frame and keeps running — but the unshipped orders still exist, accumulating quietly at the boundary as a backlog nobody raised an error about.

open as a page

What experiment during a ticket on-sale would prove a resilience claim - one failing dependency, the rest of the system unaffected - is real?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Inject a slow dependency under the on-sale load and compare two request paths: one that never touches it and one that always does. Containment means the untouched path keeps its response-time bound and the affected path still answers inside one.

open as a page

When is asynchronous data flow worth making the default style for every new service across several teams?

level: principalimportance: should knowfreq 38%

basics

~20 s

Only when most of the portfolio has the qualifying workload and the dependencies can be reached without parking a worker. The costs land on people and are largely fixed per team, so they amortise when the style is everywhere and multiply when it is a minority.

open as a page

If a component takes one message at a time from its mailbox, what does that serial handling buy inside it?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

State confinement. Only one handler touches the component's state at a time, so its internals need no locks and its invariants hold between messages — concurrency lives between components rather than inside any one of them.

open as a page

Elasticity is meant to protect responsiveness, yet scaling out mid on-sale worsened the checkout request's 99th-percentile latency - how would you settle that trade-off?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Rank the properties before arguing: responsiveness is the goal and elasticity is an instrument judged only by whether the bound holds. State the bound, measure the scale event as a cost against it, and for a scheduled spike provision ahead rather than react.

open as a page