skip to content

How do you wire the mirrored call to a shadow ticket-routing model so it cannot add latency or failures to the live request?

level: middleimportance: must knowfreq 55%

answer

  1. off the critical path, always
  2. answer first, then mirror
  3. bounded queue, drop on full
  4. separate pool, separate connections
  5. errors become counters, never retries

basics

~20 s

Keep the candidate off the critical path. Answer the live request from the incumbent first, hand a copy to a bounded background worker with its own pool and timeout, and let shadow failures only increment a counter.

solid answer

~40 s

The live reply must not depend on the candidate's latency, availability or correctness, and every wiring choice follows from that. Dispatch the mirror **after** the live response is produced, into a **bounded** queue that drops when full rather than blocking or growing. Run the shadow work on its own worker pool and its own client connections so it cannot consume live capacity. Give the shadow call a timeout even though nobody is waiting - an uncapped call pins a worker and a connection, and enough of them saturate the pool the drop rule was protecting. Catch every shadow error, count it, and never retry or page on it. Finally, shed the mirror first under load, and record which periods were shed so nobody reads the sample as uniform.

code

pseudocode · 21 lines
pseudocode
on ticket_arrived(ticket):
    features   = fetch_features(ticket)          // live path, one fetch
    live_queue = incumbent.score(features)
    respond(live_queue)                          // customer path ends here

    if shadow_enabled and hash(ticket.id) % 100 < shadow_percent:
        accepted = shadow_buffer.offer(copy(ticket, features))   // non-blocking
        if not accepted:
            metric.increment("shadow.dropped")   // shed, never block

shadow_worker:                                   // own pool, own connections
    loop:
        item = shadow_buffer.take()
        try:
            with timeout(150 ms):
                score, queue = candidate.score(item.features)
            write_paired_record(item, queue, score)
        catch timeout:
            metric.increment("shadow.timeout")   // a finding, not an incident
        catch any_error:
            metric.increment("shadow.failed")    // no retry, no rethrow

go deeper

for a junior

The point to hold on to is ordering: the live answer is produced and returned first, and the copy for the candidate is handed off afterwards. Nothing the candidate does can slow down or fail the real request.

for a middle

Be able to name the four properties and why each exists: dispatch after the reply, a bounded buffer that drops, isolated pools and connections, and errors that become counters rather than retries or pages.

for a senior

Show the operating view: shed the mirror first under load, record the shed intervals so the sample is not misread, cap the shadow call to protect the pool, and be able to switch the whole thing off from configuration without a deploy.

for a principal

The judgment call is where to attach the mirror. An in-process fork gives the cleanest comparison and the weakest isolation; a separate stack gives real isolation and real load on the feature store. Decide which risk you are buying down.

## The rule everything follows from The live reply must not depend on the candidate in any way - not on its latency, not on its availability, not on its correctness. If a single shadow failure can turn into a slow or failed ticket submission, the window is not a safety measure, it is a new outage source. ## Where the mirror hangs There are three usual places to attach it, and they test different things. | Attachment | What it tests | What it costs or hides | |---|---|---| | In-process fork, after the live decision | The candidate on exactly the feature values the live path used | Shares CPU and memory with the live process; does not exercise the candidate's own feature fetch | | Edge mirror to a second serving stack | The candidate plus its own feature fetch and its own fleet | Doubles load on the online feature store; the two stacks read a few milliseconds apart, which manufactures spurious disagreement | | Replay of the recorded request stream | Decisions, cheaply, with zero risk to the live path | Latency under replay is not production latency; if the log lacks the feature values used, you re-read today's values against yesterday's ticket, which is a point-in-time error | The third is the safest and the weakest. Teams commonly start there and move to a fork or an edge mirror once they need the latency evidence. ## The four properties the mirror needs 1. **Already answered.** Dispatch only after the live response has been produced. A mirror that runs before or alongside the reply is one slow candidate away from being on the critical path. 2. **Bounded.** The hand-off is a fixed-size queue with a non-blocking offer. When it is full, drop the copy and count the drop. An unbounded queue turns a slow candidate into a memory incident, and a blocking offer puts the candidate back on the critical path through the back door. 3. **Resource-isolated.** Its own worker pool, its own client and connection pool, ideally its own process or fleet. Sharing a thread pool with the live path means a saturated shadow starves live work using nothing but queueing. 4. **Silently failing.** Catch everything, increment a counter, do not retry, do not page. A shadow timeout is a finding about the candidate, not an incident about the desk. ## Timeouts, even though nobody waits It is tempting to let a shadow call run as long as it likes, since no user is blocked on it. Cap it anyway. An uncapped call holds a worker, a connection and the memory of its feature payload; a few hundred of them exhaust the pool that the drop rule exists to protect, and the drop rate then spikes for reasons that have nothing to do with arrival rate. Set the cap in the same order as the live scoring budget and treat timeouts as recorded data - 'the candidate exceeded 150 ms on 4% of tickets' is exactly the kind of result the window is for. ## Shedding order, and its honest cost Under pressure the mirror goes first: disable it when live p99 crosses its budget, or when the drop rate crosses a threshold. The cost is real and should be stated rather than hidden - you lose shadow coverage exactly at peak, which is when parity matters most. Record the shed intervals with the results so the latency comparison is not read as a full-traffic measurement it never was. ## Two details that bite - **Reuse the fetched features or fetch again?** Reusing the values the live path already fetched is cheaper, puts no extra load on the online feature store, and gives the cleanest comparison because both routers scored identical inputs. Fetching again exercises the candidate's own feature path, which is what you will actually run later, but it doubles online-store load and introduces read-time skew between the two stacks. Some platforms materialise features on write and others compute them on read, and the choice is more consequential on the second kind. - **One toggle, no deploy.** The mirror must be switchable off from configuration. If turning it off requires a release, the first time it misbehaves you will be shipping code under pressure. The mirror must also not write anything the live system reads - prediction caches, downstream events, shared counters. That isolation is a separate design task from this one, and it is the part teams most often skip.

  • The shadow buffer is dropping 40% of mirrored tickets during peak hours. Is that acceptable?
    It is acceptable for the live path and harmful for the comparison, and the two have to be judged separately. The drops prove the isolation works. But the surviving sample is now biased toward quiet minutes, so peak-hour latency and peak-hour disagreement rates are unmeasured. Either raise shadow capacity, lower the mirroring percentage so the buffer keeps up evenly, or record the drop intervals and restrict latency claims to the periods that were fully covered.
  • Should the shadow call reuse the feature values the live request already fetched, or fetch its own?
    Reuse them when the question is 'does the candidate decide differently', because identical inputs make every disagreement attributable to the model rather than to a read-time race. Fetch separately when the question is 'can the candidate's own serving stack hold up', since that is the path you will later run for real - but accept double load on the online feature store and some spurious disagreement from the two reads landing milliseconds apart.
  • Why should shadow errors not page the on-call engineer?
    Because no customer is affected by them, and a pager that fires for something with no user impact trains the team to ignore it. Route them to a counter and a dashboard the model's owners watch. The one exception is a shadow failure that indicates the isolation itself has broken - rising live latency correlated with the mirror, for instance - and that alert is on the live signal, not on the shadow error count.

saying these in an interview costs you the question

  • Calling the candidate in parallel and awaiting both results
  • Letting the shadow call run uncapped because nobody is waiting
  • Retrying failed shadow calls to keep the sample complete
  • Sharing one thread pool between live scoring and the mirror
  • Growing the shadow buffer instead of dropping mirrored copies
  • Paging the on-call engineer when shadow error counts rise