skip to content

Your shadow ticket router reuses the live serving code, so it also writes prediction logs, updates the prediction cache and emits routing events - what must you isolate first?

level: seniorimportance: should knowfreq 47%

answer

  1. reads real, writes contained
  2. walk the path, list the sinks
  3. cache write is a silent release
  4. tag rows before training reads them
  5. no write permission beats a flag

basics

~20 s

Every write the scoring path performs. The shadow run must not populate the shared prediction cache, emit routing events, spend shared quota or reuse idempotency keys, and its prediction rows must be tagged so training never reads them.

solid answer

~40 s

Shadow is a property of the **writes**, not of the reads. Reading live features is the whole point; writing anything the live system later reads quietly ends the experiment. Walk the scoring path and find every sink: the prediction cache a live request may hit, downstream routing events or agent notifications, shared aggregate counters that feed features back, the prediction log used to build training labels, metrics that would double-count volume, and quota or rate-limit budget on shared enrichment services. Then isolate by **capability rather than by discipline** - run the shadow worker under an identity with no write permission on those sinks, so a forgotten branch fails loudly instead of leaking. A request-scoped shadow flag checked by each writer is the weaker form: one writer that forgets it is a silent, untracked release.

go deeper

for a junior

Hold the distinction: the shadow run may read anything, but it must not change anything the live system will later read. Its answer belongs in a log, not in a cache or a queue.

for a middle

Be able to enumerate the sinks on a scoring path - cache, downstream events, shared counters, the prediction log, metrics, quota, idempotency keys - rather than naming one and stopping.

for a senior

Argue for isolation by capability over isolation by discipline, and describe the symptom of the cache leak: no errors, green dashboards, a disagreement rate that falls because the incumbent is reading the candidate's own writes.

for a principal

The standing question is who owns the invariant after you leave. A flag decays with every refactor; a credential boundary survives one. Decide which guarantee the platform should offer every future shadow window by default.

## Shadow is a property of the writes The reads are supposed to be real. The candidate should fetch the same live features, hit the same enrichment services and see the same ticket text, because that realism is what the window buys. What makes the run a shadow is that **nothing it produces changes the state anyone else observes**. Every review of a shadow design is therefore a hunt for sinks on the scoring path. ## The write classes to hunt down 1. **The prediction cache.** If the candidate writes its routing decision into the cache the live path reads, some live tickets are now routed by the candidate. That is not a shadow window any more; it is an untracked release with no record of which tickets it touched. This is the most damaging one and the easiest to miss, because the cache write usually sits inside the shared scoring function. 2. **Downstream events and notifications.** A routing event, a queue enqueue, an agent notification, a customer-facing acknowledgement. Emitting these twice routes tickets twice and shows agents work that does not exist. 3. **Shared aggregate counters.** Counters such as 'tickets scored for this customer in the last hour' are often themselves features. Double-counting them corrupts the candidate's inputs and the incumbent's at the same time, and the resulting disagreement is an artefact of the measurement. 4. **The prediction log that feeds training.** If shadow rows land untagged in the table the next training set is built from, the candidate's own outputs become labels. Tag every row with the mode and the model version, and make the tag part of the table's key so an untagged row cannot appear. 5. **Metrics, dashboards and alerts.** Mirrored calls that increment the same request counters inflate volume, distort the p99 the team watches, and can fire alerts about a load that no customer generated. 6. **Quota and rate-limit budget.** A shared enrichment or translation service with a per-minute allowance will happily spend it on shadow traffic and throttle the live path. Either give the shadow worker its own credential and allowance, or account for it in the budget explicitly. 7. **Idempotency keys.** If the mirrored call reuses the live request's key on a downstream write, the downstream service can treat the **live** call as a duplicate and swallow it. The side effect here is invisible in the shadow system and appears as a missing action in the real one. ## Two ways to isolate, one of which actually holds | Mechanism | How it works | Why it holds or does not | |---|---|---| | Request-scoped shadow flag | Every writer checks `mode == shadow` and skips | Depends on every writer, now and after the next refactor; one missed branch leaks silently and nothing alerts | | Separate identity or credential | The shadow worker has read access and no write permission on those sinks | A forgotten write fails loudly and is caught in the first minutes of the window, not by a customer | | Separate sink entirely | Shadow records go to their own store or their own partition | Removes the shared-table question completely; costs a second path to maintain | The flag is not worthless - you usually need it anyway to route records to the shadow sink - but it should never be the **only** thing standing between the candidate and production state. Prefer the arrangement where the wrong behaviour is impossible over the one where it is merely discouraged. ## The failure that looks like nothing The cache case deserves a second look because of how it presents. Nothing errors. Latency is unchanged. The dashboards are green. The only symptom is that some tickets are routed by a model nobody approved, and the paired records for those tickets say both routers agreed - because the incumbent read the candidate's own write back out of the cache. A window contaminated this way reports a **falling disagreement rate over time**, which reads like the candidate converging on the incumbent and is actually the opposite. ## A checklist before the window opens - List every write the scoring function performs, including the ones inside helper libraries. - Run the shadow worker under a read-only identity on each of those sinks and confirm it starts clean. - Confirm shadow rows in the prediction log are tagged and that the training-set query excludes them. - Confirm shadow calls carry their own idempotency keys and their own client credential. - Confirm request and error metrics separate the two modes, and that no alert routes on shadow counters. - Re-run the list after any refactor of the serving path while the window is open.

  • The disagreement rate fell steadily from 18% to 3% over a week-long shadow window. What would you check first?
    Contamination before convergence. A shared prediction cache or a shared aggregate counter lets the candidate's own output feed back into the comparison, so the two routers appear to agree more as the cache fills. Check whether the incumbent is reading rows the shadow worker wrote, then check whether the candidate's features include counters that mirrored traffic incremented. Only after both are ruled out is it worth looking for a genuine data shift.
  • Is it ever right for shadow traffic to write to a shared store?
    Yes, to one it owns. A shadow-only partition, table or cache namespace is fine and often necessary - you need somewhere to put paired records, and a shadow-only cache is the honest way to measure the candidate's cache hit rate. The rule is about sinks the live path reads. Sharing the storage engine is fine; sharing the keyspace the live path resolves against is not.
  • How do you stop shadow rows from reaching the next training set?
    Make the mode part of the record and part of the boundary: every prediction row carries mode and model version, the training-set query filters on mode explicitly rather than by omission, and the filter is covered by a test that fails if an untagged row exists. Relying on a separate table alone is weaker, because the next refactor may unify them.

saying these in an interview costs you the question

  • Letting shadow predictions populate the cache the live path reads
  • Isolating only through a flag every writer must remember
  • Assuming read-only means safe because nothing was written deliberately
  • Leaving shadow rows untagged in the shared prediction log
  • Reusing the live request's idempotency key on mirrored downstream calls
  • Alerting on shadow error counters as if customers were affected