skip to content

In a shadow window the candidate router disagrees with the incumbent on 18% of tickets - what must each paired record hold to make those diffs diagnosable later?

level: seniorimportance: should knowfreq 51%

answer

  1. aggregate rate names no cause
  2. both decisions on one row
  3. scores, not only the chosen queue
  4. model version and feature version separately
  5. inputs vanish from the online store

basics

~20 s

Enough to reproduce and attribute the diff: both decisions and both raw scores, both model versions and both feature-definition versions, the feature values used or a snapshot reference, per-call latency, and the segment fields you will group by.

solid answer

~50 s

An aggregate disagreement rate tells you the size of a problem and nothing about its cause, so the record has to carry attribution. Store both routers' **decisions and raw scores**, so a near-tie at the decision boundary is distinguishable from a confident opposite call. Store both the **model version and the feature-definition version** for each side, or a diff cannot be attributed to the model rather than to the features. Store the **feature values used**, or a reference to a point-in-time snapshot - by the time anyone investigates, the online store has moved on and the input cannot be reconstructed. Add per-call latency, null and missing-feature markers, and the segment fields (channel, language, ticket type, hour) you will slice by. Key it on the ticket id so the incumbent's real outcome can be joined in later.

code

json · 23 lines
json
{
  "ticket_id": "tkt-8831204",
  "scored_at": "2026-09-18T09:14:02.311Z",
  "mode": "shadow",
  "incumbent": {
    "model_version": "router-4.2.1",
    "feature_defs_version": "fdefs-2026-08-11",
    "queue": "billing",
    "top_scores": { "billing": 0.71, "payments-fraud": 0.22 },
    "latency_ms": 38
  },
  "candidate": {
    "model_version": "router-5.0.0",
    "feature_defs_version": "fdefs-2026-09-02",
    "queue": "payments-fraud",
    "top_scores": { "payments-fraud": 0.64, "billing": 0.61 },
    "latency_ms": 91
  },
  "agreed": false,
  "feature_snapshot_ref": "snap/2026-09-18/09/8f3c1ab2",
  "missing_features": ["customer_tier"],
  "segments": { "channel": "email", "language": "de", "ticket_type": "chargeback", "local_hour": 9 }
}

go deeper

for a junior

The idea to hold is one row per ticket carrying both routers' answers side by side, rather than two separate logs that somebody has to line up afterwards.

for a middle

Explain why raw scores matter as much as decisions: near-ties at the decision boundary and confident opposite calls produce the same disagreement rate and mean entirely different things.

for a senior

Show the attribution procedure - re-score stored features across the model and feature-definition versions - and note that it is impossible unless both versions and the point-in-time inputs were recorded at scoring time.

for a principal

The call to own is retention and scope: how much input data the comparison store keeps, for how long, and under what access rules, balanced against investigations that always arrive after the window has closed.

## What the record is for 'The routers disagree on 18% of tickets' is a headline, not a finding. Roughly 430 disagreements an hour on a desk of this size is far more than anyone can read, so the record has to support three jobs: **grouping** disagreements into a handful of classes, **reproducing** any individual one months later, and **attributing** it to the model, the features or a timing artefact. Everything below exists for one of those three. ## The fields, and what each buys - **Ticket id and scoring timestamp** - the join key to the incumbent's real outcome, and the ordering that lets you correlate a spike of disagreements with a deploy or an upstream incident. - **Both decisions** - the queue each router chose. Without both sides on one row, comparing means an expensive join across two logs whose clocks disagree. - **Both raw scores** - the single most valuable field after the decisions. An 18% disagreement rate made of near-ties just either side of the decision boundary is a calibration difference and mostly harmless. The same 18% made of confident opposite calls is a different model behaving differently, and it needs a human to adjudicate a sample. - **Model version on each side** - the artifact identity, content-addressed or semantically versioned, so a rerun scores the same thing. - **Feature-definition version on each side** - if the candidate ships a new feature pipeline alongside the new model, a diff attributable to 'the candidate' is not attributable to anything actionable. Record them separately and you can answer which one moved. - **Feature values used, or a snapshot reference** - the field teams most often omit and most often need. The online store holds the current value, not the value at scoring time, so a week later the input is gone. - **Null and missing-feature markers** - which inputs were absent. A cluster of disagreements on tickets missing one field is a wiring bug, not a model opinion. - **Per-call latency** - the candidate's and the incumbent's, on the same row, so latency parity can be sliced by the same segments as the decisions. - **Sampling rate of the stratum this ticket came from** - if the mirror is sampled unevenly, every aggregate computed from these rows is wrong without the weight, and nobody notices because the rows themselves look fine. - **Segment fields** - channel, language, ticket type, customer tier, local hour. These are what turn 430 rows an hour into 'almost all of it is German-language chargeback email at the start of the working day'. ## Decision, score, and the thing in between A router usually produces a score per candidate queue and then applies a rule - the top-scoring queue, or the top queue subject to a minimum confidence. Two routers can disagree on the decision while their score vectors are nearly identical, and they can agree on the decision while ranking everything below the top differently. Storing the top few scored queues rather than only the winner costs little and answers both cases; storing only the winner throws the evidence away at write time, which is the one point where it cannot be recovered. ## Reproducing a diff later Attribution is a subtraction exercise, and it works only if each term is recorded: 1. Re-score the stored features with the **candidate model** and the **incumbent feature-definition version**. If the disagreement disappears, the features moved, not the model. 2. Re-score with the **incumbent model** and the **candidate feature-definition version**. If it appears here, the feature pipeline owns it. 3. If neither reproduces it, look at timing: if the two routers each fetched features independently, they may have read either side of an update, and the disagreement is an artefact of the mirror's wiring rather than a property of either router. ## What not to put in it The record does not need the full ticket body to do its job, and copying customer text into a long-lived analysis store creates a retention obligation nobody planned for. Store a reference to the ticket plus the derived features, and fetch the text under the desk's existing access rules when a human adjudicates a sample. Keep the record small enough that you can afford to retain it for the whole window and a while after, because the questions arrive after the window closes.

  • The scores show most disagreements are near-ties, 0.64 against 0.61. Does that change what you do next?
    Yes. Near-ties mean the two routers largely agree on the ranking and differ only where the decision rule tips, so the change is far less risky than the headline 18% suggests, and the interesting question moves to calibration and to the decision rule rather than to the model's judgment. Adjudicate a sample of the confident disagreements, which are fewer and carry the real behavioural difference, and report the two populations separately.
  • Why record the feature-definition version separately from the model version?
    Because a candidate is usually both, and a rollback has to know which one to undo. If the new model ships with a new feature pipeline and you record only 'candidate', every diff is attributable to an inseparable bundle. With both versions on the row you can re-score the stored features in each combination and say whether the model or the features produced the change.
  • Is it enough to store a reference to the online feature store instead of the values?
    Only if the store is versioned or time-travellable. A plain reference resolves to the current value, and the current value is not the one the router scored - the online store is overwritten continuously. Either store the values on the row, or store a reference to an immutable point-in-time snapshot. Anything else means the investigation a week later reproduces a different input and reaches a confident wrong conclusion.

saying these in an interview costs you the question

  • Storing only the chosen queue and discarding the scores
  • Treating an aggregate disagreement rate as a diagnosis
  • Bundling model and feature-pipeline changes under one candidate version
  • Assuming the online store can replay the input values later
  • Logging the two routers to separate tables joined by timestamp