skip to content

In a wait-time estimator whose features are built by a nightly job and again in the checkout request path, which skew sources appear?

level: middleimportance: must knowfreq 84%

answer

  1. two code paths, one definition
  2. units, rounding, window bounds, filters
  3. the kitchen with no history
  4. a correct value read too late
  5. code disagreement versus clock disagreement

basics

~20 s

Three families: one definition implemented twice and diverging in units, rounding or window; a missing-entity default chosen differently on each path; and a serving read that returns a row older than the request it is answering.

solid answer

~40 s

Name three families. **Implementation divergence** - both paths compute `kitchen_recent_prep` from scratch, one flooring seconds to whole minutes and one rounding, with window bounds, filters and day boundaries free to drift apart as well. **Default divergence** - a kitchen with no finished orders gets the training set's median of 12 minutes offline and a hard-coded `0` on the request path, so the model meets a value its column never held. **Staleness at serving** - the open-order count is written by a stream, and when the writer lags the request scores a backlog from forty minutes ago, a pairing training never contained. The first two are code disagreements and the third is a clock disagreement, which matters because sharing one transform artifact fixes the first two and leaves the third untouched.

code

pseudocode · 13 lines
pseudocode
nightly_job(kitchen):
    rows = offline_history.finished_orders(kitchen, last = 20)
    if rows is empty:
        return 12                    // training-set median, in minutes
    secs = mean(r.finished_at - r.accepted_at for r in rows)
    return floor(secs / 60)          // whole minutes, always rounds down

request_path(kitchen):
    rows = online_store.recent_orders(kitchen, last = 20)
    if rows is empty:
        return 0                     // hard-coded fallback
    secs = mean(r.duration_seconds for r in rows)
    return round(secs / 60)          // nearest minute, rounds either way

go deeper

for a junior

Learn the three names - two implementations of one definition, a disagreeing default for a missing entity, and a stale read at serving - and one concrete example of each in a live request path.

for a middle

Explain the mechanics of each family: which incidental choices (units, rounding, window, filters, clock) make two implementations drift, why an empty aggregate becomes zero unless someone decides otherwise, and why a lagging writer is not a code bug.

for a senior

Demonstrate the ranking: which family fires on every request, which on a cohort, which under load, and which fix touches which. Say out loud that a shared artifact leaves the stale read entirely in place.

for a principal

The judgment is how much of this becomes a platform guarantee. Owning the definition centrally removes two families for every team but makes the platform a bottleneck on every new feature - that tradeoff is the real conversation.

## One definition, two implementations A wait-time estimate at checkout needs `kitchen_recent_prep`, the average preparation time of the kitchen's last twenty finished orders. The nightly training job computes it by scanning finished orders in the columnar offline store. The request path computes it from rows in the low-latency key-value tier, under a few milliseconds of budget. Two teams, two languages of expression, two sets of edge cases - and one feature name that makes them look like one thing. The question a design round is really asking is: enumerate the ways those two computations can disagree, and say which ones a shared implementation would fix. ## The three families | family | what disagrees | when it fires | fixed by one shared artifact? | |---|---|---|---| | implementation divergence | the computation itself | every request | yes | | default divergence | the value used when the entity has no rows | only for entities with no history | yes, if the default lives inside the artifact | | staleness at serving | the age of the input, not the function | under load, or when a writer lags | no | ## Family one: the same computation written twice This is the widest family, because every incidental choice is a place to differ: - **Units.** One path returns whole minutes, the other seconds. The model receives values roughly 60x larger than any it trained on for the rows that branch produces. - **Rounding.** The nightly job floors seconds to whole minutes; the request path rounds to the nearest. That is a systematic offset of about half a minute in one direction, present on every request. - **Window bounds.** Last twenty orders against last ten, or last twenty against the last hour. Both are defensible; only one matches the trained column. - **Filters.** Cancelled and refunded orders excluded offline and included online, or vice versa. - **Clock and day boundary.** An hour-of-week feature computed in the store's local time on one path and in a fixed reference time zone on the other. This is not a fourth family - it is the same family, with the clock as the thing that differs. - **Tie-breaking and ordering** when two orders carry the same timestamp, which decides which rows fall inside a last-twenty window. ## Family two: the entity with no rows A kitchen that joined this morning has no finished orders. The nightly job's definition fills the gap with the training set's median, 12 minutes. The request path's fallback fills it with `0`, because zero is what an empty average tends to produce when nobody decided otherwise. Both paths run without error; the model meets a value its training column never contained, for exactly the entities with the least operational slack. ## Family three: a correct value read too late The open-order count is materialised into the online tier by a stream. When the writer lags, the request path reads a value that was correct forty minutes ago. No code is wrong - the function is the same function, applied to an older input. Training rows never paired a small open-order count with a long actual wait on a Friday evening, so the model has no learnt behaviour for the combination it is now being asked about. How old a row is *allowed* to be, and who promises that bound, is a separate contract; the skew here is simply that the served input is not the input the definition describes. ## Which fix touches which family 1. **Compile one transform artifact and execute it on both paths.** Removes family one outright and family two if the default is expressed inside the artifact rather than in either caller. 2. **Log the served input vector and replay it against a recomputation.** Detects all three, and the shape of the diff tells you which one you have - a constant offset for family one, an all-or-nothing gap on a cohort for family two, a load-correlated gap for family three. 3. **Reduce the lag or accept it explicitly.** Only family three responds to this, and no amount of shared code substitutes for it. ## In a design round The answer that scores is the one that separates code disagreement from clock disagreement, because they have different fixes and different owners. Candidates who name only the duplicated-implementation family propose one artifact, declare the problem solved, and are surprised when the live error stays high.

  • Which source survives compiling one transform artifact and running it on both paths?
    Staleness at serving. The artifact makes both paths compute the same function, but the request path applies it to whatever row the online tier currently holds, which can be older than the request. A field the request path cannot fetch inside its latency budget survives for the same reason - identical code cannot conjure an input that is not there.
  • How would you rank the three for a checkout wait-time service?
    By how often each fires times how large the error is. Rounding divergence fires on every request with a small bias. The default fires on a small cohort with a large error, and on newly onboarded kitchens. Staleness fires under load - precisely when the estimate matters most - so it usually carries the highest product cost even though its rate is lowest.
  • Is a time-zone difference between the two paths a fourth family?
    No. It is implementation divergence, the same shape as units and rounding: the two computations of one definition disagree on a parameter. Keeping the list at three families matters because the grouping maps onto the fixes - shared code, shared default, and lag control.

saying these in an interview costs you the question

  • Claims a shared feature name guarantees the two paths agree
  • Treats skew as purely a code-duplication problem and ignores stale reads
  • Says a stale serving read is harmless because the value was once correct
  • Assumes both paths inherited the same default because nobody chose one
  • Thinks reviewing the two implementations side by side keeps them equal