skip to content

Your fix for wait-time skew is one transform artifact run by both paths - which skew does that remove, and which survives?

level: seniorimportance: should knowfreq 46%

answer

  1. same code, different input
  2. one artifact, not one specification
  3. staleness is a clock, not a function
  4. availability inside the request budget
  5. the audit covers what the artifact cannot

basics

~20 s

It removes implementation divergence - units, rounding, window bounds, filters - and the missing-entity default if that lives inside the artifact. It does not remove a stale input, an input the request path cannot fetch in budget, or a training job that quietly runs its own query.

solid answer

~50 s

A single definition means one artifact, built once, executed by both the nightly job and the request path - not one design document and two implementations kept in step by review. That removes every incidental disagreement inside the computation, and removes the missing-entity default too if the fallback is expressed inside the artifact rather than by each caller. Three things survive. The **stale input**: identical code over an older row still yields an older answer, because staleness is a clock problem, not a function problem. The **availability gap**: a field the request path cannot fetch inside its latency budget cannot be conjured by shared code. And the **bypass**: if the training job runs its own equivalent query for speed, there is again no single definition. Those are what the served-input log and the replay audit exist to catch.

go deeper

for a junior

The idea to hold on to is that computing a feature in one shared place is better than computing it twice, and that sharing the code still says nothing about how fresh the values fed into it are.

for a middle

Explain which disagreements a shared artifact eliminates - units, rounding, windows, filters, and the default when it lives inside - and why a stale input is a different category of problem altogether.

for a senior

Demonstrate that you know the fix is partial: name the survivors, name the two ways teams fake it, and pair the artifact with per-request input logging and a replay audit before calling the work done.

for a principal

The tradeoff is centralisation. One executed definition removes a failure class for every team but puts the platform on the critical path of every new feature and every backfill, and the pressure to bypass it arrives with the first slow historical rebuild.

## What one definition has to mean The phrase does a lot of work, and teams cash it at different rates. The weak reading is that both paths implement one written specification and reviewers keep them aligned. The strong reading is that one artifact is built from one source, versioned, and *executed* by both the nightly job and the request path, so that making them disagree requires a deliberate deploy rather than a typo in an edge case nobody tested. Only the strong reading buys anything durable. Two implementations agree on the cases somebody thought to write a fixture for, and drift on the ones they did not - the empty window, the cancelled order, the tie on equal timestamps, the day boundary. Those are exactly the cases that produce skew. ## What the artifact removes - **Units.** One computation, so one unit. - **Rounding.** One flooring or rounding step, at one place in the sequence. - **Window bounds and filters.** Last twenty means last twenty on both paths, with the same exclusions. - **Clock handling.** The day boundary and the time zone are properties of the definition rather than of whichever runtime happened to execute it. - **The missing-entity default**, but only if it lives inside the artifact. A default written by each caller is outside the artifact's reach, and this is the single most common place the fix is left half-done. ## What survives it | survivor | why shared code cannot touch it | |---|---| | a stale input row | the function is right and the input is old; staleness is a property of the clock, not the computation | | a field serving cannot fetch in budget | shared code cannot supply an input the request path has no time to read | | a definition change deployed to one path first | during the rollout window the two paths run different versions of the same artifact | | a training job that runs its own equivalent query | the artifact is only a single definition for the callers that actually call it | The first survivor is the one candidates most often miss. If the open-order count in the online tier was last written forty minutes ago, both paths compute the same function over different inputs, and the served vector still describes a kitchen that no longer exists. ## Two ways teams fake the fix 1. **Same source, two builds.** The definition is compiled separately into the batch environment and the serving environment. This is better than two hand-written implementations and still leaves a seam: differing numeric behaviour between runtimes, a version pinned in one place and floating in the other, a rollout that lands on one path a day before the other. 2. **The artifact for serving, a hand-written query for training.** Someone rewrites the definition as a bulk query because scanning a year of history through the shared artifact is slow. The rewrite is equivalent on the day it is written and stops being equivalent the first time the artifact changes. Both are worth naming out loud in a design round, because both are common and both look like the fix from a distance. ## The availability gap deserves a decision, not a workaround When the shared artifact needs an input the request path cannot obtain inside its budget, there is no clever way out. Either the value is precomputed and read as a stored row - which converts the problem into a freshness question with an explicit age - or the feature leaves the model's input set entirely, and the training job drops it too. What is not acceptable is letting the serving path substitute something cheaper, because that is exactly the skew the artifact was adopted to remove, reintroduced by a different door. ## What closes the remainder The artifact is a prevention; it is not a detector, and it makes no statement about the inputs it is fed. The pairing that actually holds is: - the shared artifact, so the computation cannot drift; - the served input vector logged per request, so the served values are recoverable; - a replay audit with a per-feature tolerance and a skew rate, so the survivors show up as a number on a dashboard rather than as a customer complaint. A candidate who proposes only the first and calls training-serving skew solved has answered half the question, and the half they left out is the half that fires under load.

  • What do you do when the shared artifact needs a field the request path cannot fetch inside its budget?
    Make it an explicit decision. Either precompute the value and read it as a stored row, accepting a stated age, or drop the feature from the model's input set and retrain without it. Letting the serving path substitute a cheaper approximation reintroduces exactly the skew the artifact was adopted to remove.
  • How do you stop the training job quietly diverging again after the fix?
    Make the offline rows a product of the same artifact rather than of a hand-written bulk query, version the definition and stamp that version on every served record, and keep the replay audit running so a divergence becomes a rising skew rate. Review alone fails the first time someone needs a faster backfill.
  • Does compiling the same source separately for batch and for serving count as one definition?
    Partly. It removes hand-written divergence, which is the biggest source, but it leaves a seam: two builds can pin different versions, differ in numeric behaviour between runtimes, and roll out on different days. Treat it as a good approximation that still needs the replay audit rather than as a guarantee.

saying these in an interview costs you the question

  • Declares skew solved as soon as both paths share a transform artifact
  • Treats two code copies kept in step by review as one definition
  • Assumes the training job uses the artifact while it runs its own query
  • Believes shared code can supply a field serving cannot fetch in time
  • Calls a stale online row an implementation bug in the transform
  • Leaves the missing-entity fallback in each caller after adopting the artifact