skip to content

Why ship the rule-based arrival window first and instrument it, rather than waiting until the model is ready?

level: seniorimportance: must knowfreq 55%

answer

  1. ship the promise, start the clock
  2. log the inputs as of decision time
  3. outcome joins back days later
  4. recomputed values can encode the outcome
  5. residual structure reveals the headroom

basics

~20 s

Shipping the rule first delivers the product immediately and starts the clock on the evidence a model needs: every promise logged with the inputs as they stood at decision time, joined later to the actual delivery. Without that log, a model has no honest training set and no measured bar.

solid answer

~50 s

A rule-first launch does three things a waiting project does not. It puts an arrival window in front of customers now, at the accuracy the heuristic already has. It establishes the bar on the live population rather than on an offline sample. And it manufactures the data a learned estimator will need, which does not exist yet: for each promise, the identity of the parcel, the window quoted, the rung of the ladder that answered, the table version, and **the input values as they stood at that moment** - dispatch hour, day of week, service level, hub backlog - joined weeks later to the delivered timestamp. Logging the inputs as-of the decision is the part teams skip and regret: those values move on, and rebuilding them later from current tables produces numbers the rule never saw, some of which quietly encode the outcome. Months of that log is what turns "we should build a model" into a claim with evidence behind it.

go deeper

for a junior

Recall that a model needs recorded examples of past decisions and their outcomes, and that shipping the rule first is what starts producing them while customers already get a date.

for a middle

Explain what each logged field is for - the promise, the ladder rung, the table version, the as-of inputs - and why the outcome joins back only after the parcel is delivered.

for a senior

Show the leakage judgment: values rebuilt after the decision differ from what the estimator saw and can carry the outcome, so the training set must record what was knowable at promise time.

for a principal

Treat instrumentation as the cheap option that keeps the expensive one open: a few logged fields today decide whether the build case can be argued from evidence months from now, or only asserted.

## Two jobs in one stage Staging a rule first is not a consolation prize for not having a model. It is the stage that makes the model decidable, and it carries two jobs at once. The **product job**: checkout needs a date today, the heuristic can produce one today, and a promise that is right four times in five is worth far more than no promise at all. Shipping it also surfaces everything that has nothing to do with estimation - where the window is displayed, what happens when it is missed, who answers the customer. The **evidence job**: a learned estimator needs a training set, and that training set does not exist in the warehouse. Operational tables record where parcels are now, not what any estimator believed at the moment of the promise. If you do not record it as the decisions happen, that information is simply gone. ## What the decision log has to carry For each promise, at the moment it is made: - **The parcel identifier**, so the outcome can be joined to it later. - **The promised window and the point estimate behind it**, so the rule's own error is measurable per parcel rather than in aggregate. - **The rung of the backoff ladder that answered** and the support behind it, which is both the provenance of the promise and a natural slice for measurement. - **The version of the transit table** in force, so a change in the rule can be separated from a change in the network. - **The input values as they stood at that moment** - dispatch hour, day of week, origin and destination, service level, the sorting hub's backlog, the declared weight class. - **The outcome**, joined when the parcel is delivered, along with when the join completed. The join is what makes the rest useful, and it arrives late by construction: a promise made today gets its outcome in days, so the log is a stream of half-finished rows until then. Any read of the data has to account for rows whose outcome has not matured, rather than treating the delivered ones as the whole picture. ## As-of values, not recomputed ones The most common way this stage fails is subtle. A team logs the parcel and the promise, then assumes the inputs can be rebuilt later from the warehouse by joining on parcel and date. Two things break: 1. **The values moved.** A hub's backlog at the moment of the promise is not in any table a month later; what is there is a later value, which is a different number about a different moment. 2. **The later value can encode the outcome.** A rebuilt "backlog" measured after the parcel moved reflects, in part, how the parcel's own journey went. A model trained on it looks brilliant offline and fails in production, because at prediction time that information does not exist. The rule for the stage is therefore blunt: **write down what was knowable at the moment the promise was made**, and treat anything reconstructed afterwards as suspect until proven otherwise. ## What the log tells you before a model exists After a few months, the log answers questions that are otherwise guesswork: - **How good is the rule, really** - per lane, per service level, per rung of the ladder, on live traffic. - **Where the error lives** - if residuals are structured by dispatch hour, backlog or weekday, that structure is headroom a learned estimator could claim. If the residuals look like unstructured scatter around zero, there is little for any estimator to find in the signals being logged. - **What population you actually have** - only parcels that received a promise are in the log. Lanes where checkout suppressed the window, or orders cancelled before dispatch, are absent, and a model trained on the log inherits that population whether or not anyone notices. - **Whether a new input is worth plumbing** - a signal not logged now is a signal that will delay the model by however long it takes to accumulate, which makes adding one to the log the cheapest decision in the whole project. ## What this stage is not It is not running a candidate model quietly beside the rule, and it is not a percentage ramp between two estimators - those are separate mechanics for a candidate that already exists. Here there is exactly one estimator serving, and the second job of the stage is bookkeeping. Keeping that distinction crisp matters in a design round: the rule-first stage is cheap precisely because it adds no second scoring path, no comparison at request time and no extra failure mode - only a durable record of what was promised and what actually happened.

  • What is the minimum you must log for the records to become a training set later?
    The parcel identifier, the estimate and window promised, the inputs as they stood at that moment, and enough provenance - ladder rung and table version - to interpret the row. The outcome joins later. Without the as-of inputs the row cannot be reproduced at prediction time, so it is not a usable training example.
  • You only log parcels that were quoted a date. What does that make the training population?
    Whatever the rule served, not the network. Lanes where the window was suppressed and orders cancelled before dispatch never enter the log, so a model fitted on it is fitted to the served slice. That is workable if the slice is stated and the gap is measured; it is dangerous when nobody knows it exists.
  • The team wants to log every field available, to be safe. Any objection?
    Volume and retention cost are the obvious ones, but the real risk is fields whose value at read time differs from decision time being treated as if they were captured. Prefer a deliberate list of as-of inputs, each with a known meaning, over a broad dump nobody can later vouch for.

saying these in an interview costs you the question

  • Assuming the warehouse can reconstruct what the rule saw
  • Logging the promise but not the inputs behind it
  • Treating delivered rows as the whole log while outcomes are still pending
  • Skipping the rule launch because the model is a quarter away
  • Forgetting that only promised parcels appear in the log