skip to content

A new parcel lane carries about forty shipments a week - why is its own median transit time a weak baseline?

level: middleimportance: should knowfreq 45%

answer

  1. few hundred parcels is thin evidence
  2. whole-day medians hop between integers
  3. set a minimum support per cell
  4. back off lane to region to network
  5. record which rung answered

basics

~20 s

Forty parcels a week is too little evidence to key an estimate on: transit days are whole numbers, so a small shift in the week's mix flips which day sits in the middle and the quoted median hops. A support floor with a backoff to a coarser key fixes it.

solid answer

~50 s

A lane cell computed over an eight-week window holds about 320 delivered parcels at that volume, and the mix inside it - which shippers, which dispatch hours, which day of week - turns over fast on a young lane. Because transit times are whole days, the middle order statistic moves between adjacent integers as the mix shifts, so the promise jumps from two days to three and back with nothing real behind it. The fix is part of the baseline's definition, not a patch: set a **minimum support** per cell, and where the cell is below it back off to the region pair at the same service level, then to the service level network-wide, recording which level answered. Carry-forward covers a cell that was healthy last refresh and is thin this one. A design that quotes a per-lane median without naming its support floor has not specified the baseline.

code

pseudocode · 16 lines
pseudocode
WINDOW_DAYS  = 56
MIN_SUPPORT  = 500

function transitDays(lane, serviceLevel):
    sample = deliveredParcels(lane, serviceLevel, WINDOW_DAYS)
    if size(sample) >= MIN_SUPPORT:
        return median(sample.transitDays), "lane"

    sample = deliveredParcels(regionPairOf(lane), serviceLevel, WINDOW_DAYS)
    if size(sample) >= MIN_SUPPORT:
        return median(sample.transitDays), "region"

    sample = deliveredParcels(ALL_LANES, serviceLevel, WINDOW_DAYS)
    return median(sample.transitDays), "network"

// new lane: 40 per week x 8 weeks = 320 < 500, so the region rung answers

go deeper

for a junior

Recall that an estimate computed from very few observations is unreliable, and that the design needs a defined fallback so every parcel still gets a promised window.

for a middle

Explain the mechanics: a support floor per cell, a ladder from lane to region pair to service level, and carry-forward across refreshes - and why whole-day medians move in jumps on thin cells.

for a senior

Show that you record which rung answered and the age and support of each value, so a strange promise can be traced, and treat the share of traffic answered by coarse rungs as an operational signal.

for a principal

Weigh what pooling across lanes is actually worth: a ladder buys most of it for a lookup, so the argument for a learned estimator has to rest on measured structure in the residuals, not on the existence of thin lanes.

## Why forty a week is not a key you can estimate on The transit table is keyed by (origin region, destination region, service level) and each cell holds the median transit time of parcels **delivered** inside a rolling window. On a mature lane that window contains tens of thousands of parcels and the median is stable to the hour. On a lane opened last quarter it contains a few hundred, and three things go wrong at once: - **The estimate is coarse and jumpy.** Transit times are whole days. The median is the middle observation, so it lands on an integer, and a modest change in the week's mix moves it a whole day at a time rather than a little. The quoted promise therefore hops between two and three days while the lane itself has not changed. - **The mix is not stable.** A young lane's volume is concentrated in a few shippers and a few dispatch windows. When one of them changes their cut-off time, the cell's composition changes with it, and the median follows the mix rather than the network. - **The window is truncated.** Only parcels that have already been delivered are in it. A parcel still moving slowly is not yet counted, so a lane that is degrading looks fine for as long as its slow parcels are in flight. This bites hardest exactly where volume is thin, because a single slow batch is a large share of the cell. None of this makes a heuristic baseline wrong. It makes an **unqualified per-cell median** wrong, and the qualification is the interesting part of the design. ## The backoff ladder The standard shape is a ladder of keys from specific to general, with a support floor at each rung: | rung | key | when it answers | |---|---|---| | lane | origin region, destination region, service level | cell has at least the support floor of delivered parcels in the window | | region | origin-destination region pair, service level | lane cell is below the floor | | network | service level only | region cell is also below the floor | Two properties matter more than the exact thresholds: 1. **Every query gets an answer.** Checkout has to print something. The ladder decides *which estimate* answers, never *whether* to answer. 2. **The answering rung is recorded with the estimate.** An arrival window whose provenance is unknown cannot be debugged, and the rung is also the first thing a later model would want as a feature and as a slice for measurement. With a 56-day window and a floor of 500 delivered parcels, the new lane supplies 40 x 8 = 320 observations, falls short, and is answered by its region pair. That is not a failure of the lane - it is the ladder doing its job, and the promise it produces is more defensible than a median over 320 parcels would have been. ## Carry-forward, and what it hides Carry-forward is the other half of the same idea across time: if a cell was above the floor at the last refresh and is below it now, keep the last value rather than recomputing from thin evidence. It is cheap and usually right, because transit times persist week to week. Its cost is silence. A cell holding the same value for a month looks identical whether the lane is stable or has simply stopped producing evidence, and nothing in the quoted number distinguishes them. The design should therefore carry, alongside each estimate, the age of the value and the support behind it, so an operator can see a stale cell rather than infer it. Carry-forward is a bridge over a gap in data; it is not a substitute for data. ## What this means for the case for learning Thin lanes are also the honest argument against reaching for a model on day one. A learned estimator has the same evidence problem - it cannot know a lane it has barely seen - and it answers it with pooling and shared structure, which is a real advantage but a modest one on a handful of lanes. Meanwhile the ladder already delivers most of that pooling for the price of a lookup, with no labelled training set behind it. The design answer is therefore: state the floor, state the ladder, record which rung answered, and let the thin lanes accumulate evidence under a rule that does not pretend to know more than it does. If, months later, the region rung is answering a large share of traffic and the residuals on those parcels are structured, that is a concrete argument for learning - measured on the rule's own logs rather than asserted at the whiteboard.

  • How would you pick the support floor rather than guessing at 500?
    Replay it. For a range of floors, recompute the historical promises on mature lanes and read how the error and the week-to-week movement of the estimate behave, including on cells deliberately thinned. Pick the smallest floor where the estimate stops hopping between whole days without pushing most traffic onto the coarse rungs.
  • The region rung now answers a third of all parcels. Is that a problem?
    It is a signal, not a fault. It says a large share of promises are made from a key coarser than the lane, which is where a pooled estimator would have the most to add. Track that share as an operational number: rising means the network is opening lanes faster than evidence accumulates.

saying these in an interview costs you the question

  • Quoting a per-cell median with no minimum support behind it
  • Claiming a few slow parcels drag a median upward
  • Letting thin cells return no estimate, leaving checkout with nothing
  • Carrying a value forward for months without exposing its age
  • Assuming only a model can handle a lane with little history