skip to content

A flow board's longest queue sits at one stage - how do you turn that into one measured improvement?

level: seniorimportance: must knowfreq 55%

answer

  1. An observation is not yet a diagnosis
  2. The hypothesis must name a policy
  3. Exactly one variable at a time
  4. Baseline and prediction written before the change
  5. Enough finished items for a median

basics

~20 s

Pick the single stage the data implicates, state a hypothesis about which policy causes the queue, change that one policy, and agree beforehand which measure should move, in which direction and by when. Judge it on enough finished items to mean something, then keep, adjust or revert.

solid answer

~50 s

Treat it as an experiment with one variable. Start from the observation rather than a feeling: this stage holds most of the work in progress, and items sit there longest. Form a hypothesis that names a policy - pulling into review is optional, so it happens last; the entry criteria are unwritten, so work bounces back. Change **exactly one** explicit policy, leaving everything else alone, because two changes at once make the result unattributable. Before the change, write down the measure you expect to move, its current baseline, the direction and roughly by when - most usefully the median time from start of work to delivery. Then wait for enough finished items that a median means something; three or four do not. Read the result at the service delivery review and decide honestly: keep it, adjust it, or revert it and say so.

code

pseudocode · 10 lines
pseudocode
EXPERIMENT CARD - wine-cellar inventory service

observation  review holds 7 of the 11 items in progress
hypothesis   review is optional to pull, so it happens last
change       cap the review column at 3 items; nothing else
measure      median cycle time, start of work to delivery
baseline     9.0 days over the previous 41 finished items
predicted    below 7.5 days
judged at    the next 30 finished items
outcome      6.5 days over 34 items - keep the cap

go deeper

for a junior

Know that an improvement starts from something visible on the flow board - a stage where work piles up - and that the team changes one rule and then checks whether the same number moved.

for a middle

Explain the single-variable discipline and why it matters: two changes at once make the outcome unattributable, and the baseline plus the prediction have to be written down before the change, not after.

for a senior

This is the question you are expected to answer from experience. Walk through a real change end to end - observation, policy hypothesis, the one change, the baseline and window, the result - and include one that you reverted.

for a principal

Own the credibility of the evidence: sample size, contaminated windows, definitions that shifted mid-measurement, and how you stop a team from reporting only the experiments that worked.

## From a signal to a hypothesis A queue on a flow board is an observation, not a diagnosis. The improvement work begins by turning it into a statement precise enough to be wrong. 'Review is slow' is not that statement; 'eleven of our fifteen items in progress are sitting in review, and items entering review wait longer than they spend being worked on' is, because it names a stage, a count and a comparison anyone can check on the board. The hypothesis then has to name a **policy**, not a person or a mood. Policies are the things a team can actually change: who may pull work into a stage, what a stage's entry and exit criteria are, how much work in progress a stage may hold, what happens to an item that fails a check, which class of work may jump a queue. 'People are busy' cannot be changed. 'Pulling an item into review is optional, so it happens after everything else' can. ## Change exactly one policy The single-variable discipline is what separates improvement from redecorating. Change the work-in-progress cap at the queueing stage *and* rewrite its entry criteria *and* add a lane for urgent work, and any subsequent movement in the numbers is unattributable - which matters most when the numbers get worse, because the team cannot tell which change to undo. A single change is also cheap to revert, and reverting must be a socially acceptable outcome. A team where reverting looks like failure will quietly keep bad policies alive to protect whoever proposed them. ## Decide the measure before the change Write this down before touching the board, because deciding afterwards is how a team convinces itself of anything: - **Which measure.** Usually the median time from the start of work to delivery for the affected class of work; sometimes the count of items finished per week, sometimes days lost while blocked. Name where its clock starts, since a measure timed from request and one timed from start of work answer different questions. - **The baseline.** The value over a stated number of recently finished items, not a remembered impression. - **The direction and rough size.** A prediction that cannot fail is not a prediction. - **When it is judged.** Expressed in finished items where possible, because a calendar window can contain almost no completed work. ## A worked example A team running a wine-cellar inventory service had a median cycle time - measured from the start of work to delivery - of 9.0 days across its last 41 finished items, and its board showed 7 of the 11 items in progress parked in review. The hypothesis was that review was optional and therefore always last. The change was one line of policy: cap the review column at 3 items, so nobody may start new work while review is full. Nothing else moved. The prediction was a median below 7.5 days, judged over the next 30 finished items. The actual result was 6.5 days over 34 items, and the cap stayed. Note what makes this credible: one variable, a stated baseline, a falsifiable prediction and a sample large enough that a median is not an accident. ## What makes an improvement claim fail review 1. **Too few items.** A median over four items moves wildly on its own; a change judged that way is judged by noise. 2. **A contaminated window.** A two-week holiday shutdown inside the measurement window inflates every duration that spans it. Either exclude that period or extend the window and say you did. 3. **A measure chosen afterwards.** If three measures were watched and the one that moved is reported, the result means nothing. 4. **A moved definition.** Redrawing the stage boundary while measuring the time to cross it makes before and after incomparable. 5. **No stated failure condition.** If no outcome would have caused the change to be reverted, no experiment was run. ## When the queue is not the constraint Sometimes the stage with the longest queue is only where the symptom is visible. Work piles up in review because items arrive far larger than a reviewer can absorb, and the real change is upstream in how work is split. The test is cheap: make the change, watch the measure, and if the queue simply relocates to the next stage, the hypothesis was wrong and the board just told you so. That is a successful experiment with a negative result, and saying that plainly in an interview reads far stronger than a list of changes that all supposedly worked.

  • How many finished items do you want before you believe a median moved?
    Enough that one unusual item cannot swing it - in practice a few dozen for most teams, and always stated up front rather than chosen once the answer is known. If the team finishes too little for that, judge the change on a coarser signal such as whether the queue re-forms, and say explicitly that the number is indicative rather than evidence.
  • A two-week holiday shutdown falls inside your measurement window. What do you do?
    Say so and handle it rather than quietly reporting the number. Every item whose life spans the shutdown carries two extra weeks that have nothing to do with the policy, so either exclude items spanning it, or extend the window until enough items finished outside it. Reporting the inflated median as a failed experiment would discard a change that may have worked.
  • The queue moved to the next stage instead of disappearing. Was the experiment a failure?
    No - it was informative. Relocating a queue means the constraint was not where the hypothesis put it, or that the stage downstream now receives work faster than it can absorb. Keep or revert the change on its own merits, then form the next hypothesis about the stage the work moved to. A negative result cheaply obtained is exactly what a small reversible change is for.

saying these in an interview costs you the question

  • Changes several policies at once and claims credit
  • Judges the result on three or four finished items
  • Picks the measure after seeing which one moved
  • Blames people rather than naming a policy
  • Never states what would have counted as failure
  • Ignores a holiday period inside the measurement window