How do you turn measured cycle-time percentiles into a delivery date you will commit to, and what does that ask of stakeholders?
answer
- Averages hide a long right tail
- Answer with a probability, not a date
- Read a percentile off past finishes
- The sample must stay representative
- More confidence costs a later date
basics
~20 sQuote a percentile from measured finished-item data as a probability, not a date: most comparable items finished within this span. Higher confidence buys a later date. It asks stakeholders to accept a probability, hold scope still, and stop rewriting the priority order.
solid answer
~50 sCycle-time distributions are right-skewed, so a mean sits above the median and still leaves a large minority of items outside it — promising at the average is promising to miss. A percentile is an honest statement about a measured sample: *85 percent of items like this finished within 14.8 days*. For one item, read the percentile off comparable finished work; for a set, sample measured throughput repeatedly to project how many periods the set needs. The tradeoff is explicit — **confidence is bought with calendar time**, and the 95th percentile is often so distant it is useless as a plan. The commitment is two-sided: stakeholders must accept a probability rather than a date, hold the scope still, and stop rewriting the queue order, because every reorder changes what 'an item like this' means and quietly invalidates the sample.
code
pseudocode · 8 linessample = cycle times of the last 62 comparable finished items, in days
sorted = sort(sample)
p50 = sorted[ceil(0.50 * 62)] # 6.4 days
p85 = sorted[ceil(0.85 * 62)] # 14.8 days
p95 = sorted[ceil(0.95 * 62)] # 27.3 days
forecast = "85 percent of items like this finished within 14.8 days"
# valid only while new items resemble the sampled onesgo deeper
Recall that a forecast can be stated as a probability rather than a single date, and that cycle-time data has a long tail which an average hides.
Explain how a percentile is read off measured finished-item data and why a right-skewed distribution makes the mean a poor promise. Be ready to state what an 85th-percentile figure does and does not claim.
Show that you check the sample before quoting from it — enough items, recent enough, comparable enough — and that you use the right clock for the question being asked.
Own the organisational bargain. Explain what the method asks of stakeholders, how churn in the priority order invalidates the sample, why percentile inflation destroys the method's usefulness, and how you would prevent a published forecast from being reread as a commitment.
## Why the average is the wrong summary Cycle-time distributions are not symmetric. Most items finish quickly; a long right tail is made of the items that hit a dependency, a rework loop or a decision nobody was available to make. In a right-skewed distribution the mean sits above the median and still leaves a substantial share of items outside it. Promising a date at the mean therefore means missing for a large minority of items — and the misses are not small overruns, because they come from the tail. The average also destroys the one piece of information a stakeholder actually needs: how much of the range they are being asked to accept. "About a week" and "usually four days, occasionally five weeks" can produce the same mean and are completely different promises. ## Percentiles turn the sample into a statement A percentile read off measured finished-item cycle times is a claim about the past that is honest on its face: *85 percent of items like this one finished within 14.8 days.* Nothing is modelled, nothing is estimated, and the claim can be checked against the sample it came from. | Percentile from one measured sample | The statement it supports | What it costs | |---|---|---| | 50th — 6.4 days | half of comparable items finished this fast | wrong about half the time; useful only for planning internally | | 85th — 14.8 days | most comparable items finished inside this | roughly one item in seven still misses | | 95th — 27.3 days | a date you will very rarely miss | more than four times the median, so the promise is nearly useless as a plan | The table is the tradeoff in miniature: **confidence is bought with calendar time**, and there is no percentile that is both safe and tight. Choosing one is a judgement call about which failure hurts more — a missed date, or a promised date so distant that the work looks uncompetitive. ## Forecasting one item versus a set - **One item.** Read the percentile straight off the cycle-time sample of comparable items, starting the clock at the point you are forecasting from. - **A set of items.** Cycle time will not compose, because the items overlap. Sample the measured per-period throughput repeatedly to project how many periods it takes to finish the set, and report the same kind of percentile over those runs. This needs a stable measurement window, and it needs the scope to hold still. Both rest on the same premise: the future items must resemble the sampled ones. A request unlike anything in the sample — a first integration with an unfamiliar external system, a regulatory review nobody has been through — has no percentile, and saying so is a better answer than quoting one anyway. ## What the arrangement asks of the people receiving it This is the part candidates skip, and it is what makes the question a leadership question rather than a statistics question. A percentile forecast is a two-sided deal: 1. **They must accept a probability instead of a date.** "85 percent chance by the 14th" is a different object from "the 14th", and an organisation that reflexively rounds it to the latter gets all the cost of the method and none of the benefit. 2. **They must leave the sample valid.** On the construction-site safety product, a founder rewrote the priority order most weeks. Every reorder pushes uncommitted work back and changes what "an item like this one" means, so the measured distribution stops describing the future. The forecast degrades not because the team slowed down but because the queue policy changed underneath it. 3. **They must let the scope hold still.** A set forecast is a forecast of a defined set. Adding to it silently is the same as adding to the date, and the arithmetic should be shown rather than absorbed. 4. **They must accept that the tail is real.** The 15 percent that falls outside the 85th percentile is not a failure of the team, and treating each occurrence as one destroys the incentive to report honest numbers. ## Failure modes worth naming out loud - **Percentile inflation.** Quoting the 95th for everything makes you technically reliable and practically ignored. - **False precision.** A sample of eleven finished items does not support a 95th percentile; say the sample is too thin instead. - **Stale windows.** A distribution measured before a reorganisation, a change of stages or a shift in item mix describes a system that no longer exists. - **Forecasting the wrong clock.** If the stakeholder is asking from the moment they raised the request, the answer must come from lead-time data, not from cycle-time data that starts when work begins. - **Quiet substitution.** Publishing a percentile and then being held to it as a commitment is worse than never publishing it, because the next forecast will be padded. The honest summary a lead gives is short: the method converts an argument about optimism into an argument about evidence and risk tolerance, and it works only for as long as the system it measured stays recognisable.
- A founder reorders priorities most weeks. What does that do to a percentile forecast?It erodes the sample's meaning. Each reorder pushes uncommitted work back and changes the mix of what gets started, so the measured distribution stops describing what will happen next. The forecast degrades even though the team has not slowed down, and saying so out loud is how the reordering cost becomes visible instead of being blamed on delivery.
- Which percentile would you publish, and why not simply always use the 95th?The 85th is usually the workable compromise. The 95th is technically reliable and practically ignored, because it can be four times the median, so plans built on it look uncompetitive and people quietly discount them. Choosing is a judgement about which failure costs more — a missed date, or a promise so padded that nobody uses it.
- How do you forecast a set of twenty items rather than a single one?Not by adding cycle times, because items overlap. Sample the measured per-period throughput repeatedly to project how many periods the set takes, then report a percentile across those runs. It needs a stable window and a scope that holds still — silently adding items to the set is the same as moving the date.
- A request is unlike anything in your sample. What do you tell the stakeholder?That there is no percentile for it, and why. The method is a statement about how comparable past work behaved, so a first integration with an unfamiliar external system has no comparable history. Saying the sample does not cover it, and offering to start it and measure, is more useful than quoting a number that means nothing.
saying these in an interview costs you the question
- Quotes the average as a delivery date
- Presents a percentile as a guarantee
- Computes a high percentile from a handful of items
- Forecasts from a sample measured before the workflow changed
- Uses cycle-time data to answer a lead-time question
- Pads to the 95th percentile for everything by reflex