skip to content

Estimation and Velocity

Relative sizing with story points, Planning Poker, T-shirt sizing, velocity, and capacity planning per sprint — plus why Scrum steers away from hour-based estimates. Expect to explain what velocity is good for and why it is not a productivity metric.

on this pageshow

questions

5

Why do Scrum teams size backlog items in story points instead of hours?

level: juniorimportance: must knowfreq 76%

answer

  1. Compare, do not measure
  2. One number, three ingredients
  3. Complexity, effort and uncertainty together
  4. The scale is local to one team
  5. Hours read as an individual promise

basics

~20 s

Story points size an item relative to other items, folding complexity, effort and uncertainty into one number on that team's own scale. Hours invite false precision, vary with whoever picks the item up, and turn a shared forecast into an individual promise.

solid answer

~40 s

A story point is a **relative** size: this item is about twice the one the team called a 2. The number folds together **complexity, effort and uncertainty**, so the discussion stays on the work rather than on whose day is being spent. Hour figures fail on three counts — they change depending on who does the item, they read as a commitment to a date, and they invite precision nobody has, since no one can honestly separate 13 hours from 16. Relative sizes are also faster to produce, because people compare well and measure badly. Points mean something only **inside one team**; there is no exchange rate between teams. Over several Sprints the team learns how much it typically finishes, and that history, not the raw number, is what makes a forecast possible.

go deeper

for a junior

Be ready to say what a story point bundles together — complexity, effort and uncertainty — and to state plainly that it is relative to other items rather than convertible into hours.

for a middle

Explain the mechanics: how one reference item anchors the whole scale, why the ladder of sizes widens at the top, and why the item is sized once by the group rather than once per possible assignee.

for a senior

Expect to defend the practice under pressure. Show how you stop sizes being quietly converted into dates, and what you do when an item comes back sized far away from the reference it should have matched.

for a principal

Own the organisational angle: story points are a local currency, so any report that adds or ranks them across teams is measuring nothing. Be ready to offer a delivery forecast that never needed a shared scale.

## What a story point is measuring A story point is a **relative size**. When a team calls an item a 5, it is saying: this is about as big as the other things we have called 5, and roughly two and a half times the thing we called 2. The number carries no unit, no hours, and no meaning outside the team that produced it. Three ingredients go into that single number: - **Complexity** — how intricate the work is: how many moving parts, how many places the change touches, how much design thinking is needed before anyone types. - **Effort** — sheer volume: how much repetitive, well-understood work remains once the thinking is done. - **Uncertainty** — how much the team does not yet know: an unfamiliar integration, a partner whose data shape nobody has seen, an area of the system with no tests around it. Bundling all three is deliberate. A dull but perfectly understood job and a small but unmapped one can both come out as a 5 for entirely different reasons, and nothing in planning requires the team to separate those reasons. Planning needs a comparable size, and that is what it gets. ## Why hour estimates behave badly here | | Hour estimate | Story point | |---|---|---| | Whose number is it | The person expecting to do the work | The Developers as a group | | What it implies | A date someone can be held to | A size, converted into a forecast later | | Precision suggested | To the hour | To the nearest bucket | | How it ages | Invalid as soon as someone else picks it up | Still valid; the item is still the same size | | What it needs to be useful | Nothing but confidence | Several Sprints of history | One mechanism sits behind every row: an hour figure is a personal prediction dressed up as a team plan. Two developers who agree completely about what the work is will still disagree about the hours, because one of them has worked in that area and one has not. The argument that follows is about people, not about the item. Relative sizing sidesteps this because humans compare far better than they measure. Asked to guess the height of a building, most people are badly wrong. Asked which of two buildings is taller, almost everyone is right. Sizing uses the reliable skill and skips the unreliable one. ## The scale is local, and that is the design Story points mean something only inside the team that set them, and there is no exchange rate. A team that anchored its scale by calling a small familiar item a 3 will produce numbers around 50% larger than a team that called the same item a 2, and neither team is wrong. Two consequences are worth saying out loud: 1. **Never compare teams by story points.** Any report that adds them across teams is adding quantities with different units and reporting the sum as if it meant something. 2. **Never convert points back into hours for a status report.** The conversion re-imports every problem the unit was chosen to avoid, and the result looks far more precise than anything the team ever claimed. Most teams use a widening ladder of sizes — 1, 2, 3, 5, 8, 13 — rather than every whole number. The gaps widen because confidence falls as items get bigger. The difference between a 3 and a 5 is real and worth discussing; an argument between 21 and 22 is theatre. The top of the ladder earns its place too: an item that will not fit under the largest size a team uses is not really a number, it is a signal that the item is not yet understood well enough to enter a Sprint. ## What the practice buys, and where it stops helping The benefits, in the order they actually matter: - **The conversation.** Most of the value arrives before any number is written down, at the moment two people discover they were imagining different work. - **Speed.** A team can size a dozen items in the time one defensible hour breakdown takes to build. - **A forecast.** Consistent sizes plus several Sprints of history give a range for how much the team can take on — which is what anyone wanting the hours actually wanted. Relative sizing stops helping in three situations. A brand-new team has no reference items and no history, so its early numbers are guesses with a nicer unit. Work genuinely unlike anything in the team's history cannot be triangulated at all, and the honest response there is a time-boxed investigation rather than a size. And in any environment where the numbers are harvested as a performance measure, the scale inflates — because the people being measured are the same people who set the unit — and once it inflates, the history it produces is worthless for the one job it had.

  • If points are relative, what fixes the scale before a team has any history at all?
    The team picks one small, well-understood item everyone remembers and calls it a 2 or a 3 — deliberately not a 1, so there is room to go smaller later. Everything else is sized against that reference. The absolute value is arbitrary; only the ratios matter, and the scale settles into something stable after a few Sprints of use.
  • Why do most teams use a widening sequence of sizes rather than every whole number?
    Because confidence drops as items get bigger, so the gaps should widen with it. Choosing between 5 and 8 is a real distinction a team can argue about; choosing between 21 and 22 is invented precision. Coarse buckets at the top make the imprecision visible and push the team to say 'too big to size yet' instead of inventing a number it cannot defend.

Rating hills on a route as easy, moderate or brutal is quick and reliable; predicting the exact minutes each will cost you is neither, and the ratings are what let you plan the ride.

saying these in an interview costs you the question

  • Says one story point equals a fixed number of hours
  • Sizes in points, then converts straight back to days for a date
  • Treats the number as a measure of how hard someone worked
  • Compares two teams' point values as if the scales matched
  • Gives a different size depending on who will pick the item up
  • Insists a precise hour figure is always the more honest answer
open as a page

Why is team velocity a forecasting input rather than a productivity metric?

level: middleimportance: must knowfreq 71%

basics

~20 s

Velocity counts the story points one team brought to its Definition of Done per Sprint, averaged over several Sprints. Its only job is forecasting what that team can take on next. The scale is local, so velocity compares nothing across teams — and it inflates the moment it becomes a target.

open as a page

In a story-point sizing session, why does everyone reveal their number at once?

level: middleimportance: should knowfreq 47%

basics

~20 s

Simultaneous reveal blocks anchoring. Once a senior voice says 'that is a three', the rest converge on three and the disagreement that carried the information disappears. A wide spread is the signal: it marks a hidden assumption worth surfacing before any size is agreed.

open as a page

How do you plan a Sprint's capacity when historical velocity assumes a full team?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Velocity is history from past Sprints; capacity is what this particular Sprint can hold. Scale the historical range by the person-days actually available — holidays, part-time allocation, on-call and support duty — then let the Developers select against the Sprint Goal rather than filling up to a number.

open as a page

When would you stop sizing items and forecast a delivery date from counted history?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Stop sizing when backlog items are already broadly similar in size and enough Sprints of history exist that counting finished items forecasts as well as summing story points does. The cost is the sizing conversation, where hidden assumptions surface, and any forecast for genuinely novel work.

open as a page