Tell me about a time an estimate you gave turned out to be badly wrong.
answer
- the number you gave, out loud
- what the work actually turned out to be
- the unknown you missed, named
- the day you raised the slip
- the calibration habit you kept
basics
~10 sTests calibration and honesty under uncertainty. Name the number you gave, the number it became, the specific unknown you missed, how early you raised it, and the estimating habit you changed afterwards.
how to answer
5 beats- the estimate you gave, and who was planning around itOpen with the actual number and the audience in one or two sentences. Keep this beat and the next to roughly a fifth of your airtime; the setup is not where the signal is.
- what the work turned out to actually involveName the specific thing you had not priced — an approval path, an unfamiliar system, a dependency queue. One concrete unknown beats a list of vague complications.
- the moment you knew, and what you did that dayThis is the bulk of the answer, around sixty percent. Say when you detected the slip, how you re-forecast, who you told, in writing or not, and what you did to attack the blocker rather than wait it out.
- what actually shipped, with the gap stated plainlyGive the delivered number next to the estimate without softening it, and add one outcome measure so the interviewer knows the work still mattered. Around a fifth of the airtime covers this and the reflection.
- the estimating habit you kept afterwardsClose with a change to what you inspect before quoting a number, not a promise to add buffer. Make it specific enough that the interviewer can imagine you doing it next week.
your answer
5 story prompts- Pick an estimate you personally gave out loud, ideally within the last eighteen months.
- Write down both numbers now: what you said, and what it actually took.
- Name the single unknown that explains most of the gap, in one sentence.
- Find the day you first knew, and who you told that day.
- This can be your missed-deadline story re-angled onto the estimate itself.
draft and rehearse your own answer in a learn session
go deeper
The interviewer is probing calibration and ownership under uncertainty: whether you can distinguish what you actually knew from what you assumed, whether you surface a slip early rather than absorbing it silently, and whether a miss changed your method. A strong answer proves you are safe to give an estimate to, because your estimates come with their assumptions attached.
In my first year on a seven-person infrastructure team at an early-stage startup, I picked up a ticket to move our disk-space alerts off a fixed threshold and onto a rate-of-fill check, because that single rule was producing most of our night-time noise. I estimated three days. It took eleven. Two days in I had the rule written and tested against recorded data. Then I found out the alert definitions lived in a repository owned by the data platform group, and any change there needed a review from whoever was on their rotation. Their reviewer was away, and my change sat for nine days before it merged. What I got wrong was estimating only the part I could see, which was writing the rule. I never asked where it deployed from or who had to approve it. What I got right was not sitting on it quietly: on the fourth day, once I realised the review might be slow, I posted the revised date in our on-call channel and asked our rotation lead whether it was worth escalating. The change did work. Duplicate disk pages went from twenty-three a week to four. Since then, before I give any number, I write two lists: what I will do myself, and what I need from someone outside our team. If the second list is not empty, I give a range and say out loud what would collapse it to the low end.
The signal here is the diagnosis and the fourth-day update, not the size of the miss — a junior can credibly own both. Naming the unseen approval step and converting it into a two-list habit is what makes it more than an apology. Staying silent until day eleven would downlevel it immediately.
By my second year I owned on-call quality for our infrastructure group. I committed to five weeks to move alerting for eleven services off the self-hosted alert manager we had outgrown. It landed in nine. The decomposition itself was sound — a ticket per service, and I had already done two of them as a trial run. What I had not priced was that four of those services registered their alert receivers inside a shared traffic module another squad owned. Each one needed a change request into that squad's queue, and their queue was clearing about one request a week. I saw it in the second week. I rebuilt the plan as two tracks: the seven services I could finish alone, and the four gated on that squad. Then I re-forecast in writing to my manager and our product lead as a range, six to ten weeks, with the single condition that decided which end we hit. I also asked to join that squad's planning session so the four changes could go in as one batched request instead of four separate ones. We finished at nine weeks. Paging for the group dropped from sixty-two a week to twenty-one, and the gated services landed last, as forecast. What stuck is that I now size a dependency by the throughput I can observe rather than the date someone offers me, and I put that assumption in the estimate where people can argue with it.
Middle scope shows up in the re-forecast: a written range, the deciding condition, and a move to change the other squad's batching rather than just reporting the delay. The habit at the end is a method change, not padding. Reporting the slip only at week nine would drop this a level.
One task you owned is enough. Show that you estimated the part you could see, name the part you could not, and prove you raised the new date early rather than hoping to catch up quietly.
The miss should be feature-sized and the diagnosis should reach dependencies: work you needed from people outside your control. Show a written re-forecast with a range, and show you acted on the blocker instead of only reporting it.
You were committing on behalf of a team, so the cost of the miss lands on other people. Show how you renegotiated the commitment, what you protected, and the change you made to how the team estimates afterwards.
Talk about a forecast across several teams and a systemic cause — a class of unknown nobody was pricing. The durable outcome is a mechanism others now use, not a lesson you personally learned.
saying these in an interview costs you the question
- No numbers at all — only 'it took longer than expected'
- Putting the entire miss on another team with no ownership
- Discovering the slip at the deadline and telling nobody before then
- Framing it as bad luck rather than an assumption you never checked
- Ending on 'I add more buffer now' with no change to the method
- Choosing a miss so small the interviewer learns nothing about your judgment
- When did you first know you were going to miss it?Give a specific point in the work, not a vague 'partway through', and say what the signal was. Then say who you told and through what channel. The gap between knowing and telling is the thing being measured here, so if it was long, own it plainly and say what you now do differently.
- What number would you give if you had to estimate that work again today?Answer with a range and the condition that decides it, and explain what you would check before quoting anything. This is a live calibration test: repeating your original number unchanged suggests you learned nothing, and a wildly inflated number suggests you replaced judgment with padding.
- How did the people depending on that date find out?Describe the actual update — where you wrote it, who read it, and whether you offered options alongside the new date. Strong answers show the update travelling to everyone who had planned around the old date, not just to your manager in a one-to-one.
## This prompt is a calibration test, not a confession Everyone has a bad estimate, so the size of the miss carries almost no signal. What varies between candidates is what they knew, when they knew it, and what they did the moment the picture changed. Two failure modes bracket the answer: - treating it as **a trap** and offering a miss so trivial it proves nothing (an afternoon that took a day), - or treating it as **a confession** and spending most of the airtime on how bad it felt without ever reaching the method. ## Common wordings - 'Describe an estimate that turned out to be wrong.' - 'Tell me about a project that ran over.' - 'When were you most wrong about how long something would take?' - 'Tell me about a time you missed a deadline.' These are one prompt, but the deadline wording needs more care: a missed deadline can plausibly be someone else's doing, while a missed estimate is squarely yours to explain. If you get the deadline wording, still answer as though the estimate is the subject — that is what the interviewer is actually probing. ## Weak and strong on identical facts - **Weak:** 'I said two weeks, it took five, the other team was slow.' Every cause sits outside the speaker and nothing generalises to the interviewer's team. - **Strong:** 'I said two weeks. I had estimated the code and not the review path, and the review path ran through a queue whose throughput I had never measured. On the fourth day I could see the shape of it, so I re-forecast in writing with a range and the one condition that decided it, and I went to change how that queue batched our requests rather than waiting on it. It took five.' Same events, and now the listener can predict your behaviour under a slipping plan. ## The planning fallacy is the real subject People estimate the path where nothing goes wrong, because that path is the only one they can picture in detail. The antidotes are concrete and worth naming aloud: - compare against a **reference class** of work you actually finished rather than to your mental model of the task; - separate what you will do yourself from what you need from others and price the second by **observed throughput**, not by promises; - and treat an unknown as something to buy down with a **timeboxed spike** rather than something to cover with padding. ## Evidence that carries weight 1. The original number said out loud. 2. The delivered number. 3. The one unknown that explains most of the gap, named precisely. 4. The moment of detection and the channel of the update. 5. The durable change to how you estimate. Five items, and an answer missing the fourth is the one interviewers most often downgrade, because **early honest re-forecasting** is the behaviour they are hiring for. ## Two traps in the reflection beat - **The first is the buffer-only lesson:** 'I multiply by two now.' Padding hides uncertainty instead of exposing it, and it fails the moment someone asks what the padding is for. - **The second is over-correction:** a candidate who now quotes deliberately pessimistic numbers has swapped one calibration error for another and will be caught by the follow-up asking what they would estimate today. The lesson that lands is a change to what you inspect before you quote.