skip to content

A correct fix makes today's figures disagree with every number already published — do you restate the history or fork the series?

level: principalimportance: should knowfreq 39%

answer

  1. two honest outcomes, one rule
  2. defect or redefinition
  3. can the past inputs still reproduce it
  4. labelled break versus one definition

basics

~20 s

Restate when the old figures were wrong and the past inputs can still reproduce them; fork with a labelled break when the definition changed rather than the answer being wrong. The deciding question is what readers compare across the seam.

solid answer

~50 s

There are three honest outcomes and no fourth. **Restate**: recompute the affected past periods with the new logic and replace the stored figures, so the whole series sits on one definition. **Fork**: leave published history as it stands, start the new definition at a dated boundary and label the break loudly enough that nobody computes a change across it. **Publish both** for a transition window, then retire the old — the most expensive and the least disruptive. The first question is whether the old numbers were *wrong* or merely *differently defined*: a defect argues for restatement, a redefinition argues for a labelled break. The second is whether history can be reproduced at all — if the inputs for those periods have since changed, a recompute moves the figures for two reasons at once and you can no longer attribute either.

go deeper

for a junior

Recall the two basic options: recompute the past so everything sits on one definition, or leave the past alone and mark the point where the definition changed. Silence about which one happened is the worst outcome.

for a middle

Explain why a defect and a redefinition call for different remedies, and why recomputing past periods assumes the inputs for those periods still exist unchanged.

for a senior

Demonstrate the attribution step: recompute the old logic over today's inputs first, so input drift is separated from the fix before anyone looks at a restated number. Then say where the seam gets recorded.

for a principal

This is your call to own. Weigh past use of the wrong figure, comparability across the seam, reproducibility of history and copies you cannot rewrite; choose; and convert the choice into a standing rule so the next occurrence is decided rather than debated.

## The situation A change is correct. Everyone agrees it is correct. It also means that the figure for last Tuesday, and every Tuesday before it, is different from what was published and from what people have already quoted in their own reports. This is not a deployment problem and it does not have a technical answer. Someone has to choose what the series is going to say, and defend that to the people who used the old number. ## Three outcomes 1. **Restate.** Recompute the affected past periods with the new logic — running today's logic again over input from days already processed, so the stored result for those days is replaced — and publish the corrected figures. The series ends up on one definition throughout. How that recompute is scheduled, and how a long stretch of history is pushed through a continuous job, are separate subjects; the decision that it is owed is this one. 2. **Fork.** Leave the published history untouched, start the new definition at a stated boundary, and mark the break in the data itself so that any comparison spanning it is visibly crossing a definition change. 3. **Publish both for a window.** Run both definitions in parallel until readers have moved their own derived numbers, then retire the old. It costs the most and disrupts the least, and it only works with an end date agreed at the start — otherwise you have permanently doubled the surface you maintain. ## The first question: wrong, or merely different? This single distinction decides most cases. - If the old figures were **wrong** — a defect, a mis-joined record, a filter that dropped rows it should have kept — then people made decisions on numbers that were not true, and leaving them standing is a choice to keep publishing something false. That argues for restatement. - If the **definition changed** — the measure now excludes internal traffic, or counts a different event, or uses a different denominator — then the old figures were correct answers to a different question. Overwriting them destroys a true record and makes the past unreproducible. That argues for a labelled break. The common failure is treating these as the same kind of change because both produce a different number. ## The second question: can history be reproduced at all? Restatement quietly assumes the past can be recomputed, and that assumption is often false: - the inputs for those periods must still exist, be retrievable, and be **unchanged** - if upstream inputs accept late corrections, the input you would read today is not the input the original run read — so recomputing moves the figure for two reasons at once, the fix and the input drift, and neither can be attributed - reference data that the logic joins against may itself have changed, silently re-deriving history under a present-day view of the world The diagnostic is simple and worth naming in an interview: first recompute the **old** logic over today's inputs. The gap between that and the originally published figure is input drift; whatever remains after the new logic runs is the fix. If the drift is large, restatement is not the clean act it appeared to be. ## The third question: what do readers do with the series? | how the series is read | restating | forking with a labelled break | |---|---|---| | period-over-period change | preserves comparability; the natural choice | every comparison spanning the break is silently wrong | | current level only | expensive for little gain | cheap and adequate | | external or contractual reporting | may be required, and may itself require disclosure | usually unacceptable without agreement | | feeding downstream calculations | forces those downstream numbers to move too | leaves them consistent but on a stale definition | One more consideration cuts across all of them: **the number lives in places you cannot rewrite.** It is in a deck, a spreadsheet, a board pack, somebody's own copy of your table. A restatement in your store does not reach any of those, so restating without telling anyone produces a population of readers holding two irreconcilable numbers and no idea why. What notice is owed, and to whom, is a contract question owned elsewhere; that it is owed is part of this decision. ## Set the standing rule Whatever you choose, the expensive part is re-arguing it next quarter. Write the rule down: 1. Classify every such change as a **defect** or a **redefinition** before anyone argues about remedy. 2. Bind the remedy to the classification in advance — defects restate, redefinitions fork — so the argument happens once. 3. Keep the superseded figures readable rather than deleting them, so what was published remains recoverable. 4. Record the seam or the restatement with the numbers themselves, never only in a message. And know what is not yours to settle here: the notice owed to downstream consumers and the contract it sits in, tracing who across the platform is affected, and how restated history is represented in the model with rows carrying their own validity dates are each owned by other disciplines. This decision is the one that comes first and tells them what to do.

  • You restate, and the recomputed figures differ from the published ones by more than the fix explains. What happened?
    The inputs themselves changed: late-arriving records, upstream corrections, reference data re-derived under today's view. Separate the causes by first recomputing the old logic over today's inputs — the gap from the published figure is input drift, and only the remainder is attributable to the fix.
  • Why is just fork it, nobody looks at last year a dangerous default?
    Because the reader who does look is almost always computing a year-over-year change, which is precisely the comparison a break destroys. And unless the break is marked in the data, it is invisible: the comparison returns a number, it is simply the wrong number.
  • When is publishing both definitions in parallel worth its cost?
    When the series feeds external or contractual reporting, or when readers maintain their own derived figures and need time to migrate. It costs running and explaining two definitions at once, so it needs a retirement date fixed at the start or it becomes permanent by inertia.
  • Does it matter whether anyone has already acted on the wrong figure?
    Yes, and it usually dominates. A wrong figure nobody used is a tidy-up; a wrong figure that drove a decision, a payment or an external statement makes restatement far harder to decline. The blast radius of past use, not the size of the numeric difference, is the real weight.

A sports league finds that a scoring rule was applied wrongly all season. It can re-score every past match so the table is consistent throughout, or leave the results as they were played and apply the corrected rule from next season with a note in the record book. Re-scoring is defensible when the rule was misapplied; starting fresh is defensible when the rule itself changed. What is never defensible is a table where half the matches used one rule and nobody wrote down which.

saying these in an interview costs you the question

  • Restates history without checking the past inputs still reproduce it
  • Forks the series and never labels the break in the data
  • Assumes nobody compares periods across the break
  • Deletes the superseded figures instead of keeping them readable
  • Treats a redefinition and a defect as the same kind of change
  • Calls it a deployment decision rather than a decision about the numbers