skip to content

How far back should late-arriving data be allowed to restate a warehouse's published history?

level: principalimportance: should knowfreq 28%

answer

  1. measure the lag before choosing a number
  2. the tail matters, not the average
  3. what has the business already closed?
  4. accuracy versus reproducibility
  5. write it down and apply it everywhere

basics

~20 s

Set an explicit lateness window per subject area, driven by each source's observed arrival lag and by which periods the business has closed. Inside the window, restate; outside it, book corrections forward and never silently change a signed-off number.

solid answer

~50 s

Treat lateness as a measured property of each source, not a global constant. Instrument the gap between event time and load time per feed, look at the tail rather than the average, and set a window that covers the realistic lag — then publish it as part of the mart's contract, so consumers know that last week's numbers may still move and last quarter's will not. Align the boundary with whatever the business already treats as closed: financial periods that have been signed off are frozen by policy, and the warehouse reflects that rather than arguing with it. Inside the window, restate and log what changed. Outside it, apply the correction forward as an adjustment. The rule matters less than its consistency: differing implicit behaviour between marts is what actually confuses people, far more than any particular cutoff.

code

text · 6 lines
text
-- example lateness policy, published with the mart contract
subject area   observed lag (tail)  restatement window   outside the window
orders         ~2 days              7 days               restate + notify
payments       ~6 hours             2 days               restate + notify
gl_postings    ~3 days              until period close   adjusting entry only
web_events     ~15 minutes          1 day                drop, counted + alerted

go deeper

for a junior

Know that late data forces a choice between changing already-published numbers and leaving them alone, and that the choice should be a stated rule rather than a pipeline accident.

for a middle

Explain how to measure arrival lag per source and why the tail of the distribution, not the average, determines a usable restatement window.

for a senior

Demonstrate the operational consequences: stale aggregates and delivered extracts, restatement logging so changed numbers can be explained, and alerting on data arriving outside the window.

for a principal

Own the tradeoff and the governance — reproducibility versus accuracy, alignment with the financial close, per-subject-area windows published as part of the mart contract, and resisting the post-incident drift toward restating everything forever.

## Why a policy is needed at all Without an explicit rule, the answer to "how far back do we restate?" is decided accidentally, by whatever each pipeline's incremental window happens to be. One mart rebuilds seven days, another three, another only today. Consumers cannot reason about any of them, and the same late payment shows up in one report and not another. The value of a lateness policy is less about picking the perfect number and more about there being a number that is written down and applied the same way everywhere. ## Step one: measure, do not guess Every fact feed should record both the event timestamp and the load timestamp, which makes lateness a queryable property rather than folklore. Look at the distribution per source and specifically at its tail — a mean lag of four hours tells you nothing if one percent of rows arrive nine days later, because that one percent is precisely what a restatement window has to cover. Sources cluster into recognisable shapes. Streaming telemetry is late by seconds, with an occasional offline-client tail of days. Batch files from partners are late by a predictable cycle. Financial settlements and reconciliations are late by design, sometimes weeks, and are the ones people actually argue about. Manual adjustments and human-entered corrections have no bound at all. ## Step two: let the business boundary dominate The measured lag proposes a window; the organisation's own close process disposes. If accounting closes a month on the eighth working day of the next month and signs the numbers, those numbers are frozen — restating them silently is not a technical decision available to the platform. In that world the warehouse's lateness window for financial subject areas ends at close, and everything later becomes an adjusting entry in the open period. That is not a compromise; it is the same discipline the ledger itself uses, and mirroring it makes the warehouse reconcilable to the ledger instead of perpetually five thousand off it. Analytical subject areas with no formal close have more freedom, and there the argument is between accuracy and reproducibility. ## The real tradeoff: accuracy versus reproducibility **Restating** maximises accuracy — the warehouse always reflects the best currently-known truth. Its cost is that a report is not reproducible: run it twice and it may differ, with nothing visible to explain why. Every restatement also invalidates derived aggregates, delivered extracts and cached dashboards for the affected period, so the true cost is well beyond the update itself. **Freezing** maximises reproducibility — a period, once closed, answers the same way forever, and anyone can archive a report and trust it. Its cost is that the frozen numbers are knowingly slightly wrong, and the corrections live somewhere else, so anyone comparing the warehouse to the source has to understand the adjustment mechanism. Neither is universally right, and a mature platform runs both: short windows and freezing for regulated or closed subject areas, longer windows and restatement for exploratory analytics. What is never acceptable is choosing implicitly. ## What the published policy should contain For each subject area: the observed lag characteristics, the restatement window, what happens to data arriving outside it, and how a restatement is announced. Consumers of a mart need to know whether last week's number is provisional, because that determines whether they should be building a board pack on it. An illustrative shape: orders restate for seven days with notification; payments for two days; general-ledger postings not at all after close, with corrections booked forward; high-volume web telemetry restates for one day and anything later is dropped with an alert, because the analytical value of a two-week-late page view does not justify rebuilding a quarter. That last case is worth stating plainly in an interview: **dropping genuinely worthless late data is a legitimate policy**, provided the drop is counted, alerted on, and disclosed. It is only unacceptable when it is silent. ## Make restatement observable Whatever the window, every restatement should leave a record: which entity, which period, which run, how many rows moved, and how much the headline measures changed. Two things follow. Consumers get an answer to "why did March change?" without a forensic investigation. And the platform gets a feedback signal — a subject area that restates constantly has an upstream delivery problem that no window setting will fix, and the right response is to fix the feed rather than widen the window. ## Governance, not just configuration The window is owned jointly: data engineering knows the lag, the business owns the close and the tolerance for moving numbers. Set it together, review it when a source changes, and resist the drift toward "restate everything forever" that follows any single embarrassing miss. Widening a window is a cheap-feeling response to one incident that permanently raises the cost and unpredictability of every load afterwards — the sort of decision worth making deliberately rather than in the week after an outage.

  • How do you decide the window rather than guessing it?
    Record event time and load time on every fact and look at the lag distribution per source, focusing on the tail. A four-hour mean is irrelevant if one percent of rows land nine days later, because that tail is what a restatement window has to cover. Then let the business close boundary override the measurement wherever a period is formally signed off.
  • Is it ever acceptable to simply drop data that arrives outside the window?
    Yes, when the analytical value genuinely does not justify the rebuild — a two-week-late page view rarely does. The conditions are that the drop is counted, alerted on, and documented in the mart's contract. Silent dropping is never acceptable, because consumers then cannot tell a quiet loss from a real decline.
  • What should a consumer of the mart be told about the policy?
    Which periods are provisional and which are final, how long the provisional window is, what happens to data arriving after it, and how a restatement is announced. Without that, someone builds a board pack on a number that is still moving, and nobody finds out until the two versions are compared in a meeting.
  • What does a subject area that restates constantly tell you?
    That the upstream feed is broken, not that the window is too narrow. Widening the window hides the symptom and permanently raises the cost and unpredictability of every load. The restatement log is the signal: persistent restatement volume is a case for fixing delivery at the source, or renegotiating the SLA with whoever owns it.

saying these in an interview costs you the question

  • Pick a round number of days with no measurement
  • Restate everything forever to be maximally accurate
  • Change closed-period numbers without notice
  • Set one global window for every subject area
  • Drop late data silently with no count or alert

context