skip to content

A rerun over stored records and a continuous job disagree on the same hour's total: what must the reconciliation step decide?

level: seniorimportance: must knowfreq 62%

answer

  1. two implementations of one rule
  2. ask which answer a reader sees
  3. precedence and a cutover moment
  4. corrections must be visible, not silent
  5. detect the gap before a stakeholder does

basics

~20 s

Reconciliation decides which of the two answers a reader sees for each period, at what moment one supersedes the other, whether a published number is allowed to change, and how a disagreement is detected rather than reported by a stakeholder.

solid answer

~50 s

The setup is a **two-path design**: the same business rule written twice — once as a job that recomputes over stored records and once as a job that never stops — plus a step that decides which of the two answers a reader sees. The two drift for several reasons at once: two code paths maintained by hand, differently cut inputs, records that showed up after the live path had already emitted, and reference data read at different moments. Reconciliation is not a bug fix; it is the permanent job of assigning precedence per period, fixing the moment the recomputed answer supersedes the live one, deciding whether a reader may see a number change and how that is signalled, and running a comparison that catches a gap before someone downstream does. Automatic precedence for the rerun is a policy, not a fact — a partial or broken rerun can overwrite a correct answer.

go deeper

for a junior

Recall that running the same rule twice — once over stored records, once continuously — produces two answers, and that something has to decide which one a reader is shown.

for a middle

Explain the causes of divergence in mechanism terms: two hand-maintained code paths, differently cut inputs, late-arriving records and reference data read at different moments.

for a senior

Show that reconciliation is owned work: precedence per period, a stated cutover rule, a visible correction path, and a scheduled comparison that alerts before a stakeholder notices.

for a principal

Frame it as a standing cost the organisation signed up for — two implementations, two on-call surfaces, twice the compute — and be able to say what would justify continuing to pay it.

## The shape being described A **two-path design** is the same business logic written twice: one implementation recomputes over records already stored, the other runs continuously over arrivals, and a third piece decides which of their two answers a reader sees. It is adopted for an honest reason — the live path is fresh but provisional, the recomputation is slower but works over a settled input — and the price is that the organisation now maintains two answers to every question. ## Why the two answers differ - **Logic drift.** Two hand-maintained implementations diverge the first time a fix is applied to one and not the other, and nothing in either path notices. - **Differently cut inputs.** The recomputation reads a stored body of records; the live path read arrivals. Anything dropped, duplicated or delivered twice on one side only will show up as a difference in the total. - **Records that turned up afterwards.** The live path emitted a number and then more records for that period arrived. What that should mean for the published figure is a semantics question with its own owner; here it is simply a cause of divergence. - **Reference data read at different moments.** The live path joined against the categories as they stood at the time; the recomputation joined against them as they stand now. Both are defensible and they do not agree. - **Different failure histories.** One path had a machine die and redid work; the other did not. Where writes are not idempotent, that alone produces a gap. ## What reconciliation has to decide 1. **Precedence per period.** Which path's number is the published one for a given hour or day, and on what evidence. 2. **The cutover moment.** From when the recomputed answer replaces the live one, stated as a rule rather than as whenever the run happens to land. 3. **Whether a reader may see a value change.** Either published numbers are immutable and the live figure is labelled provisional, or they are revisable and a correction has to be visible to anyone who already acted on the old one. 4. **How disagreement is detected.** A scheduled comparison of the two outputs per period, with a tolerance, an owner and an alert — otherwise the detector is a stakeholder in a meeting. 5. **What happens when the recomputation is the wrong one.** A partial input, a bad deployment or a reference table that moved can all make the slower path the incorrect one, and a policy that publishes it blindly overwrites a correct answer with a wrong one. ## What each arrangement gives a reader | arrangement | what a reader gets | what it costs | |---|---|---| | live path only | fresh, provisional, occasionally wrong at the edges | no second implementation, but no settled figure either | | recomputation only | settled and consistent, hours late | nobody can act on the current period | | both, reconciled | fresh now, settled later | the reconciliation itself, forever | ## The standing price of the two-path design - Every change to the business rule is two changes, two reviews and two test suites — and a third test that the two agree. - Two operational surfaces to run, alert on and be paged for, with different failure units and different capacity profiles. - Roughly twice the compute for one answer, since both paths process the same period. - A number a reader quotes depends on when they looked, so every conversation needs a timestamp attached. - Anyone joining the team must learn both implementations before they can safely change either. ## What varies, and what does not Some engines in this class expose one program surface that can run over a finite stored input or continuously, which removes the hand-maintained duplication. That shrinks the drift; it does not remove it. The two executions still read differently cut inputs, still publish under different contracts — a settled answer once against a figure revised as more arrives — and still read reference data at different moments. Reconciliation survives the shared program. ## What an interviewer is listening for That you treat reconciliation as a permanent product decision with an owner, not as a migration task or a script. The weak answer is that the recomputation is correct by definition and the live path is a preview. The strong answer names precedence, a cutover rule, a visible correction path and a detector — and can say what it would do on the day the slower path is the one that is wrong.

  • What detects the disagreement before a stakeholder does?
    A scheduled comparison of both paths' outputs for the same period, with a stated tolerance, an owner and an alert when the gap exceeds it. Publishing the gap as an ongoing measurement also makes drift visible as a trend rather than as a single bad morning.
  • Is the recomputed answer always the one to publish?
    No. It is usually more complete, but it can be the broken one — a partial input, a bad deployment, or reference data that moved after the fact. Precedence should be a rule supported by checks on the run's own inputs, not an axiom that the slower path wins.
  • Does sharing one program between both paths remove the need for reconciliation?
    It removes the largest cause — two hand-maintained copies drifting — but not the others. The two executions still read differently cut inputs, publish under different contracts, and join against reference data as of different moments, so the answers can still differ.
  • How should a reader be told that a published figure has been revised?
    By making the state explicit in what they read: a provisional marker while the live path owns the period, a settled marker afterwards, and a record of the previous value. Silent replacement is what turns a normal correction into a trust incident.

saying these in an interview costs you the question

  • Says the recomputed answer is authoritative by definition, whatever produced the gap
  • Assumes identical logic in both paths guarantees identical numbers
  • Treats reconciliation as a one-off migration rather than permanent work
  • Blames every difference on records that arrived after the period closed
  • Thinks sharing one program between the paths removes the drift entirely
  • Lets a published number change silently when the slower answer lands