skip to content

Rerunning Against History

Pushing months of stored input through logic written for live data: the assumptions that only hold in real time, the flood at the destination, and running the correction alongside.

on this pageshow

questions

5

A job written for a live feed is replayed over six months of stored input - which of its parts depend on the wall clock?

level: juniorimportance: must knowfreq 58%

answer

  1. the code reads the present
  2. machine clock, not record timestamp
  3. timeouts measured in real seconds
  4. six months arrive in one hour
  5. pass the present in as a parameter

basics

~20 s

Anything the code reads from the present: the current time written into output, timeouts and inactivity gaps measured in real seconds, rate thresholds per minute, expiry of held state, and relative date filters. Compressed into one hour, none behaves as it did live.

solid answer

~50 s

The replay changes one thing the logic never expected: six months of records now pass in an hour, while the machine's clock still says today. Every decision the code makes from that clock is therefore wrong. A generated load-date column stamps today onto a record from March. A grouping that closes after thirty idle minutes of machine time never closes during the run, then closes everything at once when the job finally idles. A `per minute` rate threshold sees the whole history as one continuous burst. State that expires after a real day expires nothing. Worst of all, a filter written as `the last thirty days` against the current date selects nothing and the run reports success. The repair is to derive every one of these from the timestamp carried on the record, and to pass the present in as a parameter of the run rather than reading it.

go deeper

for a junior

Know that a replay does not replay the original timing: the machine's clock still says today and the input arrives thousands of times faster. Be able to point at a generated load-date column as the obvious casualty.

for a middle

Explain the difference between a decision made from the machine's clock and one made from the timestamp carried on the record, then walk through what an inactivity gap of thirty real minutes does when the whole history arrives in seconds.

for a senior

Audit a job for wall-clock dependence before committing to a replay, thread a single as-of parameter through instead of reading the present, and validate by replaying one day whose live output you still hold.

for a principal

Decide whether replayability is a standing requirement for every job the team writes - no direct clock reads, the present supplied as a parameter - or whether some jobs can only ever be rerun approximately, which is a statement you owe the consumers of their numbers.

## What a wall-clock dependence is A replay pushes stored input through logic that was written while that input was still arriving. The logic therefore contains decisions that were made about **the present**, and the present has moved. The clock in question is the machine's own clock, read at the moment a worker process - one process on one machine that runs pieces of the job and owns the memory they use - reaches a record. That reading is commonly called **processing time**, as against **event time**, the timestamp carried on the record itself. The second one is not something an input inherently has: a pipeline **assigns** it, from a field in the record, from the moment of arrival, or from metadata the source supplies. A pipeline that never assigned one has been grouping by arrival all along, and a replay makes that visible within minutes. Compression is the whole problem. Six months that originally took six months to arrive now pass in an hour - roughly four thousand times faster than the logic has ever seen. ## Where the dependence usually hides - **Stamping.** A generated column holding the current date or timestamp: a load time, a partition column, an audit field. Every replayed row claims today. - **Timers and inactivity gaps.** A grouping closed after thirty idle minutes on the machine's clock; a match that gives up after two minutes of waiting; a periodic flush every ten seconds. - **Rate and threshold logic.** *More than one hundred events for this key in a minute*, computed over elapsed real time rather than over the records' own timestamps. - **Expiry.** Held state dropped once it has been retained for a real day; a cached lookup refreshed hourly. - **Relative filters.** A predicate written as *the last thirty days*, evaluated against the current date, which in a replay of last spring matches nothing at all. ## What each one does under compression | Depends on | While live | Replaying six months in one hour | |---|---|---| | current time written into output | correct by coincidence | every row dated today | | inactivity gap of thirty machine-clock minutes | closes groups at real gaps | closes nothing during the run, then closes everything at once when the job idles | | per-minute rate threshold | fires on real bursts | fires continuously; the whole history looks like one burst | | state expiry after a real day | bounded | nothing expires, so the retained set grows toward the whole replay | | *last thirty days* filter | selects recent data | selects nothing, and the run succeeds empty | The last row is the dangerous one, because the job does not fail. It produces a clean, empty, plausible result. ## What varies between engines How much help the runtime gives you differs sharply across this family, and a candidate who has used only one of them will generalise wrongly: - Where the runtime processes each record as it arrives and folds it into state held under its key, a timer can usually be registered against **either** clock, so the dependence may be a one-word choice in the code - and equally a one-word mistake. - Where a continuous job is built from repeated small finite runs, the slice boundary is itself a machine-clock artefact: an aggregate emitted *every thirty seconds* emits once per slice no matter whether that slice covered thirty seconds or three weeks of history. - In the two-phase disk-to-disk model, where each phase writes its whole output to shared storage before the next reads it, there are no runtime timers at all. Every wall-clock dependence is in code someone wrote, which makes it easier to find and gives you no runtime affordance to fix it with. ## Making the logic replay-safe 1. Give the run a single **as-of** parameter - the moment the run is pretending to be - and thread it everywhere the code would otherwise read the present. In the live job it is simply set to now. 2. Derive grouping, expiry and thresholds from the timestamp carried on the record, not from elapsed real time. 3. Never write a relative date filter into logic that will be replayed; bound the run by an explicit period passed in. 4. Validate by replaying one day whose live output you still have, and diffing. The diff finds the dependences an audit of the code missed. ## What throttling does and does not fix A common instinct is to slow the replay to roughly the original arrival rate. That does repair the dependences measured in elapsed real time: timers, inactivity gaps and per-minute thresholds start behaving as they did. It repairs nothing that reads the current date, and it costs you the original six months of runtime. Throttling is a workaround for logic you cannot change, not a fix for logic you can.

  • Does throttling the replay to roughly the original arrival rate fix these dependences?
    Only the ones measured in elapsed real time: timers, inactivity gaps and per-minute thresholds start behaving as they did live. Anything reading the current date still says today, and expiry keyed to real duration still drops nothing. You also pay the original runtime back in full, so throttling is a workaround for logic you cannot change rather than a repair.
  • How would you find every wall-clock dependence in a job before replaying it?
    Search for reads of the current time, timers registered against the machine's clock, and retention or caching keyed to elapsed real duration, and make each take its value from one run parameter instead. Then replay a single day for which the live output still exists and diff the two outputs. The diff catches the dependences the search missed.

saying these in an interview costs you the question

  • Replay is safe because the code was not changed
  • Assumes an inactivity gap measured in real minutes still closes groups correctly
  • Thinks stamping today's date on a record from March is cosmetic
  • Believes slowing the replay to the original rate makes the logic correct
  • Expects the job to fail loudly when a relative date filter matches nothing
open as a page

A year of history must be pushed through a job whose output lands in a live serving database - what sets the replay's duration?

level: seniorimportance: must knowfreq 52%

basics

~20 s

The write rate the destination can sustain while still serving its live traffic, not the cluster. Beyond that point extra workers only produce rejected writes, retries and contention. Plan the replay backwards from the destination's spare capacity.

open as a page

A replay enriches each stored record with a customer tier read from a table that is overwritten nightly - what is wrong with the output?

level: middleimportance: should knowfreq 50%

basics

~20 s

Every historical record is labelled with today's tier rather than the one in force at the time, so the replay quietly rewrites history. Reproducing the past needs reference data that retains its own change history, looked up as of each record's timestamp.

open as a page

Six months of stored records are pushed through a job that groups by event time in two hours - how do the results differ from the live run's?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The job's running claim about how far time has advanced is derived from the data, so it races through six months in seconds. Groups close back to back, and the replay can end up either more complete than the live run or less, depending on the order history is read in.

open as a page

How would you prove a corrected job is right over a year of history before any consumer sees its output?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Run it beside the existing output: the same stored input, the same period, written to a side destination nobody reads, then compared against what the live job already produced. The comparison, not the run, is the deliverable.

open as a page