A job written for a live feed is replayed over six months of stored input - which of its parts depend on the wall clock?
answer
- the code reads the present
- machine clock, not record timestamp
- timeouts measured in real seconds
- six months arrive in one hour
- pass the present in as a parameter
basics
~20 sAnything the code reads from the present: the current time written into output, timeouts and inactivity gaps measured in real seconds, rate thresholds per minute, expiry of held state, and relative date filters. Compressed into one hour, none behaves as it did live.
solid answer
~50 sThe replay changes one thing the logic never expected: six months of records now pass in an hour, while the machine's clock still says today. Every decision the code makes from that clock is therefore wrong. A generated load-date column stamps today onto a record from March. A grouping that closes after thirty idle minutes of machine time never closes during the run, then closes everything at once when the job finally idles. A `per minute` rate threshold sees the whole history as one continuous burst. State that expires after a real day expires nothing. Worst of all, a filter written as `the last thirty days` against the current date selects nothing and the run reports success. The repair is to derive every one of these from the timestamp carried on the record, and to pass the present in as a parameter of the run rather than reading it.
go deeper
Know that a replay does not replay the original timing: the machine's clock still says today and the input arrives thousands of times faster. Be able to point at a generated load-date column as the obvious casualty.
Explain the difference between a decision made from the machine's clock and one made from the timestamp carried on the record, then walk through what an inactivity gap of thirty real minutes does when the whole history arrives in seconds.
Audit a job for wall-clock dependence before committing to a replay, thread a single as-of parameter through instead of reading the present, and validate by replaying one day whose live output you still hold.
Decide whether replayability is a standing requirement for every job the team writes - no direct clock reads, the present supplied as a parameter - or whether some jobs can only ever be rerun approximately, which is a statement you owe the consumers of their numbers.
## What a wall-clock dependence is A replay pushes stored input through logic that was written while that input was still arriving. The logic therefore contains decisions that were made about **the present**, and the present has moved. The clock in question is the machine's own clock, read at the moment a worker process - one process on one machine that runs pieces of the job and owns the memory they use - reaches a record. That reading is commonly called **processing time**, as against **event time**, the timestamp carried on the record itself. The second one is not something an input inherently has: a pipeline **assigns** it, from a field in the record, from the moment of arrival, or from metadata the source supplies. A pipeline that never assigned one has been grouping by arrival all along, and a replay makes that visible within minutes. Compression is the whole problem. Six months that originally took six months to arrive now pass in an hour - roughly four thousand times faster than the logic has ever seen. ## Where the dependence usually hides - **Stamping.** A generated column holding the current date or timestamp: a load time, a partition column, an audit field. Every replayed row claims today. - **Timers and inactivity gaps.** A grouping closed after thirty idle minutes on the machine's clock; a match that gives up after two minutes of waiting; a periodic flush every ten seconds. - **Rate and threshold logic.** *More than one hundred events for this key in a minute*, computed over elapsed real time rather than over the records' own timestamps. - **Expiry.** Held state dropped once it has been retained for a real day; a cached lookup refreshed hourly. - **Relative filters.** A predicate written as *the last thirty days*, evaluated against the current date, which in a replay of last spring matches nothing at all. ## What each one does under compression | Depends on | While live | Replaying six months in one hour | |---|---|---| | current time written into output | correct by coincidence | every row dated today | | inactivity gap of thirty machine-clock minutes | closes groups at real gaps | closes nothing during the run, then closes everything at once when the job idles | | per-minute rate threshold | fires on real bursts | fires continuously; the whole history looks like one burst | | state expiry after a real day | bounded | nothing expires, so the retained set grows toward the whole replay | | *last thirty days* filter | selects recent data | selects nothing, and the run succeeds empty | The last row is the dangerous one, because the job does not fail. It produces a clean, empty, plausible result. ## What varies between engines How much help the runtime gives you differs sharply across this family, and a candidate who has used only one of them will generalise wrongly: - Where the runtime processes each record as it arrives and folds it into state held under its key, a timer can usually be registered against **either** clock, so the dependence may be a one-word choice in the code - and equally a one-word mistake. - Where a continuous job is built from repeated small finite runs, the slice boundary is itself a machine-clock artefact: an aggregate emitted *every thirty seconds* emits once per slice no matter whether that slice covered thirty seconds or three weeks of history. - In the two-phase disk-to-disk model, where each phase writes its whole output to shared storage before the next reads it, there are no runtime timers at all. Every wall-clock dependence is in code someone wrote, which makes it easier to find and gives you no runtime affordance to fix it with. ## Making the logic replay-safe 1. Give the run a single **as-of** parameter - the moment the run is pretending to be - and thread it everywhere the code would otherwise read the present. In the live job it is simply set to now. 2. Derive grouping, expiry and thresholds from the timestamp carried on the record, not from elapsed real time. 3. Never write a relative date filter into logic that will be replayed; bound the run by an explicit period passed in. 4. Validate by replaying one day whose live output you still have, and diffing. The diff finds the dependences an audit of the code missed. ## What throttling does and does not fix A common instinct is to slow the replay to roughly the original arrival rate. That does repair the dependences measured in elapsed real time: timers, inactivity gaps and per-minute thresholds start behaving as they did. It repairs nothing that reads the current date, and it costs you the original six months of runtime. Throttling is a workaround for logic you cannot change, not a fix for logic you can.
- Does throttling the replay to roughly the original arrival rate fix these dependences?Only the ones measured in elapsed real time: timers, inactivity gaps and per-minute thresholds start behaving as they did live. Anything reading the current date still says today, and expiry keyed to real duration still drops nothing. You also pay the original runtime back in full, so throttling is a workaround for logic you cannot change rather than a repair.
- How would you find every wall-clock dependence in a job before replaying it?Search for reads of the current time, timers registered against the machine's clock, and retention or caching keyed to elapsed real duration, and make each take its value from one run parameter instead. Then replay a single day for which the live output still exists and diff the two outputs. The diff catches the dependences the search missed.
saying these in an interview costs you the question
- Replay is safe because the code was not changed
- Assumes an inactivity gap measured in real minutes still closes groups correctly
- Thinks stamping today's date on a record from March is cosmetic
- Believes slowing the replay to the original rate makes the logic correct
- Expects the job to fail loudly when a relative date filter matches nothing