How do you verify a one-off data backfill will finish inside its window before running it in production?
answer
- Do not divide a total and multiply
- The rate falls as the run proceeds
- Project from the slowest chunks observed
- Rehearsal cannot reproduce live contention
- Checkpoint thresholds and a practised abort
basics
~20 sRehearse it against a volume-representative copy, measure the rate chunk by chunk instead of only the total, and extrapolate with the non-linear effects stated. Then set a go/no-go checkpoint during the real run and practise the abort.
solid answer
~50 sDo not extrapolate from a total. Run the backfill against a copy sized and shaped like production, instrument the elapsed time and rows processed per chunk, and look at the trend across the run: rate usually falls as the target grows, as caches stop covering the working set, and as the run reaches denser or skewed regions. Project from the slowest observed rate, not the mean, and add margin for the effects a rehearsal cannot reproduce, chiefly contention with live traffic. Then treat the window itself as an assertion: define a checkpoint — by this fraction of the window, this fraction of rows must be done — decide in advance what happens when the checkpoint fails, and rehearse that abort so it is a practised action rather than an improvisation at three in the morning.
code
pseudocode · 9 lineschunks = rehearsal_log.chunks # id, rows, elapsed, rows_per_second
slowest_decile_rate = percentile([c.rows_per_second for c in chunks], 10)
remaining = 12_400_000 - rehearsal_log.rows_done
pessimistic_seconds = remaining / slowest_decile_rate
with_contention_haircut = pessimistic_seconds / 0.75 # stated estimate
assert with_contention_haircut < window_seconds * 0.65
checkpoints = { 0.25: 0.30, 0.50: 0.60 } # window fraction -> rows fractiongo deeper
Be ready to say that dividing a small sample's time and multiplying up is unreliable, because a long data job usually gets slower as it proceeds rather than holding a steady rate.
Explain what to instrument — rows and elapsed per chunk, not just the total — and name two reasons the rate falls, such as the target growing as you write into it and the working set leaving cache.
Show the decision discipline: project from the slowest observed rate with a stated contention allowance, require the projection to fit well inside the window, and define checkpoint thresholds with a branch for each outcome.
Own the risk position. Decide the margin the organisation accepts, whether the change is split across several windows, who is authorised to abort mid-run, and how that decision is made without the migration's author in the room.
## Why the naive projection is wrong The common approach is: process a sample, divide, multiply. Take a parcel-tracking gateway that must populate a derived field across 12.4 million rows inside a 3-hour-40-minute window. The rehearsal did 1,600,000 rows in 22 minutes, which extrapolates to roughly 2 hours 50 minutes and looks comfortable with 50 minutes to spare. It finished in 3 hours 55 minutes and blew the window. The reason is that the rate is not constant. Per-chunk instrumentation showed it starting at about 1,210 rows per second and drifting to about 740 by the final chunks. Extrapolating from the average of the first few minutes is extrapolating from the fastest part of the run. ## Where the non-linearity comes from Name these in an interview; they are the substance of the answer. - **The target grows as you write into it.** Every maintained secondary structure gets larger and more expensive to update as the run proceeds, so late chunks cost more than early ones. - **Cache falloff.** The first chunks often touch recently written, still-warm data. As the run sweeps into colder regions, more of every chunk comes from storage. - **Uneven regions.** If chunks are cut by an identifier range, the rows are not evenly distributed across ranges. A skewed region processes far more rows per chunk than an average one, and if it sits late in the ordering it lands exactly where you have least slack. - **Growing companion structures.** Journals, change logs and replication streams accumulate during the run and their own maintenance competes with it. - **Live traffic the rehearsal did not have.** A quiet copy has no concurrent readers or writers. Production does, and contention is the effect a rehearsal reproduces worst. - **Feedback limits.** Many runs are deliberately throttled by an observed signal such as replication lag; when that signal degrades, the run slows itself down, which is correct behaviour and ruins a linear projection. ## Instrument the rehearsal properly Record per chunk, not per run: chunk identifier, rows processed, elapsed, rows per second, and any lag or queue signal you throttle on. Then read it as a series: 1. **Plot rate against progress.** A flat line means projection is safe. A downward trend means project from the tail, not the head. 2. **Take the slowest decile of chunks and assume the remainder runs at that rate.** This is a deliberately pessimistic projection and it is the number to plan against. 3. **Check the region distribution.** Confirm the rehearsal actually covered the dense, skewed part of the key space. A rehearsal that only processed the first tenth of the range may have processed the easy tenth. 4. **State the contention allowance explicitly** as a percentage haircut, and say it is an estimate rather than a measurement, because it is. A reasonable planning rule is to require the pessimistic projection to finish inside about 60 to 70 percent of the window. If it does not, the answer is not to hope; it is to change something — narrow the scope of what gets backfilled, split it across several windows, or increase the resources available to the run. ## Turn the window into a runtime assertion The rehearsal gives a prediction; the real run needs a decision rule, and this is what separates a senior answer from a plausible one. Before starting, write down: **at 25 percent of the window, at least 30 percent of rows must be complete; at 50 percent of the window, at least 60 percent.** Those thresholds come from the pessimistic projection with its margin. Then define the branch: if a checkpoint is missed, do you continue and accept overrunning, pause and resume in the next window, or stop and revert? Whichever you choose, **rehearse it**. Practise stopping the run mid-flight on the copy and confirm three things: that progress is durably recorded so a resume does not redo completed work or skip anything, that stopping leaves the data in a state the application tolerates, and that someone other than the author can execute the stop from the written procedure. On an 11-person team the person on call during the maintenance window is frequently not the person who wrote the migration, which is precisely why the abort has to be a documented, practised action rather than institutional knowledge. ## Report it as a decision, not a duration The output of this exercise is not "the backfill takes about three hours". It is: measured slowest sustained rate; pessimistic projection with margin; the assumptions the rehearsal could not test; the checkpoint thresholds; the abort procedure and who runs it. That is a package a reviewer can challenge on specifics, and challenging it is far cheaper than discovering the shortfall halfway through the window.
- Your rehearsal copy is one quarter of production size. Does that make the exercise useless?No, but it changes what you can claim. A quarter-size copy still shows the rate trend, the shape of the per-chunk curve and whether the skewed regions are expensive, which are the qualitative findings. What it cannot give you is a trustworthy absolute duration, because the effects that bend the curve get worse with size rather than scaling with it. Report the trend as measured and the duration as a lower bound, and say plainly that the real run will be slower.
- What single number from the rehearsal would you put in front of a reviewer?The slowest sustained rate observed, with the projection built from it. A mean rate invites the reader to believe the optimistic case, whereas the slowest sustained rate is the one the end of the run will actually deliver. Pair it with the window fraction it consumes, so the reviewer sees margin rather than a bare duration, and state the contention allowance separately so it can be argued with.
saying these in an interview costs you the question
- Extrapolates the whole run from the first few minutes
- Reports a mean rate when the rate is clearly falling
- Assumes a quiet copy predicts production contention
- Starts with no checkpoint and no defined abort
- Never confirms the rehearsal touched the dense regions
- Treats an untested resume as a safe fallback