skip to content

Two writes update the same column of a table in sequence, and the second one's condition was computed before the first ran — what goes wrong, and what changes if it is recomputed?

level: seniorimportance: should knowfreq 38%

answer

  1. the condition is a snapshot
  2. the write moves the population
  3. captured before, or recomputed after
  4. the classic promotion cascade
  5. order-dependent reclassification

basics

~20 s

A condition captured before a write is blind to the rows that write changed, so the second write addresses the old population. Recomputed, it does the opposite: it can cascade over rows the first write just created.

solid answer

~40 s

A condition column — one outcome per row, computed for the whole table — is a **snapshot of the rows as they stood when it was computed**, not a standing rule that tracks the column underneath. Capture it before a write and it still names the pre-write population, so the second write misses every row the first write reclassified. Recompute it after, and it now includes those rows, so the second write cascades over the first write's own output. Neither behaviour is wrong in itself; they answer different questions, and the defect is failing to choose. The repair is to compute every target population from the same pre-write state, or to write the outcome into a fresh column so nothing the writes do can move the ground under the conditions.

code

pseudocode · 12 lines
pseudocode
# the condition is computed ONCE, against the column as it stands now
already_silver <- rows_where(table, tier equals "silver")

write_cells(table, rows: points >= 100, column: "tier", value: "silver")

# `already_silver` still names the rows that were silver BEFORE the line above,
# so this write misses every row that line just reclassified
write_cells(table, rows: already_silver, column: "bonus", value: 50)

# re-evaluating instead gives the opposite behaviour: rows the previous write
# created are now included, and the bonus cascades to them too
write_cells(table, rows: tier equals "silver", column: "bonus", value: 50)

go deeper

for a junior

Remember that a condition is worked out once, over the rows as they are at that moment. It does not keep itself up to date when a later write changes those rows.

for a middle

Explain both directions: a captured condition names the pre-write population, and a re-evaluated one includes rows the earlier write created. Show why the two give different answers.

for a senior

Diagnose it from the symptom — a band with more members than anyone intended, or rows that were silently skipped — and fix it by fixing the populations up front or by writing into a separate column.

for a principal

Decide how the team expresses multi-step reclassification at all, and whether order-independence should be structural rather than left to each author's care.

## A condition is a snapshot, not a rule A condition column holds one outcome per row, computed for the whole table at a moment in time. Nothing about it is live. Once it exists it is data: a fixed set of outcomes that will keep describing the rows as they were, however much the underlying column changes afterwards. This is easy to say and easy to forget, because the code reads like a rule. A line that writes into *the rows where the tier is silver* sounds like a standing policy. It is not. It is a one-off addressing of whichever rows satisfied that test at the instant the test ran. ## Two failure shapes, opposite to each other | how the condition was obtained | what it names when the write runs | characteristic bug | |---|---|---| | computed into a name before an earlier write | the population as it stood **before** that write | the later write misses every row the earlier one reclassified | | re-evaluated after the earlier write | the population **including** rows the earlier write created | the later write cascades over the earlier write's own output | The cascade is the more famous of the two because it is the classic tier-promotion bug: a first write moves a band of rows up one level, and a second write, testing for that level, then moves some of them up again — rows that were never supposed to reach the top band arrive there because they passed through the middle one on the way. The stale-capture shape is quieter. Nothing cascades; a set of rows just fails to be updated, and they look exactly like rows that legitimately did not qualify. ## Why it is so hard to see - **Nothing is raised.** Every write is well-formed and completes. - **No rows appear or disappear.** Only the values are wrong, and only for a subset. - **Both orderings of the writes run without complaint**, so the sequence looks commutative when it is not. - **A small test set often hides it**, because the bug needs rows that fall into both populations, and a handful of sample rows may contain none. - **The wrong rows are plausible.** A row in the top band is not obviously an error; it becomes one only when someone counts the bands and the total looks off. ## Making the intent explicit Three approaches, in rough order of preference: 1. **Compute every target population up front, from the same pre-write state, and keep them fixed.** Every write then addresses the world as it was before the sequence began, and the result no longer depends on which write ran first. This is the right default when the writes are meant to be a single reclassification expressed in parts. 2. **Write the outcome into a new column.** The conditions read the original column, which nothing modifies, and the destination is separate. Order-independence is then structural rather than a discipline. 3. **Order the writes so the populations provably cannot overlap**, and say so in the code. This works, but it is the fragile option: it survives only until someone adds a fourth band. Whichever you pick, name it. A comment saying *these populations are as of before the sequence* costs nothing and is the difference between a reader trusting the code and a reader re-deriving it. ## Proving it afterwards The check is the same one that catches every silent write defect: re-read the affected rows from the table and compare their values against what the sequence was supposed to produce. For a reclassification, the useful form is to read back the rows in each outcome band and confirm that the rows in the top band are the ones that were meant to be there — not merely that the write ran, and not merely that some rows changed. A sequence of writes that individually succeeded can still produce a classification nobody specified, and only looking at the resulting values will tell you.

  • How do you make a sequence of reclassifying writes order-independent?
    Compute every target population from the same pre-write state before any write runs, and hold those populations fixed while applying them. Alternatively write the outcome into a new column so the conditions read something nothing modifies. Either turns a sequence whose result depends on ordering into one that does not.
  • Which of the two behaviours — captured or re-evaluated — is the correct one?
    Neither in general; they answer different questions. A captured condition means the rows that satisfied this before we started; a re-evaluated one means the rows that satisfy it now, including ones we just created. The defect is not picking. Write the one you mean and make the choice visible.
  • Why does a small test set so often miss this?
    The bug needs rows that fall into both populations — rows the first write moves into the set the second write addresses. A sample of a few dozen rows frequently contains none, so every write appears to do exactly what it says. Construct a case with a row in the overlap deliberately.

saying these in an interview costs you the question

  • Treats a sequence of writes as one atomic reclassification.
  • Assumes a stored condition tracks the column it was computed from.
  • Believes re-evaluating the condition each time is always the safe choice.
  • Thinks the write raises if the addressed rows have already changed.
  • Reorders the writes without stating which population each one means.
  • Concludes the sequence is fine because every write ran without error.