skip to content

How do you prove across a demand surge that no unit of work was lost or done twice?

level: seniorimportance: should knowfreq 44%

answer

  1. The bad outcomes leave no failures behind
  2. Identifiers, generated by the submitter
  3. Two lines of arithmetic, not one
  4. Totals balance while both faults exist
  5. Refused is a term, not a failure

basics

~20 s

Tag every submitted unit with a unique identifier before demand rises, then reconcile after the observation window: submitted equals completed plus refused plus pending. Separately compare distinct completed identifiers against total completions, because totals alone hide loss and repetition together.

solid answer

~50 s

Both faults are invisible in the obvious figures. A unit accepted and then dropped when a buffer overflowed produces no failure, because the caller was already told it was accepted; a unit processed twice produces two successes that each look correct. So the evidence has to be counts over identifiers, generated on the submitting side and recorded before the call rather than after it. After the post-surge observation window closes, reconcile `submitted = completed + explicitly refused + still pending` to catch loss, and compare distinct completed identifiers against total completions to catch repetition. **You need both lines**: one unit lost and one processed twice leaves the totals exactly balanced. Refusal is a separate term rather than a failure, so a system that says no honestly is not scored the same as one that accepted everything and silently discarded the excess.

code

pseudocode · 15 lines
pseudocode
# submitting side, before the call
id = new_unique_identifier()
record submitted(id, at = now)
send(work, id)

# after the post-surge observation window closes
loss        = submitted_count - (completed_count + refused_count + pending_count)
repetition  = completed_count - distinct(completed_ids)

assert loss == 0        else list missing identifiers
assert repetition == 0  else list identifiers seen more than once

# second pass, on the data itself
for id in distinct(completed_ids):
    assert effects_for(id) == 1

go deeper

for a junior

Be ready to recall that a surge can cause work to be lost or processed twice without producing any visible failure, and that unique identifiers on each unit of work are what make either one detectable afterwards.

for a middle

An interviewer expects the reconciliation stated as arithmetic — submitted against completed, refused and pending — plus the point that distinct identifiers, not totals, are what expose repetition.

for a senior

Show you designed the accounting into the run rather than bolting it on: submitter-generated identifiers recorded before the call, explicit boundaries for in-flight work, reconciliation only after the window closes, and a second check on effects per identifier.

for a principal

Own how much per-unit accounting a product's runs should carry given its cost, and where the line sits between systems that must prove conservation of work and those where a sampled check is proportionate.

## Why counts are the only evidence Across a surge, the failures that matter most are the ones that leave nothing behind. A unit of work that was accepted and then discarded when a buffer overflowed produces no failed response, because the caller was already told it had been accepted. A unit that was processed twice because something re-sent it produces two successes, each of which looks perfectly correct on its own. Neither appears in a response-status summary, in a response time distribution, or in a failure log. What does reveal them is arithmetic across both sides of the event, on identifiers rather than on totals. ## The reconciliation identity Before demand rises, every submitted unit is given a unique identifier that travels with it and is recorded on both sides. Then, once the post-surge observation window has closed: ``` submitted = completed + explicitly refused + still pending total completions = count of distinct completed identifiers ``` The first line catches loss. The second catches repetition. You need both, and this is the part candidates miss: **the totals can balance while both faults are present at once.** One unit lost and one unit processed twice leaves the completed total exactly where it was expected to be, and every rate and percentage derived from it looks healthy. | Symptom | What it means | What proves it | |---|---|---| | submitted exceeds completed plus refused plus pending | work vanished | the missing identifiers, listed | | total completions exceed distinct identifiers | work repeated | identifiers appearing more than once | | both faults present in equal measure | loss and repetition together | totals balance; identifiers do not | ## Designing the accounting into the run - **Generate the identifier on the submitting side**, never inside the system under test. A unit the system never received cannot be shown to be missing if only the receiver assigns names. - **Record the submission before the call, not after it.** A submission recorded only when the call returns successfully can never evidence a loss, because the losses are exactly the calls that did not return. - **Bound the run explicitly.** Everything counted must fall inside the window; work in flight at either boundary is counted separately and either drained or excluded on purpose rather than by accident. - **Define what an explicit refusal looks like before the run**, so the refused term is populated from a real signal rather than inferred. - **Reconcile only after the observation window closes.** Reconciling at the end of the peak reports every still-queued unit as lost and produces a spectacular false failure. ## Refusal is not loss A system that refuses work it cannot take, and says so to the caller, has behaved correctly under a surge even though fewer units completed than were submitted. The reconciliation has to be able to express that outcome, which is why the refused count is a term in the identity rather than a footnote. Collapsing refusals into "failures" produces a run that punishes the correct behaviour and rewards a system that accepted everything and dropped the excess in silence — the exact inversion of what the run is for. ## Repetition and its effects The counts alone do not tell you whether repetition is a defect. Repeating a read has no consequence; repeating a state change may have a very large one. So the accounting is followed by a second question the run must answer: for each identifier that completed more than once, did the repeat produce a second effect — a second stored record, a second charge, a second outbound message — or was it absorbed? That is checked on the far side of the run by counting effects per identifier rather than counting completions. A system that repeats attempts but leaves one effect each has held the property that matters; a system whose totals balance while its data carries two effects for one identifier has failed it while looking entirely healthy from the outside. ## Practical traps - **Sampling.** A run that records only a sample of submissions can prove neither claim. Reconciliation needs every unit, which is a real cost and has to be budgeted for. - **Clock skew.** Aligning the two sides by timestamp goes wrong at exactly the moment of interest. Align on the identifier instead. - **Aggregated counters.** Counters that reset, roll over, or are collected on an interval will lose a burst's worth of events. Count from the itemised record, not from a periodic figure. - **Identifier rewriting.** If an intermediate layer re-sends a unit under a fresh identifier, the reconciliation reads it as one loss plus one unrelated extra completion. The identifier has to survive whatever re-sends the work. - **Deduplicating stores.** A store that discards a repeat on write hides the repetition from the effect check while the completion counts still show it, so check both sides rather than one.

  • Why must the identifier be generated by the submitting side rather than assigned when the work arrives?
    Because the losses are exactly the units that never arrived. If names are assigned on receipt, a unit that vanished in transit has no name and cannot appear as missing in any reconciliation — it simply never existed as far as the receiving side is concerned. Submitter-generated identifiers, recorded before the call rather than on success, are what make an absence provable.
  • The identifier counts show repetition, but the stored data shows one effect per identifier. What do you report?
    Report both: the system repeated attempts under surge, and the property that matters — one effect per unit of work — held anyway. That is a materially different finding from data carrying two effects for one identifier, and it changes who acts on it. The repetition still costs capacity during recovery, so it belongs in the result even though nothing was corrupted.

Counting parcels into and out of a depot balances perfectly even if one was lost and another delivered twice. Only the tracking numbers can tell those two situations apart.

saying these in an interview costs you the question

  • Concludes work was intact because submitted and completed totals matched
  • Records a submission only after the call returns successfully
  • Counts still-queued work as lost by reconciling before the window closes
  • Treats explicit refusals as identical to silently dropped work
  • Reconciles a sample of submissions rather than every unit
  • Aligns the two sides by timestamp instead of by identifier