skip to content

Which invariants would you let a submitted program enforce at the volatile tier, and which would you keep at a durable system of record?

level: principalimportance: nice to knowfreq 32%

answer

  1. one unit, one node, nothing else
  2. atomic is not durable
  3. what if it were lost, or doubled
  4. survivable and re-derivable belongs here
  5. degrades as latency, not errors

basics

~20 s

Enforce invariants whose violation is survivable and re-derivable: admission decisions, claims, dedupe marks, operational counts. Keep at a durable record anything that must never double or vanish, since the program gives atomicity and nothing else.

solid answer

~50 s

A submitted program gives exactly one guarantee: a multi-step decision applied as one uninterleaved unit, on one node's keyspace. It does not give durability — that is whatever the tier's posture and replication provide, and stores differ on whether a write is acknowledged before a replica holds it. It does not give rollback, so a program that errors part-way leaves its earlier writes. It does not span entries that are not co-located. And it binds only callers that go through it. So the test is: **if this effect were lost on a failover, or applied twice, what happens?** If the answer is "someone gets one extra attempt" — admission, claims, dedupe marks, live counts — the tier is a fine enforcement point. If it is "money moved twice", the invariant belongs at a durable record.

go deeper

for a junior

Recall the boundary: a program at the store makes a decision happen as one step, but it is not the place where the permanent copy of important data lives.

for a middle

Explain the four limits — no durability of its own, no rollback, one node's keyspace, and only binding on callers that use it — and give an example of an invariant each side of the line.

for a senior

Apply the test to a real workload: what a lost or doubled effect costs, whether every writer really goes through the program, and how contention will show up in production signals.

for a principal

Own the contract. Say which invariants are enforced at the tier by choice, what the durability posture makes that worth, who else may write those entries, what bounds each program, and where authority sits when the tier is unavailable.

## What the program actually guarantees One sentence, and it is worth being precise about, because every architecture decision here follows from it: **a multi-step decision, applied as one unit that no other caller interleaves, over entries on a single node's keyspace.** That is genuinely valuable — it closes a race no client-side sequence can close without a retry loop or a claim — and it is much narrower than it sounds when someone proposes making the tier the coordinator for a system. ## The four things it does not give you - **Durability.** Atomicity and durability are separate properties. Whether the write survives a restart or a failover is whatever the tier's durability posture and replication give, and stores in this class differ — some acknowledge only once a replica holds the write, some acknowledge immediately, some offer the choice per call. - **Rollback.** There is no isolation level, no save point and no undo. A program that errors after its second write leaves those writes behind, and repairing them is the caller's job. - **Scope beyond one node.** Where the keyspace is split, the unit is one node's. Entries that are not co-located cannot be in the same unit, and "make them co-located" is a placement decision with its own costs. - **Exclusivity.** It enforces the invariant against callers that go through it. A migration job, an operations console or a second service writing the same entries directly is not covered by anything at all. ## The test to apply Ask two questions about the effect, not about the mechanism: 1. **If this effect vanished on a failover, what would the business notice?** 2. **If this effect were applied twice, what would the business notice?** When both answers are "a user retries" or "a number is briefly wrong and then re-derived", the tier is an appropriate enforcement point and the program is the right tool. When either answer is "a payment moved twice", "an identity was reused", or "nobody would ever find out", the invariant belongs at a durable record — and the tier may still hold a fast approximation in front of it, with the record as the authority. | Fits a program at this tier | Belongs at the durable record | |---|---| | Admission decisions that may occasionally allow one extra attempt | Financial movement, and anything else that must happen exactly once | | A claim marking a unit of work as being worked on now | Uniqueness that must hold permanently and be provable afterwards | | A short-lived deduplication mark against a retry window | State a customer, auditor or regulator reads as final | | Live operational counts that are re-derived from the source | Anything whose violation is undetectable later, so it can never be repaired | ## What operators see when it degrades A store-side decision under contention has a distinctive failure mode that has to be designed for: **it degrades as latency, not as errors.** Callers queue behind the units, the success rate stays at a hundred percent, and the service's error-rate dashboard shows nothing while every dependant is timing out. So the contract includes what will be visible: - the latency distribution of calls to the tier, at high percentiles rather than the mean, because the stall is in the tail first; - how many callers are waiting, and for how long, rather than only how many failed; - the size of any collection a program walks, because that is the number that silently turns a fast program into a stall; - an alert on the durability posture, since an effect that was atomic and then lost on a failover looks like a bug in the caller, not like a lost write. ## Write the contract down The deliverable is not "we use programs at the tier". It is a short statement, reviewable by people who did not write the program: 1. **Which invariant is enforced where**, naming the store-side ones explicitly and listing them as store-side by choice. 2. **What the tier's durability posture is**, and therefore what losing an acknowledged effect does to each of those invariants. 3. **Who else may write those entries**, and how that is prevented or accepted. 4. **What bounds every program's work**, because the program is shared capacity and a two-second unit is a two-second outage for every other caller of the node. 5. **Where the authority lives when the tier is unavailable** — whether the system fails closed, fails open, or falls through to the durable record. The principal-level answer is not a preference between the tier and the record. It is being explicit that putting a decision in the store buys atomicity and buys nothing else, and then choosing the invariants for which that trade is honest.

  • A team argues the tier is now the source of truth because the program makes the decision atomically. What is wrong with that?
    Atomicity is about interleaving, not survival. The decision is applied as one unit and can still be lost on a failover or a restart, depending on the tier's durability posture; some stores acknowledge before a replica holds the write. Source of truth is a durability and recoverability claim, and nothing about applying a decision as one unit supplies it.
  • The program enforces the invariant, but a nightly job writes the same entries directly. What do you do?
    Either route that job through the same program, or accept in writing that the invariant is advisory. A store-side decision constrains only the callers that go through it — there is no constraint the store itself enforces on other writers. The usual failure is that the second writer was added later by someone who never saw the program.
  • How should the system behave when the tier holding these decisions is unreachable?
    That has to be a stated choice, not an emergent one. Fail closed — refuse the work — when a doubled effect is unacceptable; fail open when refusing is worse than an occasional duplicate; or fall through to the durable record where one can answer, accepting the latency. The worst outcome is each caller's timeout handler deciding it independently.

saying these in an interview costs you the question

  • Treats a store-side program as a substitute for a durable record.
  • Assumes an atomic effect survived because it was acknowledged.
  • Forgets that other writers bypass the program entirely.
  • Alerts only on error rate for a decision that degrades as latency.
  • Proposes the tier as the coordinator across entries on different nodes.
  • Treats the choice as a preference rather than a stated contract.