A nightly sweep compares your driver-hours accounts against each member firm's employment records, so what may it change on its own authority?
answer
- converge, do not just replay
- narrowing is safe, widening is not
- zero rows means everyone left
- quarantine, never purge
- told versus discovered
basics
~20 sGrade it by reversibility and direction: a sweep may narrow access by itself — disable accounts and drop derived grants — but must only report anything that widens it or destroys data, and must refuse to act at all when the employment answer looks empty or stale.
solid answer
~60 sThe sweep exists because a per-change provisioning feed drops, duplicates and reorders, so a periodic full comparison is the only thing that converges. The design question is not whether to run it but how much authority it holds. Split discrepancies by direction: **narrowing** — the firm says a driver has gone but the account is active, or a derived grant is no longer implied — fails safe and can be applied automatically; **widening** — re-enabling an account or restoring a grant — must be a report a human acts on, because a sweep able to grant access on machine authority is a privilege-escalation path that runs every night. Destruction is never the sweep's: an account with no counterpart upstream is quarantined and reported, never deleted, because it is as likely to be a broken feed or an out-of-band provisioning route as a real orphan. And the sweep needs its own kill switch: an employment query returning zero rows is indistinguishable from everyone having left, so guard on absolute and relative change thresholds and on the freshness of the answer, and abort rather than apply.
code
pseudocode · 23 linesrun_reconciliation(firm):
upstream = source_of_truth.employed(firm) # may be empty, stale or truncated
local = accounts.for_firm(firm)
# ---- guard BEFORE any write ----
if not upstream.isFresh():
alert("stale employment answer", firm); return
departures = local.active() - upstream
if departures.size > ABSOLUTE_LIMIT
or departures.size > local.active().size * RELATIVE_LIMIT:
alert("implausible drift, aborting", firm, departures.size); return
# ---- narrowing: applied automatically ----
for person in departures:
accounts.disable(person, discovered = true)
for person in upstream:
accounts.revokeDerived(person, no_longer_implied(person))
# ---- widening and destruction: reported only ----
report("disabled locally, still employed", upstream & local.disabled())
report("no upstream counterpart - quarantined", local.unmatched())
metrics.record(firm, examined = local.size, drift = departures.size)go deeper
Recall that a system fed by messages can miss one, and that a periodic comparison against the original records is how the gap is noticed at all.
Explain why the sweep applies the changes that reduce access but only reports the ones that restore it, and what an account with no upstream counterpart might actually mean.
Show the guards that run before any write — freshness, absolute and relative change thresholds, a dry-run diff — and the told-versus-discovered distinction recorded on every transition.
Set the departure-to-no-access window the whole design must live inside, decide how much authority an unattended job may ever hold, and say who owns the report it produces when nobody is on call for it.
The cooperative's hours system learns about employment changes from each member firm as they happen, and that channel is lossy in every ordinary way: a message is dropped during an outage, delivered twice, or applied out of order so a departure lands before the arrival it follows. So a scheduled full comparison against each firm's employment records runs anyway. That much is unremarkable. The decision worth an interview is **how much authority that sweep holds**, because the same job that can silently fix everything can silently break everything. ## Grade authority by direction and reversibility Sort every class of discrepancy by which way it moves access and how hard it is to undo. | discrepancy | sweep's authority | why | |---|---|---| | firm says the driver has gone, account is active | disable automatically | narrows access, trivially reversible, matches the stated leaver policy | | a derived grant is no longer implied by current employment | remove automatically | narrowing, and the target-set model already owns the removal | | firm lists the driver as employed, account is disabled | report only | widening access on machine authority is an escalation path | | a grant exists upstream that the account does not hold | report only | same direction; the missing grant may be a deliberate local decision | | account exists with no counterpart in any firm's records | quarantine and report | as likely a broken feed or an out-of-band account as a real orphan | | a change arrives for a person no record has ever mentioned | hold and report | creating identities from an unverified source is how strangers get accounts | The asymmetry is the whole answer: **removal fails safe and addition does not**. A wrongly disabled driver phones the depot and is re-enabled in minutes with a human in the loop. A wrongly re-enabled account is an access grant nobody approved, made by a job nobody watched. ## The empty answer is the failure mode that matters A query to a firm's employment system that returns zero rows looks exactly like every driver at that firm having left. A sweep with autonomous deactivation and no guard will faithfully disable an entire firm overnight, and it will look like it worked. Before applying anything the sweep must clear: 1. **Freshness** — the answer is from a live, current export, not a truncated or stale file; 2. **Absolute and relative thresholds** — more than *n* departures, or more than *x* per cent of one firm's population in one run, aborts instead of applying; 3. **A readable diff** — the run can be executed dry, producing the change list a human can scan; 4. **Abort-and-alert as the default failure** — a sweep that cannot confirm its input does nothing, which is the correct behaviour precisely because doing nothing leaves access unchanged rather than deleted. Phrase the rule as a principle: **a job authorised to correct everything is authorised to destroy everything, so its input must be validated before its output is trusted.** ## Told versus discovered Record how each state change reached you. A deactivation you were **told** about carries a time, an actor and a reason from the firm. One you **discovered** carries only "not present in today's answer", which is a weaker fact — it could equally be an export bug. Keeping the distinction in the account's transition history means an incident can ask "did anything ever tell us this person left?", and it lets you treat discovered changes more conservatively than told ones. ## The records that do not fit Three classes deserve a named destination rather than a silent drop: - **Out of order** — a change for a person who does not exist locally, or a departure preceding the arrival. Hold it briefly and re-evaluate rather than discarding it, because the create may be seconds behind. - **Orphans** — a local account with no upstream counterpart. This is evidence of something: a feed that stopped, a person provisioned by hand outside the channel, or a firm that left the cooperative. Quarantine disables access while preserving the row, which is the reversible half of the answer. - **Duplicates** — the same change applied twice. Harmless if every write is idempotent, which is the reason to make them so rather than to detect duplicates. ## Prove the sweep is alive A reconciliation job that reports no drift for months is either healthy or dead, and the two look identical from outside. Emit the population it compared and the number of records it examined on every run, not just the discrepancies; alert on a run that examined implausibly few; and track drift found per run as a signal about the *event channel* rather than about people — a rising count means the per-change feed is degrading, which is what the sweep is really there to tell you. Finally, the leaver window the cooperative commits to is bounded by whichever path is slower, so a departure that only the nightly sweep catches inherits the sweep's period; if the stated window is shorter than that, the event channel is load-bearing and its failures are incidents, not background noise.
- Why not let the sweep re-enable an account when the employment record still lists the driver as working?Because it grants access with no human in the loop, every night, on the strength of one system's answer. The disabled state may also have been set deliberately — a suspension, a security hold — which the employment record knows nothing about. Reporting keeps the discrepancy visible and makes restoring access a decision somebody owns.
- The sweep finds an account with no counterpart in any member firm's records. What do you do with it?Quarantine and report; never purge. The account is evidence of something you do not yet understand — a feed that stopped, a person provisioned outside the channel, or a firm that has left the cooperative. Quarantine removes access while keeping the row and its attribution intact, which is the reversible half, and the report is what turns it into an answer.
- The sweep has reported zero drift for four months. Is that good news?Not by itself, because a healthy sweep and a dead one produce the same report. Emit the population compared and the record count examined on every run and alert when either is implausibly low. Treat drift found as a metric about the per-change feed: a rising count means that channel is degrading, which is exactly what the sweep exists to surface.
- How does the sweep's period interact with the departure-to-no-access window the cooperative commits to?A departure caught only by the sweep inherits the sweep's period plus however long the revocation it fires takes to bite. If the committed window is shorter than that, the per-change feed is load-bearing rather than an optimisation, and its outages are incidents. Measure the window as a distribution across real departures, because the tail is what a reviewer asks about.
saying these in an interview costs you the question
- A reconciliation sweep should correct every discrepancy it finds
- Re-granting access automatically is as safe as removing it automatically
- An account missing upstream is stale data and can be deleted
- An empty answer from the employment system means everybody left
- A sweep reporting no drift proves the pipeline is healthy
- The sweep replaces the per-change feed rather than backstopping it