Forty long-lived replicas of a chat gateway serve two different versions of an edited rate-limit table — how do you detect and bound that split?
answer
- every propagation path is per-instance
- a stale replica is not a failing replica
- report what was adopted, not what was delivered
- distinct-version count as one alertable series
- prefer edits both versions serve correctly
basics
~20 sMake each replica publish a checksum of the value set it actually adopted and the time it adopted it, then alert when more than one version is live beyond the expected window. Bound the split by designing the edit so both versions are simultaneously acceptable.
solid answer
~60 sPropagation is per-instance and asynchronous — each host refreshes its projection on its own clock, each process re-reads on its own trigger, and a replacement walks the fleet instance by instance — so a split fleet is the normal transient state of every configuration change, not an anomaly. The problem is that it is invisible by default: a replica on the old table answers plausibly and logs nothing unusual. Make it visible by having every replica report the identity of the values it is serving, a checksum of the bytes it adopted plus the adoption time, on its status surface, in a log field and as a metric labelled by version; then the distinct-version count across the fleet is one number you can alert on when it stays above one for longer than the expected window. Bounding it is a design choice: prefer edits that both versions can serve correctly at once, and where that is impossible, replace the fleet rather than trusting reload and accept the rollout as the cost.
code
pseudocode · 14 lines# in each replica, at the moment a candidate is adopted
limits = candidate
configVersion = checksum(candidate.rawBytes) # adopted bytes, not disk bytes
adoptedAt = now()
exportMetric("config_version_info", labels = {version: configVersion})
# fleet check, run from outside
seen = emptyMap
for each replica in workload:
v = fetch(replica, "/status").configVersion
seen[v] = seen[v] + 1
if count(seen) > 1 and elapsedSinceEdit > expectedWindow:
alert("fleet split across " + count(seen) + " value versions", seen)go deeper
Know that replicas pick up a changed value independently, so for a while after any edit some of them answer using the old value and some using the new one.
Explain where the window comes from — per-host projection refresh, per-process detection, instance-by-instance replacement — and why a stale replica looks perfectly healthy from the outside.
Design the visibility: each replica reports the checksum and time of the values it adopted, exported as a labelled metric, with an alert on the distinct-version count staying above one beyond the expected window.
Own the rule that configuration changes must be safe under a mixed fleet, and that incompatible changes are split into compatible steps rather than coordinated. Simultaneity across a fleet is a cost most edits should never incur.
## A split fleet is the normal state, not the exception Every mechanism that propagates a value to a running fleet is **per-instance and asynchronous**: - the projection onto each host refreshes on that host's own clock; - each process detects the change on its own watch or receives its own reload signal; - if the change is carried by replacing instances, the replacement walks the fleet a few at a time; - an instance that was unreachable, restarting or newly created at the wrong moment starts from whichever bytes were current for it. So for some window after every edit, part of the fleet enforces the old table and part enforces the new one. The engineering question is never "how do I avoid that?" — you cannot, short of stopping the whole fleet — but "how long is the window, how do I see it, and is the change safe inside it?" A worked bound for the gateway, stating the assumptions: if the change is carried by replacement, a few instances at a time, and each new instance needs roughly twenty seconds to start and be ready, then forty replicas take on the order of several minutes end to end, and **both tables are live for that entire period**. If the change is carried by a re-read, the window is instead the worst-case projection refresh plus the detection delay on the slowest host — usually shorter, but with no natural completion event to observe. ## Why it is invisible by default A replica on the stale table is not failing. It answers every request, returns correct-looking results, passes its health checks and logs nothing unusual. The only symptom is at the edges: - two clients doing the same thing get different answers, depending which replica they landed on; - a limit that was lowered is still not enforced for a fraction of traffic, and the fraction shrinks over minutes; - reverting the edit "fixes" it for some users and not others, which sends the investigation in exactly the wrong direction. That is the real cost of a split fleet: not that it exists, but that nobody can tell it exists. ## Make the version a first-class fact The fix is for each replica to publish **what it is actually serving**, not what was delivered to it: 1. On every successful adoption, compute a short checksum over the raw bytes the process parsed. 2. Store it with the adoption timestamp, and expose it on the instance's status surface. 3. Emit it as a structured field on a log line at adoption, so the history is queryable. 4. Export it as a metric labelled by version, so the fleet-wide count of distinct versions is one series. 5. Alert when that count stays above one for longer than the window you expect for your platform. The checksum must be taken from the bytes the process **adopted**, not from the bytes on disk — otherwise a replica whose reload failed reports the new version while serving the old table, which is worse than reporting nothing. | Signal | What it proves | What it misses | |---|---|---| | Bytes present under the mount | Delivery reached this instance | Whether the process ever read them | | A reload log line | The process attempted a re-read | Whether it adopted or rejected the result | | Adopted-version checksum and time | Which table this replica is serving now | Nothing, if it is emitted only on change and the process later froze | | Distinct-version count across the fleet | Whether a split exists right now | Which replicas, unless it is labelled per instance | Where the workload cannot be changed to report anything, the fallback is behavioural: send probe requests that are shaped to produce a different answer under each version, one per replica, and infer the version from the response. ## Bounding the split by design Detection tells you the split exists; design decides whether it matters. - **Prefer both-tolerant edits.** Lowering a numeric limit, adding a new entry, widening an allowance — both versions serve correctly and the window is merely a period of mixed strictness. - **Split incompatible edits into compatible steps.** To rename a key, first ship readers that accept both names, then change the value set, then remove the old name. Each individual step is safe with a mixed fleet. - **Where the change must be uniform, buy it with a replacement** rather than trusting each process to re-read. A replacement has a definite completion you can observe; a fleet-wide reload has no natural end. - **Do not try to synchronise the moment.** Coordinating forty processes to switch together is a distributed-consensus problem that a configuration edit does not deserve, and its failure mode is worse than the split it prevents. ## What to do while the split is live If a split is discovered during an incident, the first question is which version is correct, and the second is whether waiting closes it. A split created by replacement closes on its own and is observable; a split created by a re-read may be permanent on any replica whose watcher died, and that replica will sit on the old table indefinitely with no error. That is the case where the version report earns its whole cost: without it, the stuck replica is found weeks later by a customer.
- Why report a checksum of the adopted bytes rather than reading the bytes under the mount?Because the mount tells you what was delivered, and the question is what is being enforced. A replica whose reload failed validation, or whose watcher died, has the new bytes on disk and the old table in memory. Reporting the disk content would mark it as up to date and hide exactly the replica you need to find. Report the bytes the process actually parsed and swapped in.
- The split closes on its own after a few minutes. Is it still worth alerting on?Alert on duration, not on existence. A count above one is the expected state during any change, so the useful alert is that it stayed above one for longer than your platform's worst-case window. That threshold catches the failures that never close — a dead watcher, an instance nobody signalled, a stalled replacement — without firing on every routine edit.
saying these in an interview costs you the question
- Treats a mixed fleet as a bug rather than the normal transient state
- Reports the bytes on disk as the version being served
- Assumes replicas adopt a change at the same moment
- Tries to coordinate a simultaneous switch across the fleet
- Renames a configuration key in one step with a mixed fleet live
- Alerts on any version mismatch and drowns in noise during every edit