skip to content

At 3am one replica of a five-replica device-telemetry ingester exits and nobody is paged - what replaces it, and what does the replacement not inherit?

level: juniorimportance: must knowfreq 75%

answer

  1. count gap, not an event
  2. replaced, never repaired
  3. same spec, new instance
  4. new address, empty writable layer
  5. alert on rate, not on one

basics

~20 s

The platform's control loop sees four copies where the spec declares five and starts a fifth. The replacement is a new instance built from the same spec: new address, empty writable layer, cold caches. Nothing from the dead copy is recovered.

solid answer

~40 s

A cluster platform keeps comparing the declared replica count against the copies it can observe. When one exits, the observed count drops to four, the declared count is still five, and the loop closes the gap by starting a new copy somewhere with room. It is a **replacement, not a repair**: the new instance inherits the workload spec - image reference, command, configuration, reservations - but not the dead instance's identity. It gets a different address, an empty writable layer, an empty in-process cache and fresh connections. Anything the dead copy held only in memory or only on its own local disk is gone. Nobody is paged because this is the designed path, not an incident; what deserves an alert is the *rate* of replacements, not any single one.

go deeper

for a junior

Recall the one-line shape: the platform compares declared copies against running copies and starts a new one when they differ. Say clearly that the new copy is fresh, not the old one revived.

for a middle

Explain why this is a continuous comparison rather than a reaction to a death notification, and list precisely what a fresh copy does and does not carry over from the one it replaces.

for a senior

Show what the model demands of the application: state kept outside the instance, callers reaching it by name, and a warm-up cost paid on every replacement. Then say what you would alert on, given that the count self-corrects.

for a principal

Frame it as a contract the platform offers and the application must earn. Where a workload cannot be interchangeable, be explicit about what you are buying instead and what it costs to keep the guarantee honest.

## The loop that notices A cluster platform is not waiting for an event called "a replica died". It repeatedly compares two things: the **declared state** written in the workload spec (five copies of this thing) and the **observed state** it reads back from the cluster (four are running). When the two differ, it acts on the difference. The death of one copy is not a notification routed to a human - it is simply the moment the two numbers stop matching. That is what makes the behaviour **level-triggered**: the loop reacts to the current gap, not to a message about a past event. If a notification were dropped, the gap would still be there on the next pass and the next pass would still close it. Nothing about the recovery depends on anyone, or any message, having witnessed the exit. ## Replacement, not repair The fifth copy that appears is not the dead one brought back to life. The platform does not reattach memory, does not replay what the old process was doing, and does not carry forward anything the old process kept to itself. It reads the same spec and starts a fresh instance from it. | The replacement inherits | The replacement does not inherit | |---|---| | The workload spec exactly as declared: image reference, command, environment, reservations, ceilings | The dead instance's network address - it is given a new one | | Whatever durable storage the spec asks the platform to attach | Anything written to the dead container's own writable layer | | The same selector, so callers reaching the workload by name find it | Warm in-process caches, pre-opened connections, in-flight work | | The same declared count it is restoring - five, not six | Logs and scratch files left on the dead instance, unless they were shipped off the host | The practical order of events, from the outside: 1. The copy exits; the observed count falls to four. 2. The loop that owns this workload's count notices the gap on its next pass. 3. A placement decision picks a host with room for the declared reservation - a separate decision with its own rules. 4. The image is fetched if that host does not already have it, then the process starts. 5. The new copy is added to routing once it reports itself able to serve - a separate signal again. Steps 3 and 5 are deliberately somebody else's subject; what self-healing owns is steps 1, 2 and 4 - the gap being noticed and closed without a human. ## What the model assumes about your workload Automatic replacement is only safe because the platform is allowed to treat one copy as **interchangeable** with another. That assumption imposes real requirements on the application: - **No durable state inside the instance.** If it matters after the process exits, it was written to storage the platform can reattach, or sent somewhere else. The container's own writable layer is not that place. - **No caller holds an instance address.** Callers reach the workload through a name that resolves to whatever copies currently exist. A client that cached the old address keeps dialling something that is gone. - **No copy is special.** "The one with the warm cache" or "the one we logged into and patched" is not a concept the loop has. It counts copies that match the selector. - **Start-up cost is paid every time.** A replacement re-fetches the image if the host lacks it, re-establishes connections, and refills whatever it caches. A workload with a five-minute warm-up is degraded for five minutes after every replacement, and the count says it is fine. ## Why silence is the correct response Paging on every replacement trains people to ignore the page. A single replacement at 3am is the platform doing its job; the interesting signal is a **change in rate**. Useful alerts on this material are the ones the count cannot express: replacements per hour crossing a normal band, a workload that has not held its declared count for some minutes, or a replacement that never reaches a serving state. The count returning to five tells you the loop converged. It does not tell you the workload is healthy - only that it has the number of copies it was told to have. ## Where the boundary is A copy that keeps crashing and being started again in the same place is a different mechanism, with its own escalating delay, and it belongs to the lifecycle material. Deliberately emptying a host for maintenance is a planned operation with its own caps. And which host the replacement lands on is a placement question. Self-healing is the narrow claim in the middle: **the declared number of copies is restored, automatically, and nothing else is**.

  • If nobody is paged, how would anyone find out the ingester has been losing a copy every few minutes all night?
    Not from the replica count, which keeps returning to five. The signals that expose it are the replacement rate over time, the age distribution of the running copies (all of them minutes old is the tell), and the fraction of time the workload spent below its declared count. Alert on those rather than on any single replacement.
  • Does the replacement necessarily run on a different host from the one that died?
    No. The loop asks for one more copy; a separate placement decision picks a host with enough unreserved capacity, and that can be the same host the dead copy ran on, if the host is healthy and has room. Replacement means a new instance, not necessarily a new machine.
  • What does this model ask of an ingester that buffers a few seconds of readings in memory?
    Those readings are lost when the copy exits, so the design must tolerate it: acknowledge upstream only after the data is durable, let senders retry, or keep the buffer small enough that the loss is acceptable. The platform makes no attempt to drain or recover an instance's memory.

A rental counter does not repair the car that broke down on you; it hands you another of the same model, with an empty glovebox, a cold engine and a different plate.

saying these in an interview costs you the question

  • Assumes the dead copy is resumed with its memory and local files intact
  • Expects a page for every replacement and treats routine recovery as an incident
  • Assumes callers holding the old copy's address keep working
  • Thinks files the dead copy wrote locally come back with the replacement
  • Believes the replacement starts warm, so no latency change is expected