After a full outage, why do the copies of a data-owning workload come back one at a time in a fixed order?
answer
- the group needs a first member
- all at once means nobody joins
- each start gated on the previous reporting ready
- reverse on the way down
- serial bring-up costs recovery time
basics
~20 sOrdered bring-up gives the group a first member. Each copy starts only after the previous one reports ready, so later copies join something that already exists instead of several copies each initialising their own empty data set. Shutdown runs in reverse.
solid answer
~50 sA data-owning group is not a set of equals that can appear in any order - at least during bring-up, one member has to exist before the next one can attach to it. If every copy starts at once after a full outage, each finds nothing, each takes the initialise-from-empty path its start-up logic offers, and you end up with several members that each believe they are the group. Ordered start-up makes the platform start the lowest-numbered copy, wait for it to report ready, and only then start the next, so every later copy finds an existing member to join. Shutdown runs in reverse for the mirror-image reason: the member others were started against is the last one still running, and each copy gets its termination signal and grace period on its own rather than all of them at once.
code
pseudocode · 10 lineson bring_up(workload):
for index in 0 .. workload.copies - 1:
start copy[index]
wait until copy[index].readyCheck passes # not merely running
# copy[index + 1] now finds a live member and takes the join path
on shut_down(workload):
for index in workload.copies - 1 .. 0:
send termination signal to copy[index]
wait until exited or gracePeriod expiredgo deeper
Recall that some workloads must be started in a set order rather than all at once, and that the platform can be told to do this.
Explain the sequencing guarantee: each copy waits for the previous one to report ready, so later copies join an existing member instead of each initialising from empty. Say why shutdown reverses it.
Diagnose the case where ordering was declared and the split still happened, by reading what the readiness signal actually measures, and price the serial cold start against the copy count.
Decide which workloads are allowed to need ordering at all, since it converts a partial-capacity outage into a no-progress one and grows recovery time linearly with width.
## What parallel start does to a group that owns data For an interchangeable workload, starting all copies at once is the right behaviour - it is the fastest path back to full capacity, and no copy cares whether the others exist. For a workload whose copies own data, the same behaviour produces a specific, ugly failure. Every data engine's start-up logic has to answer one question: *am I joining something, or am I the first?* It answers it by looking - for an existing member to contact, or for its own data on disk. After a full outage, if all copies are started simultaneously, each of them looks at the same moment, each sees no reachable peer, and each takes the branch for "I am the first". The result is not a crash. It is several copies, all healthy, each convinced it is the group, each accepting writes into its own data. Reconciling that afterwards is a data-recovery exercise, not an operational one. ## The guarantee ordered start-up actually makes Ordered bring-up is a promise about *sequencing*, not about speed: - The copies have a fixed order derived from their durable per-copy names: the lowest first, then the next, and so on. - A copy is not started until the previous one reports **ready** - not merely started, not merely running. - Therefore every copy after the first begins its start-up with at least one live member it can find, and takes the "join" branch rather than the "initialise" branch. That second bullet is where the guarantee is won or lost. The ordering is only as strong as the signal it waits on. If the readiness signal returns success as soon as the process has bound a port - before it has opened its data, replayed anything pending and made itself joinable - then the platform starts copy 1 against a copy 0 that is not yet answerable, and you are back to the parallel failure with extra steps and a longer timeline. When ordered start-up "did not work", the readiness definition is the first thing to read. ## Shutdown in reverse Stopping runs the order backwards: the highest-numbered copy is stopped first and the lowest last. | | Start | Stop | |---|---|---| | Direction | lowest name first | highest name first | | Gate between steps | previous copy reports ready | previous copy has exited, or its grace period expired | | What it protects | later copies find an existing member | the member others depend on outlives them | | Failure if skipped | several copies initialise their own data | in-flight writes cut off mid-flush, all at once | The second column matters more than teams expect. Each copy is sent a termination signal and given a grace period to finish in-flight work and close its data cleanly; taking them down one at a time means that work is serialised and observable, rather than a simultaneous scramble where one slow copy's grace period expires while everything else is already gone. ## The costs you accept 1. **Recovery time becomes serial.** Bring-up is now the sum of the copies' start times plus the readiness waits, not the slowest single start. Nine copies that each take two minutes to become ready are an eighteen-minute cold start, and the number grows linearly with the copy count - a real argument against very wide data-owning groups. 2. **One stuck copy stops the line.** If copy 1 never reports ready, copy 2 is never started. The platform is doing exactly what you asked; the group is simply parked. Ordered bring-up therefore converts a partial-capacity problem into a no-progress problem, which is easier to notice but harder to route around. 3. **The order is fixed, not clever.** The sequence follows the names, not which copy happens to hold the most recent data. If the member that was furthest ahead is not the first one in the order, an engine that trusts whoever starts first can come back on a stale branch. Engines that handle this do so themselves, by comparing data on start-up rather than by trusting arrival order - the platform's ordering cannot substitute for that. ## Where ordering does not apply Not every data-owning workload needs it. A single-copy workload has no ordering question at all. A group whose members are independent - each owning a disjoint slice with no need to contact each other at start-up - can come back in parallel safely, and paying serial bring-up for it is a pure loss. The rule is about start-up *dependencies* between members, not about owning data as such: ask whether a copy's start-up behaviour changes depending on whether another member is reachable. If it does, order the bring-up. If it does not, do not.
- A workload declares ordered start-up and still comes back with two members that each initialised from empty. What do you look at first?The readiness definition. Ordered bring-up waits for the previous copy to report ready, so if that signal fires when the process binds a port rather than when its data is open and it can be joined, the next copy is started against a member that cannot answer yet. The ordering held; the signal it was gated on was meaningless.
- When is serial bring-up the wrong choice for a workload that owns data?When the members have no start-up dependency on each other - each owns a disjoint slice and never needs to contact a peer to start correctly. Then ordering buys nothing and costs a cold start that grows linearly with the copy count. A single-copy workload is the same case: there is no order to impose.
saying these in an interview costs you the question
- Thinks every workload gets ordered start-up by default
- Says parallel start is only a resource-contention problem
- Calls the ordering decoration because copies sort themselves out
- Treats an open port as the copy being ready
- Stops every copy at once and still expects a clean flush