After a cold restart, a store cannot unwrap its protecting key because the credential for that call is itself stored inside it — how do you break the loop?
answer
- opening it needs something not inside it
- the chain must terminate outside the store
- rolling restarts never expose this
- cold start is the only honest test
- node identity from the platform, not the store
basics
~20 sThe input that opens a store must be obtainable while the store is closed. Move the unwrap credential to something the node holds independently — an identity its platform signs for the machine, or material its hardware holds — and rehearse a genuinely cold start.
solid answer
~40 sThe loop is that opening the store requires a call, the call requires a credential, and the credential is a record the store will only release once it is open. Nothing resolves it from inside; the chain has to terminate in something that exists while the store is shut. In practice that means the node proves its own identity to the unwrap path using something it did not get from the store — an identity document its platform signs for the machine, key material the host's hardware holds, or a bounded credential placed by the process that provisions the node. Then draw the start-up dependency list explicitly and enforce one rule: nothing on it may itself depend on the store. Finally, rehearse a cold start, because this defect is invisible in every other test.
code
pseudocode · 12 lineson start_up:
wrapped_root_key = read_local_blob() # ciphertext, useless on its own
cred = load_unwrap_credential() # the question is: from where?
if cred.source == THIS_STORE:
halt("cold start impossible: unwrap credential lives inside the store")
if not key_service.reachable():
halt("cold start blocked: unwrap path unavailable")
root_key = key_service.unwrap(wrapped_root_key, cred)
begin_serving(root_key)go deeper
Recall that a store has to be opened before it can hand anything out, so anything needed to open it cannot be one of the things it holds.
Walk the four steps of the cycle out loud and name where it must terminate: an input the node can obtain while the store is closed.
Demonstrate you would find this before an incident — draw the start-up dependency list, run a genuinely cold start, and log every call made before the first read is served.
Own the rule that the unwrap path must sit in a different failure domain from the store, and make cold-start rehearsal a scheduled obligation rather than a one-off verification.
## The shape of the loop A store that delegates unwrapping keeps the key protecting its contents on disk in wrapped form and asks something else to open it — a device that will not export keys, or a separate key service. To make that call it must authenticate. The loop appears the moment somebody answers the question *where does the credential for that call live?* with *in the store*, which is a very natural answer, because the store is where credentials live. Now the sequence cannot complete: 1. The store needs the protecting key before it can read any record. 2. Getting the protecting key needs a call to the unwrap path. 3. That call needs a credential. 4. Reading that credential needs the store to be open — which is step 1. No amount of retrying resolves a cycle. The chain that opens a store must **terminate outside it**, in something that exists while the store is shut. ## Why it hides until the worst possible moment This defect is invisible in almost every test anyone runs, for a reason worth stating plainly: the credential is readable exactly when the store is already open. - **Rolling restarts pass.** While one node is still serving, the credential is obtainable, so the restarted node opens normally and the dependency looks satisfied. - **Single-node restarts pass** if anything else can still answer for the store. - **The configuration review passes**, because every line of it is individually sensible: the store holds credentials, and this is a credential. - **It fails only on a cold start** — the whole cluster down, a rebuilt host, a new environment — which is the one scenario nobody rehearses and the one that arrives at 2am after a host failure. The variant that catches experienced teams is indirect. The unwrap path authorises the store by an identity issued somewhere else, and *that* system's own credential is in the store. The cycle is two hops long instead of one and survives review for years. ## Breaking it The fix is not clever, it is structural: give the node something to authenticate with that it did not get from the store. - **An identity the platform signs for the machine.** The environment the node runs in vouches for it, the unwrap path trusts that issuer, and nothing about the exchange requires the store to be open. - **Material the host's hardware holds.** The node can prove itself locally, and the wrapped key on disk stays useless on any other host. - **A bounded credential placed at provisioning time** by whatever builds the node, with a short life and a single permitted use. Weaker than the first two, and acceptable where the alternative is the cycle. - **A quorum of shares as the fallback path**, accepting that the cold start is attended while the warm path stays automatic. What unites them: each is available with the store closed. ## The dependency you took on instead Breaking the cycle does not remove the dependency, it relocates it. The unwrap path is now on the store's start-up critical path, so it inherits a rule: **it must not depend on the store, and it should not fail for the same reason the store failed.** Write the start-up list down — the wrapped key's storage, the thing that proves the node's identity, the unwrap service — and check each entry against both conditions. A delegated unwrap that runs in the same failure domain as the store converts one outage into an unopenable one. ## Proving it, since reasoning is not enough The only convincing evidence is a cold start where nothing is serving: 1. Build a node, or a whole environment, from scratch with no member of the store's cluster running. 2. Let it start with no operator supplying anything by hand. 3. Record every outbound call it makes before it serves its first read. 4. Check that list against the rule: nothing on it may resolve through the store. Step 3 is where the surprises are. Teams routinely discover a call they did not know the start-up made, and a credential path they assumed was local. Run it on a schedule, not once — the loop is usually reintroduced later by somebody doing the obviously right thing and moving a credential into the store.
- Why does a rolling restart never expose this?Because at least one node stays open, and the unwrap credential is readable from it. The cycle only closes when nothing is serving, so every routine restart, deploy and node replacement succeeds. The defect is discovered on the first genuine cold start, which is usually an incident rather than a test.
- What belongs on the store's start-up dependency list?Everything that must already answer before the store serves its first read: the storage holding the wrapped key, whatever proves the node's identity, and the unwrap path itself. Two rules apply to every entry — none may resolve through the store, and none should share the failure that took the store down.
- Is putting the unwrap credential on the host's disk an acceptable fix?It breaks the cycle, and it is the weakest option: anything that can read the host can now open the store's contents, so the wrapped key on disk gains little. It is defensible as a bounded, provisioning-time credential with a short life and one permitted use, and poor as a permanent arrangement.
The key to the office is locked in the office safe. Every day it works, because somebody already inside can fetch it — and it works right up to the first morning nobody is inside.
saying these in an interview costs you the question
- Suggests retrying the unwrap call until the store comes up.
- Says a rolling restart proves the start-up path works.
- Treats the store as the right home for its own unwrap credential.
- Misses the two-hop version, where the identity issuer depends on the store.
- Ignores that the unwrap service is now a start-up dependency itself.