skip to content

Availability as a Dependency

A store the whole estate starts against has a topology of its own: followers that serve reads, a partition that halts issuance while plain reads continue. Asked because nothing starts without it.

on this pageshow

questions

4

Consumers report that the secret store is down — how do you tell a slow store from a read-only one and from one that cannot decrypt its own contents?

level: seniorimportance: must knowfreq 58%

answer

  1. one complaint, three different failures
  2. slow still completes the work
  3. read-only rejects fast, reads fine
  4. nothing served since this start-up
  5. different fix for each state

basics

~20 s

Three failures share one complaint. Slow means requests still complete, past the caller's timeout. Read-only means reads answer instantly while writes and issuance are rejected instantly. Unable to decrypt means every request fails alike, from the moment the process started.

solid answer

~40 s

Ask three questions in order. Has this node served anything at all since it started? If not, and every request fails identically, the store is holding ciphertext it cannot open — the protecting key was never supplied, and the clock starts at the process restart, not at the incident. If plain reads of an existing value succeed and succeed *fast*, the store is serving and only state changes are failing: read-only, usually a lost write path. If requests succeed but arrive after the caller gave up, it is slow. The tell between the middle two and the last is that rejection is **immediate and deterministic** while slowness is variable and partial. The three have unrelated fixes, so calling all of them "down" costs you the first half of the outage.

code

pseudocode · 11 lines
pseudocode
classify(node, observations):
    if not node.servedAnythingSinceStart and observations.allRequestsFailAlike:
        return CANNOT_DECRYPT_CONTENTS      # protecting key never supplied

    if observations.plainReadsFast and observations.stateChangesRejectedImmediately:
        return READ_ONLY                     # write path unreachable

    if observations.requestsCompleteAfterCallerTimeout:
        return SLOW                          # capacity, or a retry pile-up

    return KEEP_LOOKING                      # not one of the three

go deeper

for a junior

Know that a secret store can fail in more than one way and that "down" is not a diagnosis. Being able to say what you would check first — do plain reads work, do writes work — is enough at this stage.

for a middle

Explain what separates the three mechanically: a slow store still completes work, a read-only one answers reads and rejects state changes instantly, and one that cannot open its own contents has served nothing since it started.

for a senior

Show the triage under time pressure: the order you would ask the questions in, what each state leaves working for consumers, why rejection speed is the discriminator, and why the read-only case has a deadline of its own.

for a principal

Talk about what the estate is owed. Decide which of the three states the store is permitted to enter, who is paged for which, and how the operating contract tells every consumer team what still works in each of them.

## One word for three unrelated failures "The store is down" is what the estate reports, and it is never a diagnosis. Three states hide behind it. They differ in what still works, in the signal they produce, in who can end them, and in what they cost the consumers waiting on them. Working out which one you are in is most of the outage. | State | What a caller sees | What still works | What ends it | |---|---|---|---| | **Slow** | Requests complete, but past the timeout the caller set; patient callers succeed, impatient ones do not | Everything, eventually | Relieving whatever is consuming the store's capacity — real load, a retry pile-up, a starved disk | | **Read-only** | Reads of stored values answer at normal speed; writes, issuance, renewal and withdrawal are rejected the instant they arrive | Reads of what was already there | Restoring the write path — healing the partition, or making some reachable node the one that accepts writes | | **Cannot decrypt its own contents** | Every request fails the same way, and has done since this process started | Nothing | Supplying the key that protects the store, which is a subject of its own | ## Reading the symptom Three questions, in this order, because the first one is cheap and rules out the most confusing case: 1. **Has this node served anything at all since it started?** If the answer is no, and every request — including a plain read of a value you know exists — fails identically, the store is holding ciphertext it cannot open. The symptom begins at a process restart, not at a network event, and that timing is the giveaway. 2. **Do plain reads of an existing value succeed, and succeed quickly?** Then the store is serving. The failures being reported are state-changing requests, and you are in read-only. The consumers complaining are the ones that write, renew, withdraw, or ask for a credential to be generated. 3. **Do requests succeed but too late?** Latency rather than rejection is the signature of slow, and it is the only one of the three where the caller's own timeout decides whether it sees an error at all. Two consumers with different timeouts will disagree about whether there is an incident. The sharpest discriminator between the second and third rows is **how the failure arrives**. A store that has lost its write path rejects immediately and identically every time. A store under pressure fails raggedly: some calls succeed, latency moves, and the error rate tracks load. ## Why the distinction decides the response - **The remedies do not transfer.** Adding capacity does nothing for a rejected write. Restoring the write path does nothing for a store that cannot open its own contents. Supplying the protecting key does nothing for a partition. - **Retrying helps in one case and hurts in another.** Against a slow store, callers retrying are part of what is keeping it slow. Against a deterministic rejection, retrying multiplies the error count and changes nothing else. - **Consumers are hurt differently.** A process that only reads a stored value is untouched by read-only and badly hurt by slow. A process that needs a generated credential fails instantly in read-only and merely struggles in slow. - **Security actions travel with writes.** Withdrawing a credential is a write, so a read-only store cannot withdraw anything: an availability failure that blocks a security response. - **Read-only carries a hidden deadline.** Credentials already issued keep expiring and cannot be renewed, so the estate degrades further at a rate set by the shortest remaining validity in the fleet — which may be much shorter than the outage. - **Monitoring lies differently in each.** A read-only store passes a read probe. A slow store passes a probe with a generous timeout. Only the third fails everything, which makes it the easiest to see and the rarest to meet. ## What designs genuinely differ on Not every store can enter all three states. Some are unlocked at start-up by a service they call, and an operator may never see an "unable to open its contents" state at all; others require the protecting key to be supplied by people, and show that state on every unattended restart. Some stores let a follower forward a write to the node that accepts them, so read-only arrives only when that node is unreachable from everywhere; others reject at whichever node was asked. The triage above survives all of it, because it is built on what a caller observes rather than on any internal design. ## What good sounds like in an interview Give each state a one-line signature, then say what you would do in the first five minutes: freeze the rate of change in the estate, establish which of the three you are in from the read path and the restart time, and tell consumers which of their operations still works. A candidate who jumps straight to a fix has not yet said which failure they are fixing.

  • Two teams disagree about whether there is an incident at all. Which of the three states does that point at?
    Slow. It is the only one where the caller's own timeout decides whether it sees a failure, so a team with generous timeouts sees elevated latency while a team with tight ones sees outright errors. Read-only and a store that cannot open its contents both reject deterministically, and every caller agrees about them immediately.
  • The store went read-only forty minutes ago and the estate is still serving traffic. What are you actually racing?
    The expiry of credentials that were already issued. Renewal is a write, so nothing in the fleet can extend what it holds, and consumers fall over one by one as their own validity windows end. Your deadline is the shortest remaining validity anywhere, not the length of the partition, so the first thing worth knowing is which consumer's window closes first.
  • Why is 'add more capacity' a reasonable first move in one of these states and a wasted one in the other two?
    Capacity only addresses slowness, where requests are completing and the constraint is throughput. A store that has lost its write path rejects instantly at any scale, and a store that cannot open its own contents serves nothing however many nodes you give it. Capacity changes the number of requests handled, not whether an operation is possible.

A shop can be understaffed (everyone is served, far too late), open with the stockroom locked (you may take what is on the shelves, but nothing new is brought out), or shuttered because nobody has the keys. All three get reported as "the shop is closed", and each one needs a different person called.

saying these in an interview costs you the question

  • Reports 'the store is down' as if that were a diagnosis
  • Assumes a rejected write means the store is overloaded
  • Thinks a store answering reads cannot be in trouble
  • Retries hard against a store that rejects instantly and deterministically
  • Treats a store that cannot open its contents as a capacity problem
  • Misses that a read-only store cannot withdraw a credential
open as a page

During a network partition, why does a secret store's follower keep answering reads of a stored value while requests to generate a new credential fail?

level: middleimportance: should knowfreq 42%

basics

~20 s

Reading a stored value is served from the copy the follower already holds; generating a credential is a write — the store must record the new credential and its expiry before returning it, and a follower has no write path.

open as a page

Why does a secret store partition that steady-state services barely notice break a service scaling out and a fleet on a rolling restart?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The store is a start-up dependency, not a per-request one. Already-running processes hold what they fetched and never call it, so a degraded store is invisible until something new starts — and scaling out and rolling restarts are the two activities that manufacture new start-ups.

open as a page

Every service now starts against one secret store — how do you decide how much of the estate's availability to stake on it?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Two levers exist: make the store more available, or make the estate need it less. Decide which operations must survive a partition — plain reads — and which may fail, then publish that as the contract every consumer designs against.

open as a page