skip to content

In a pipeline whose clients resolve schemas from a central registry, what determines whether a registry outage stops traffic?

level: seniorimportance: should knowfreq 48%

answer

  1. control plane, not data plane
  2. steady state needs no call
  3. cold start empties the cache
  4. consumers break on another team's change
  5. fail open loses the guarantee durably

basics

~20 s

Whether clients still need a resolution they have not already cached. Warm clients publishing and reading known schema versions keep running; cold starts, a first sighting of a new version, and register-on-publish producers all need the registry and stall or fail without it.

solid answer

~40 s

Treat the registry as a control-plane dependency that clients hit rarely, not a data-plane hop per record. A client caches what it resolves, so in steady state an outage is invisible: producers keep writing with identifiers they already hold, and consumers keep decoding versions they already fetched. The outage becomes visible the moment something needs a resolution the cache lacks — a process restarting with an empty cache, a consumer meeting a version written during the outage, or a producer that registers its schema as part of starting up. Two decisions then set the blast radius: whether clients fail closed or write unchecked, and whether a deploy or scale-out event coincides with the outage, because a restart converts every warm client into a cold one at once.

go deeper

for a junior

Recall that clients cache what they resolve, so a registry is contacted rarely rather than once per record.

for a middle

Explain which operations genuinely need the registry — registering, learning the current version, resolving an unseen identifier — and which need nothing.

for a senior

Reason about the co-occurrence that actually causes the incident: an outage during a deploy, when every process is cold and needs the registry at the same moment.

for a principal

Own the fail-open decision and its asymmetry: a short outage against durable unchecked records, and who is entitled to make that trade for a given stream.

The interesting property of a registry is that it is a **control-plane** dependency pretending, at first glance, to be a data-plane one. Getting that distinction right is most of the answer. ## Two different dependencies - **Per record**: nothing. A producer writes an identifier it already holds; a consumer looks up an identifier it has already resolved. Neither touches the registry for a record whose schema version it has seen. - **Per new fact**: a call. Registering a candidate, learning the current version of a name, and resolving an identifier seen for the first time all require the registry. So the question is never 'is the registry up?' but **'does anything in flight need a fact this process does not already hold?'** ## What a client caches A client typically caches identifier-to-schema in one direction and schema-to-identifier in the other, plus the current version of the names it publishes to. Two properties of that cache decide the outage's shape: - **It is per process and in memory.** A restart empties it. Scaling out creates instances that have never held it. - **Negative results are the dangerous entry.** If a failed resolution is cached at all, it must be cached briefly, or a process poisons itself for the rest of its life over a blip. ## Producer side versus consumer side | Situation | Producer | Consumer | |---|---|---| | Steady state, known versions | Unaffected | Unaffected | | Process restart, empty cache | Cannot obtain an identifier; publishing stalls | Cannot resolve incoming identifiers; consumption stalls | | New schema version appears | Cannot register; the new deploy is blocked | Meets an unknown identifier and stalls on those records | | Registry degraded, not down | Latency on the first publish per version | Latency and retry storms on first sighting | Note the asymmetry: a producer needs the registry when **it** changes, a consumer needs it when **somebody else** changes. A consumer fleet can therefore be knocked over by a registry outage it has no involvement in, simply because a producer deployed a new version an hour earlier. ## Fail closed or fail open When a client cannot reach the registry it must choose, and the choice belongs to the platform, not to a default: 1. **Fail closed on the write path.** Refuse to publish. Correct by default: it preserves the guarantee that everything on the stream was judged, at the price of backpressure into the producing service. 2. **Fail open on the write path.** Write with some fallback shape. This silently withdraws the guarantee the whole registry exists to provide, and the damage outlives the outage because the records are durable. 3. **Fail closed on the read path.** Stall or park undecodable records rather than guessing. Nearly always right; the records are still on the stream and can be reprocessed once resolution works again. ## The subtler failure: staleness, not unavailability A cache that answers is not necessarily a cache that is right. A client holding 'the current version of this name' from before a change will keep producing the older shape long after the change was registered — no error, no alert, just a stream that quietly disagrees with the contract. An outage is loud and short; staleness is quiet and long, and it is the one worth designing a refresh policy for. ## Reducing the blast radius - **Resolve at startup, not at first record**, so a process that cannot warm its cache fails its readiness check instead of failing traffic. - **Register schemas as part of the change that introduces them**, not as a side effect of a producer booting; the registry then sits out of the restart path. - **Ship the schemas an artefact needs with the artefact** as a warm-start seed, keeping the registry as the authority rather than the only source. - **Bound negative caching and add jitter to retries**, or a restart stampede after an outage becomes a second outage. - **Stagger restarts**, because the real incident is almost always 'registry down' co-occurring with 'we were mid-deploy'. The summary to give an interviewer: in steady state the registry is off the critical path, and every mitigation is about keeping it there — resolving early, registering earlier still, caching deliberately, and never letting the read path guess.

  • Why is a registry outage that coincides with a deploy so much worse than one that does not?
    A deploy replaces warm processes with cold ones, converting a dependency nobody was using into one every instance needs at once. The same outage is invisible in steady state and total during a rollout, which is why staggering restarts and readiness-gating on a warm cache matter more than availability numbers.
  • Is failing open on the write path ever defensible?
    Rarely, and only for a stream whose records are low value and short lived, with a loud signal that it happened. The cost is asymmetric: an outage lasts minutes, but unchecked records written during it are durable and will be read by someone who trusted the guarantee.
  • What symptom distinguishes a stale client cache from a registry outage?
    An outage produces errors and stalls; staleness produces none. The signature is a producer still writing an older shape, or a consumer resolving a version that is no longer current, while every health check is green — you find it by comparing what is being written against the registered head of that name.

saying these in an interview costs you the question

  • Thinks every record read or written contacts the registry
  • Assumes a green health check means client caches are current
  • Defaults the write path to open on registry failure
  • Ignores that restarts turn a warm fleet cold at once
  • Caches failed resolutions for as long as successful ones