In a pipeline whose clients resolve schemas from a central registry, what determines whether a registry outage stops traffic?
answer
- control plane, not data plane
- steady state needs no call
- cold start empties the cache
- consumers break on another team's change
- fail open loses the guarantee durably
basics
~20 sWhether clients still need a resolution they have not already cached. Warm clients publishing and reading known schema versions keep running; cold starts, a first sighting of a new version, and register-on-publish producers all need the registry and stall or fail without it.
solid answer
~40 sTreat the registry as a control-plane dependency that clients hit rarely, not a data-plane hop per record. A client caches what it resolves, so in steady state an outage is invisible: producers keep writing with identifiers they already hold, and consumers keep decoding versions they already fetched. The outage becomes visible the moment something needs a resolution the cache lacks — a process restarting with an empty cache, a consumer meeting a version written during the outage, or a producer that registers its schema as part of starting up. Two decisions then set the blast radius: whether clients fail closed or write unchecked, and whether a deploy or scale-out event coincides with the outage, because a restart converts every warm client into a cold one at once.
go deeper
Recall that clients cache what they resolve, so a registry is contacted rarely rather than once per record.
Explain which operations genuinely need the registry — registering, learning the current version, resolving an unseen identifier — and which need nothing.
Reason about the co-occurrence that actually causes the incident: an outage during a deploy, when every process is cold and needs the registry at the same moment.
Own the fail-open decision and its asymmetry: a short outage against durable unchecked records, and who is entitled to make that trade for a given stream.
The interesting property of a registry is that it is a **control-plane** dependency pretending, at first glance, to be a data-plane one. Getting that distinction right is most of the answer. ## Two different dependencies - **Per record**: nothing. A producer writes an identifier it already holds; a consumer looks up an identifier it has already resolved. Neither touches the registry for a record whose schema version it has seen. - **Per new fact**: a call. Registering a candidate, learning the current version of a name, and resolving an identifier seen for the first time all require the registry. So the question is never 'is the registry up?' but **'does anything in flight need a fact this process does not already hold?'** ## What a client caches A client typically caches identifier-to-schema in one direction and schema-to-identifier in the other, plus the current version of the names it publishes to. Two properties of that cache decide the outage's shape: - **It is per process and in memory.** A restart empties it. Scaling out creates instances that have never held it. - **Negative results are the dangerous entry.** If a failed resolution is cached at all, it must be cached briefly, or a process poisons itself for the rest of its life over a blip. ## Producer side versus consumer side | Situation | Producer | Consumer | |---|---|---| | Steady state, known versions | Unaffected | Unaffected | | Process restart, empty cache | Cannot obtain an identifier; publishing stalls | Cannot resolve incoming identifiers; consumption stalls | | New schema version appears | Cannot register; the new deploy is blocked | Meets an unknown identifier and stalls on those records | | Registry degraded, not down | Latency on the first publish per version | Latency and retry storms on first sighting | Note the asymmetry: a producer needs the registry when **it** changes, a consumer needs it when **somebody else** changes. A consumer fleet can therefore be knocked over by a registry outage it has no involvement in, simply because a producer deployed a new version an hour earlier. ## Fail closed or fail open When a client cannot reach the registry it must choose, and the choice belongs to the platform, not to a default: 1. **Fail closed on the write path.** Refuse to publish. Correct by default: it preserves the guarantee that everything on the stream was judged, at the price of backpressure into the producing service. 2. **Fail open on the write path.** Write with some fallback shape. This silently withdraws the guarantee the whole registry exists to provide, and the damage outlives the outage because the records are durable. 3. **Fail closed on the read path.** Stall or park undecodable records rather than guessing. Nearly always right; the records are still on the stream and can be reprocessed once resolution works again. ## The subtler failure: staleness, not unavailability A cache that answers is not necessarily a cache that is right. A client holding 'the current version of this name' from before a change will keep producing the older shape long after the change was registered — no error, no alert, just a stream that quietly disagrees with the contract. An outage is loud and short; staleness is quiet and long, and it is the one worth designing a refresh policy for. ## Reducing the blast radius - **Resolve at startup, not at first record**, so a process that cannot warm its cache fails its readiness check instead of failing traffic. - **Register schemas as part of the change that introduces them**, not as a side effect of a producer booting; the registry then sits out of the restart path. - **Ship the schemas an artefact needs with the artefact** as a warm-start seed, keeping the registry as the authority rather than the only source. - **Bound negative caching and add jitter to retries**, or a restart stampede after an outage becomes a second outage. - **Stagger restarts**, because the real incident is almost always 'registry down' co-occurring with 'we were mid-deploy'. The summary to give an interviewer: in steady state the registry is off the critical path, and every mitigation is about keeping it there — resolving early, registering earlier still, caching deliberately, and never letting the read path guess.
- Why is a registry outage that coincides with a deploy so much worse than one that does not?A deploy replaces warm processes with cold ones, converting a dependency nobody was using into one every instance needs at once. The same outage is invisible in steady state and total during a rollout, which is why staggering restarts and readiness-gating on a warm cache matter more than availability numbers.
- Is failing open on the write path ever defensible?Rarely, and only for a stream whose records are low value and short lived, with a loud signal that it happened. The cost is asymmetric: an outage lasts minutes, but unchecked records written during it are durable and will be read by someone who trusted the guarantee.
- What symptom distinguishes a stale client cache from a registry outage?An outage produces errors and stalls; staleness produces none. The signature is a producer still writing an older shape, or a consumer resolving a version that is no longer current, while every health check is green — you find it by comparing what is being written against the registered head of that name.
saying these in an interview costs you the question
- Thinks every record read or written contacts the registry
- Assumes a green health check means client caches are current
- Defaults the write path to open on registry failure
- Ignores that restarts turn a warm fleet cold at once
- Caches failed resolutions for as long as successful ones